IMAGE PROCESSING SYSTEM, IMAGE PROCESSING METHOD, AND PROGRAM
The image processing system enhances tilt correction accuracy for document images with mixed handwritten and printed characters by using a neural network to separate handwritten characters and estimate tilt angles, addressing the limitations of existing technologies.
Patent Information
- Application Number
- JP2021067356
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-04-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-04-12
AI Technical Summary
Existing tilt correction technologies struggle with accurately estimating the tilt angle of document images containing both handwritten and printed characters, especially when there are many handwritten characters with variations in line spacing, pitch, and angle, or when there is no ruled line information.
An image processing system that uses a neural network to separate handwritten characters from document images, generating an image without handwritten characters for tilt angle estimation, and then performs tilt correction based on the estimated angle.
Improves the accuracy of tilt correction for document images with mixed handwritten and printed characters by effectively isolating the influence of handwritten characters on tilt angle estimation.
Smart Images

Figure 0007689439000001 
Figure 0007689439000002 
Figure 0007689439000003
Abstract
Description
[Technical field]
[0001] The present invention relates to an image processing system, an image processing method, and a program for performing tilt correction on a document image containing both handwritten characters and printed characters. [Background technology]
[0002] Conventionally, there is a technology that performs optical character recognition processing (hereinafter, OCR processing) on document image data scanned by an image forming device to extract character strings in the image data as digital data. However, if the scanned document image is tilted, correct OCR processing may not be possible. Therefore, there is a technology (hereinafter, tilt correction) that estimates the tilt angle of the document image as a preprocessing of the OCR processing and corrects it to the correct angle (for example, Patent Document 1, Patent Document 2, and Patent Document 3).
[0003] The technology described in Patent Document 1 measures the variance of pixels as a function of the rotation angle of the document image, and performs skew correction at the document rotation angle (skew angle) where the variance is maximum. In addition, the technology described in Patent Document 2 detects table areas, and then performs skew correction of the input image based on the skew of ruled lines. In addition, the technology described in Patent Document 3 detects the edges of the document image, thereby performing skew correction without checking the contents of the image. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 3-268189 [Patent Document 2] Japanese Patent Application Publication No. 8-44822 [Patent Document 3] JP 2020-53931 A Summary of the Invention [Problem to be solved by the invention]
[0005] However, in Patent Document 1, when there are many handwritten characters with variations in line spacing, pitch, and angle, it is difficult to estimate an accurate tilt angle. In particular, when the ratio of handwritten characters to type characters is high, or when the handwritten characters are dense (have a large luminance difference from the type characters), there is a risk that an accurate tilt angle cannot be estimated. In addition, in Patent Document 2, when there is no ruled line information in the document image, there is a risk that tilt correction cannot be performed. In addition, in Patent Document 3, when edges cannot be detected or when the document manuscript is not square (torn, etc.), there is a risk that accurate tilt correction cannot be performed.
[0006] The present invention has been made in consideration of the above circumstances, and has an object to provide an image processing system that improves the accuracy of tilt correction for document images that contain both handwritten characters and printed characters. [Means for solving the problem]
[0007] In order to achieve the above object, an image processing system according to the present invention performs skew correction on a document image, the image processing system comprising: a document image acquisition unit that acquires a document image including a mixture of handwritten characters and printed characters; Using a neural network that has learned the characteristics of handwritten characters from an image including handwritten characters, pixels of handwritten characters contained in the document image are determined, and the determined pixels of handwritten characters are From the document image Removal By doing so, The above The document image processing apparatus includes a separation unit that generates an image other than handwritten characters, a tilt angle estimation unit that estimates a tilt angle using the generated image other than handwritten characters, and a tilt correction unit that corrects the tilt of the document image in which handwritten characters and printed characters are mixed based on the estimated tilt angle. It is characterized by: Effect of the Invention
[0008] According to the present invention, it is possible to improve the accuracy of tilt correction for a document image containing both handwritten characters and printed characters. [Brief description of the drawings]
[0009] [Figure 1] FIG. 1 is a block diagram showing an example of an image processing system according to an embodiment of the present invention. [Diagram 2] 5 is a flowchart showing the procedure of image processing in the first embodiment. [Diagram 3] 5A to 5C are views showing an example of a document image and a processing result in the first embodiment. [Figure 4] 10 is a flowchart showing the procedure of image processing in a second embodiment. [Diagram 5] FIG. 11 is a diagram showing an example of learning data for handwritten character separation in the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] Hereinafter, the embodiments of the present invention will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the present invention, and not all of the combinations of features described in the embodiments are essential to the solution of the present invention. Note that the same components are given the same reference numbers and descriptions are omitted.
[0011] <Image formation system overview> Fig. 1 is a block diagram showing an example of an image processing system according to an embodiment of the present invention. As shown in Fig. 1, the image processing system includes an image forming apparatus 100, a host computer 170, and a server 191 (which may be a cloud server).
[0012] In this embodiment, the image forming apparatus 100 will be described as a multifunction printer (MFP) that integrates multiple functions such as a print function, a read function, and a FAX function. The server 191 will be described as having a document management function. The image forming apparatus 100, the host computer 170, and the server 191 are connected to a network such as a LAN (Local Area Network) 190 so that they can communicate with each other. A plurality of image forming apparatuses 100, the host computers 170, and the server 191 may be connected, or other devices may be connected. In this embodiment, the network will be described as a LAN 190, but it may be a wired network, a wireless network, or a combination of these.
[0013] The image forming apparatus 100 includes a control device 110, a reader device 120, a printer device 130, an operation unit 140, and a storage device 150. The control device 110 is connected to each of the reader device 120, the printer device 130, the operation unit 140, and the storage device 150.
[0014] The control device 110 is a control board (controller) that performs overall control of the image forming apparatus 100. The control device 110 includes a CPU 111, a ROM 112, a RAM 113, and an image processing unit 114.
[0015] The CPU 111 controls each block in the control device 110 via a system bus (not shown). For example, the CPU 111 executes the functions of the image forming device 100 by reading and executing a program stored in the ROM 112, the RAM 113, the storage device 150, or another storage medium.
[0016] The ROM 112 stores, for example, a control program, and tables and setting data necessary for executing the functions of the image forming apparatus 100. The RAM 113 is used as a work memory for the CPU 111, for example.
[0017] The image processing unit 114 performs various image processing such as conversion, correction, editing, compression / decompression, etc., on the read image data generated by the reader device 120 and image data received from the outside. The image processing unit 114 may be configured as hardware or may be realized as software.
[0018] The reader device 120 has a scanner engine configuration, performs document scanning processing to optically read a document, and generates read image data (document image) from the optically read document. The document scanning processing may be a method of optically reading a document set on a document table, or a method of optically reading a document fed from an automatic document feeder (ADF).
[0019] The printer device 130 has a printer engine configuration compatible with various recording methods such as the inkjet recording method, electrophotography, etc. In this way, the printer device 130 forms an image on a recording medium.
[0020] The operation unit 140 includes operation keys for receiving user operations, and a liquid crystal panel for performing various settings, displaying a user interface screen, etc. The operation unit 140 outputs information received through user operations, etc., to the control device 110.
[0021] The storage device 150 stores user information, such as image data, device information such as modes and licenses, an address book, and customization.
[0022] 1, and may include other components depending on the executable functions of the image forming apparatus 100. For example, the image forming apparatus 100 may include a component required to execute a FAX function or a component enabling short-distance wireless communication.
[0023] The server 191 includes a control device 198, an operation unit 195, a storage device 196, and a display unit 197. The control device 198 is connected to each of the operation unit 195, the storage device 196, and the display unit 197.
[0024] The control device 198 is a control board (controller) that performs overall control of the server 191. The control device 198 includes a CPU 192, a ROM 193, and a RAM 194.
[0025] The CPU 192 controls each block in the control device 198 via a system bus (not shown). For example, the CPU 192 executes the functions of the server 191 by reading and executing a program stored in the ROM 193, the RAM 194, the storage device 196, or another storage medium.
[0026] The ROM 193 stores, for example, various control programs such as an operating system program (OS), and tables and setting data necessary for executing the functions of the server 191. The RAM 194 is used as a work memory for the CPU 192, for example.
[0027] The operation unit 195 includes a keyboard, a pointing device, etc. for receiving user operations, and outputs information on the received user operations, etc. to the control device 198. The storage device 196 stores, for example, various application programs, data, user information, device information, etc. The display unit 197 is, for example, a liquid crystal display, and displays various user interface screens and information.
[0028] The host computer 170 is connected to the image forming apparatus 100 and the server 191 via a LAN 190. With this configuration, the image forming apparatus 100 and the server 191 can also be operated based on operations and instructions from the host computer 170.
[0029] A specific embodiment will be described below by taking the image processing system configured as above as an example. Note that "handwritten characters" used in the following embodiment refer to characters handwritten and input by a human being.
[0030] [First embodiment] In the case of a document that contains both handwritten and printed characters, conventional tilt correction may fail due to the influence of handwritten characters with uneven character spacing and pitch. In this embodiment, when handwritten characters are present, an image for tilt angle estimation is generated that excludes the influence of the handwritten characters, and tilt angle estimation is performed.
[0031] Fig. 2 is a flowchart showing the procedure of image processing in the first embodiment. Fig. 3 is a diagram showing an example of an input document image and a processing result in the first embodiment. In the following description, "tilt" refers to the tilt angle of the input document image with respect to a reference line L (see Fig. 3) that extends in the left-right direction and serves as the basis for tilt.
[0032] The image processing procedure will be described below with reference to Fig. 2, and as necessary, the image processing procedure will be described with reference to Fig. 3. The processing in Fig. 2 is realized, for example, by the CPU 111 reading a program stored in the ROM 112 into the RAM 113 and executing the program.
[0033] In step S201, an input document image acquisition process is performed. The input document image is a document image 300 input to the image processing system. In the input document image acquisition process, when the CPU 111 receives an instruction for document scanning process from a user via the operation unit 140, the CPU 111 instructs the reader device 120 to scan and executes scanning. As a result, read image data (document image) corresponding to the document is acquired. The document image 300 shown in FIG. 3 is an example of an input document image acquired in the document image acquisition process in step S201. If the document is set at an angle on the document table, the document image 300 acquired by the reader device 120 will also be inclined. Even if the document is read using the ADF, the acquired document image may be inclined due to the way the document is set or the speed difference between the left and right sides of the conveying motor. Therefore, it is necessary to identify and correct the inclination of the acquired document image. In this way, the CPU 111 functions as a document image acquisition unit that acquires a document image in the image processing system.
[0034] In step S202, handwritten character separation processing is performed. In the handwritten character separation processing, CPU 111 performs processing to separate the portions of handwritten characters from the read image data generated in step S201. This generates an image of handwritten characters and an image of non-handwritten characters. The image of non-handwritten characters generated here is used for estimating the tilt angle after this separation processing. In this way, CPU 111 functions as a handwritten character separation unit in the image processing system that separates an image of handwritten characters determined to be handwritten characters from an image of non-handwritten characters that is not determined to be handwritten characters.
[0035] In the handwritten character separation method of this embodiment, first, a neural network (NN) is made to learn handwritten character regions and other background regions in an image. Next, based on the learning of the neural network, it is judged whether each pixel is handwritten or not. As a result, if there is a match with the image characteristics of handwritten characters, it is judged as a handwritten character and it is possible to extract the pixel. For example, by performing this process on the scanned image data of the document image 300 in FIG. 3, pixels such as those shown by pixels 310-312 are judged as handwritten characters. Next, the pixels 310-312 judged as handwritten characters are removed to obtain a document image 301 other than handwritten characters. Note that the handwritten character separation method of this embodiment is merely an example, and the method for separating handwritten characters is not limited to the method of this embodiment.
[0036] Conventional handwritten character separation was aimed at inputting the data into OCR processing specialized for each character type. Note that OCR processing refers to the extraction of character data using optical character recognition (OCR).
[0037] In contrast, in this embodiment, the handwritten characters are separated so as not to interfere with the inclination recognition. In other words, the handwritten characters are removed by this separation process to generate an image to be used for the subsequent inclination angle estimation. By using an image from which handwritten characters have been removed, such as the document image 301 other than handwritten characters, it is expected that the estimation accuracy of the inclination angle can be improved.
[0038] In step S211, CPU 111 determines whether handwritten characters are mixed in with the scanned document image 300. In step S202, if the number of pixels extracted as handwritten characters is equal to or greater than a certain amount, it is determined that handwritten characters are present (Yes), and processing proceeds to step S202. On the other hand, in step S202, if the number of pixels extracted as handwritten characters is less than the certain amount, it is determined that handwritten characters are not present (No), and processing proceeds to step S212.
[0039] In this embodiment, in determining whether or not there is a handwritten character in step S211, the presence or absence of a handwritten character is determined based on the ratio of pixels in the image of the handwritten character separated in step S202 to pixels in the image other than the handwritten character. If the number of pixels extracted as handwritten characters is less than a certain amount, there is a high possibility that the image is noisy, and there is almost no effect on the inclination angle estimation. Alternatively, even if a handwritten character is truly extracted, there is almost no effect on the inclination angle estimation as long as the number of pixels other than the handwritten character, such as a printed character, is more than a certain percentage. Therefore, when determining whether or not there is a handwritten character, it is determined that there is a handwritten character (Yes) if the number of pixels in the image of the handwritten character separated in step S202 is more than a certain percentage of the pixels in the image other than the handwritten character.
[0040] In step S203, a tilt angle estimation process is performed. In the tilt angle estimation process, CPU 111 estimates the tilt angle using document image 301 other than handwritten characters generated in step S202. By excluding handwritten characters with variations in line spacing, pitch, and angle and performing tilt angle estimation using document image 301 other than handwritten characters, the accuracy of the tilt angle estimation is improved. In this way, CPU 111 functions as a tilt angle estimation unit in the image processing system that estimates the tilt angle of an image other than handwritten characters.
[0041] The method for estimating the skew angle (rotation angle) used in this embodiment utilizes the fact that character strings and lines in a document image are aligned horizontally in the data before printing. For example, the skew angle can be estimated by taking projection histograms in various directions and selecting the angle corresponding to the histogram in which the peaks and bottoms of the histogram oscillate greatly in a short cycle. This is because, if the projection is in the correct direction, horizontal lines such as character strings on the same line and ruled lines in the same direction are voted for the same bin on the histogram, and nothing is voted for the parts between the lines, so that large amplitudes occur in the cycle between characters.
[0042] The angles estimated by the methods described above do not take into account the orientation of the characters, and there is an uncertainty of 180 degrees. The orientation of the characters can be determined using the likelihood information of the characters obtained by performing a simple character recognition process. This makes it possible to calculate angle information that also takes into account the orientation of the characters. This skew angle estimation method is effective for documents that are mainly composed of type and ruled lines, in which the line spacing is uniform, the inter-line gap is greater than a predetermined gap, and the horizontal strokes are horizontal. Therefore, in documents based on type, such as the document image 301 other than handwritten characters, the skew angle can be accurately determined. The horizontal direction mentioned above refers to a direction parallel to the reference line L in FIG. 3. The reference line L is a line that extends in the left-right direction on the paper surface and is perpendicular to the up-down direction on the paper surface.
[0043] In this embodiment, by performing the inclination angle estimation process on the document image 301 other than the handwritten characters, it is possible to obtain the inclination angle α with respect to the reference line L. However, the method for identifying the inclination angle of the image is not limited to a specific method.
[0044] In step S212, the CPU 111 performs a tilt angle estimation process on the document image determined to have no handwritten characters in step S211. The tilt angle estimation process is similar to the process performed in step S203.
[0045] In step S213, it is determined whether the document image is skewed based on the skew angle estimated in steps S203 and S212. If the skew angle is equal to or greater than a certain angle, it is determined that the document image is skewed (Yes), and the process proceeds to skew correction processing in step S204. On the other hand, if the skew angle is less than the certain angle, it is determined that the document image is not skewed (No), skew correction is skipped, and the process proceeds to OCR processing in step S205.
[0046] In step S204, CPU 111 performs a tilt correction process on the document image acquired in S201 using the tilt angle estimated in step S203 or step S212. In the tilt correction in this embodiment, rotation coordinate conversion is performed using the tilt angle estimated in step S203 and step S212. Note that the correction means is not limited to this. In this embodiment, the tilt correction process is performed on document image 301 other than handwritten characters using tilt angle α shown in FIG. 3, thereby obtaining a corrected image 302 after tilt correction. After the tilt correction process, the process proceeds to OCR process in step S205. In this way, CPU 111 functions as a tilt correction unit that corrects document image 300 based on tilt angle α in the image processing system.
[0047] In step S205, the CPU 111 performs OCR processing on the corrected image 302 corrected in step S204. In this embodiment, OCR processing specialized for handwritten characters and printed characters is performed on the handwritten characters separated in step S202 and the document image 301 other than the handwritten characters, respectively. Then, a process is performed to merge the OCR result of the handwritten characters and the OCR result of the document image 301 other than the handwritten characters.
[0048] In addition, in this embodiment, character string areas are determined before OCR processing, and OCR processing is performed on each area that is a character string area to obtain the character code of the character string in the character string area. This area determination eliminates the need to process areas other than the character string area. As a result, the processing load can be reduced and the accuracy of character recognition can be improved. Note that various methods have been devised for OCR processing, and the method is not limited to the method of this embodiment.
[0049] In step S206, CPU 111 adds the text information obtained in step S205 to document image 300 and corrected image 302, registers the data in storage device 150, and ends this process. When registering the data, the document image may not be stored as image data, but may be converted into a document format such as full-text searchable PDF using the results of OCR processing.
[0050] In this embodiment, all the processes are performed on the image forming apparatus 100, but the present invention is not limited to this. For example, in order to distribute the processing load, the scanned image data generated in step S201 may be transmitted to the server 191 via the LAN 190, and the server 191 may perform processes other than accepting operations from the user.
[0051] [Second embodiment] In this embodiment, a method is described for maintaining high accuracy in separating handwritten characters even for document images with a large inclination in the process of separating handwritten characters in the first embodiment (the process of step S202 in FIG. 2). For document images with a large inclination, a process that maintains accuracy in separating handwritten characters is executed, and for document images with a small inclination, a simple process that can provide sufficient accuracy is executed. FIG. 4 is a flowchart showing the steps of image processing in the second embodiment. Below, the second embodiment will be described, mainly focusing on the differences from the first embodiment.
[0052] In step S410, CPU 111 determines the range of inclination of the document image acquired in step S201. The range of inclination refers to the range of possible inclinations of the image that may be inclined. For example, a document image acquired by setting on a document table has a higher degree of freedom in how to place the document than a document image acquired by an ADF, so the range of inclination can be said to be larger. Even with an ADF, the document image may be inclined due to the way the document is set or the difference in speed between the left and right sides of the conveying motor. In particular, when an ADF that can handle documents of multiple sizes is used, the range of inclination is larger than when an ADF that can handle a specific document size is used. In this way, CPU 111 functions as an inclination angle range determination unit that determines the range of possible inclination angles in the image processing system.
[0053] For example, the ADF in the image forming apparatus 100 used in this embodiment automatically detects the document size, such as a small size (postcard, receipt, etc.). If the detected document size is smaller than the maximum document size that can be fed, the document is likely to be tilted due to a misalignment of the set position, etc. In this case, it is determined that the range of tilt is equal to or larger than the specified value (Yes). Also, if the document is read and acquired from the document platen, it is determined that the range of tilt is equal to or larger than the specified value (Yes). On the other hand, if the document size detected by the ADF is the maximum document size that can be fed, it is determined that the range of tilt is small and is smaller than the specified value (No). Thus, in step S410, if the possibility that the document image is tilted is high and the range of tilt is equal to or larger than the specified value (Yes), the process proceeds to step S401. On the other hand, if the range of tilt is smaller than the specified value (No), the process proceeds to step S402.
[0054] The handwritten character separation method used in this embodiment is a method in which a neural network (NN) is trained on handwritten character regions and other background regions in an image, and each pixel is judged to be handwritten or not. Below, we will explain the case depending on whether the range of the inclination angle of the input document image is equal to or larger than a specified range.
[0055] In step S401, since it is determined in step S410 that the range of the inclination angle of the input document image is equal to or larger than the prescribed range, CPU 111 performs handwritten character separation processing for documents with large inclination angles that can accommodate that range. The neural network used in the processing in step S401 is trained with a plurality of pattern images with different inclination angles of handwritten characters as image data of handwritten characters. FIG. 5 is a diagram showing an example of training data for handwritten character separation in the second embodiment. As shown in FIG. 5, by training handwritten characters at various angles, it is possible to maintain the extraction accuracy of handwritten characters even for input document images with large inclination angles.
[0056] In this embodiment, the inclination angles of handwritten characters are varied, but the present invention is not limited to this. For example, the range of inclination angles of the learning images of the neural network may be limited, and in this process, handwritten characters may be extracted while varying the inclination angles of the input document image, thereby covering all possible inclination angles.
[0057] In step S402, since it is determined in step S410 that the range of inclination of the input document image is below the prescribed range, the CPU 111 performs handwritten character separation processing for documents with small inclination that can accommodate that range. The neural network used in the processing of step S402 is trained with images of fewer rotation patterns than the neural network used in the processing of step S401 as image data for handwritten character separation. When the same accuracy is aimed for, it is possible to reduce the inference cost by using a neural network with a simple network structure with fewer training patterns. Therefore, when the inclination range is considered to be small, a simple network structure that is expected to provide sufficient accuracy is used. For the neural network used in this processing, only images with normal orientation, such as image 501 in FIG. 5, are used for training.
[0058] In this manner, in this embodiment, for document images with a large range of inclination, a neural network that has learned multiple rotated patterns of handwritten characters is used to perform handwritten character separation processing. This makes it possible to maintain high accuracy in separating handwritten characters. For document images with a small range of inclination, a neural network with fewer learning patterns is used to perform handwritten character separation processing. This makes it possible to achieve sufficient accuracy with simple processing.
[0059] [Other embodiments] Although the present invention has been described in detail above based on the preferred embodiments, the present invention is not limited to these specific embodiments, and various forms within the scope of the gist of the present invention are also included in the present invention. In addition, parts of the above-mentioned embodiments may be appropriately combined. In particular, in the above-mentioned embodiment, the CPU 111 of the image forming apparatus 100 is exemplified as the CPU that performs image processing, but the CPU 192 of the server 191 may also be used.
[0060] The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions. [Explanation of symbols]
[0061] 100...Image forming apparatus 110...Control device 120...Reader device 191…Server 198...Control device
Claims
1. In an image processing system for performing skew correction on a document image, a document image acquisition unit that acquires a document image including a mixture of handwritten characters and printed characters; a separation unit that uses a neural network that has learned features of handwritten characters from images including handwritten characters to determine pixels of handwritten characters included in the document image, and removes the determined pixels of handwritten characters from the document image to generate an image other than the handwritten characters; a tilt angle estimation unit that estimates a tilt angle by using the generated image other than the handwritten character; a tilt correction unit that corrects the tilt of the document image including both handwritten characters and printed characters based on the estimated tilt angle.
1. An image processing system comprising:
2. The separation unit applies a method corresponding to a range of inclination angles of the document image to remove handwritten characters from the document image, thereby generating an image other than the handwritten characters.
2. The image processing system according to claim 1.
3. The separation unit is if the range of the inclination angle of the document image is equal to or greater than a specified value, a first neural network is used to determine pixels of handwritten characters contained in the document image, and the determined pixels of handwritten characters are removed from the document image to generate an image other than the handwritten characters, if the range of the inclination angle of the document image is smaller than the specified value, a second neural network is used to determine pixels of handwritten characters contained in the document image, and the determined pixels of handwritten characters are removed from the document image to generate an image other than the handwritten characters; the inclination of the document image used for training the second neural network is smaller than the inclination of the document image used for training the first neural network; 2. The image processing system according to claim 1.
4. The separation unit is Using the neural network that has learned the characteristics of handwritten characters from images of handwritten characters with different inclination angles, pixels of the handwritten characters are judged, and the judged pixels of the handwritten characters are removed from the document image, thereby generating an image other than the handwritten characters.
2. The image processing system according to claim 1.
5. The separation unit is Using the neural network that has learned the characteristics of handwritten characters from a document image including a plurality of handwritten characters with different inclination angles, pixels of the handwritten characters are judged, and the judged pixels of the handwritten characters are removed from the document image, thereby generating an image other than the handwritten characters.
2. The image processing system according to claim 1.
6. 1. An image processing method for performing skew correction on a document image, comprising: a document image acquisition step of acquiring a document image including a mixture of handwritten characters and printed characters; a separation step of determining pixels of handwritten characters contained in the document image using a neural network that has been trained on features of handwritten characters using images including handwritten characters, and removing the determined pixels of handwritten characters from the document image to generate an image other than the handwritten characters; a tilt angle estimation step of estimating a tilt angle by using the generated image other than the handwritten character; and a skew correction step of correcting the skew of the document image including both handwritten characters and printed characters based on the estimated skew angle.
13. An image processing method comprising:
7. A program for causing a computer to function as the image processing device according to any one of claims 1 to 5.
Citation Information
Patent Citations
Discrimination method of document skew
JP1991268189A
Character recognizing device
JP1996044822A
Image processing system, image scanner, and image processing method
JP2020053931A
Method and system for document segmentation
US20030215136A1