A Handwritten Character Erasure Method Based on Deep Learning
Through the handwriting erasing method based on deep learning, the handwriting area in the document image is recognized and filled with the full convolutional neural network model, which solves the problems of low handwriting removal efficiency and high labor cost in the prior art, and achieves an efficient and automated handwriting erasing effect.
Patent Information
- Application Number
- CN202210401782.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-18
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-04-18
AI Technical Summary
The prior art is difficult to automatically and efficiently remove handwritten in document images, especially when handwritten overlaps with printed words, and the existing methods require significant labor costs.
Using a handwritten erasing method based on deep learning, we can identify and fill the handwritten area in the document image by making training samples, establishing a fully convolutional neural network model and training.
It realizes efficient and high-quality handwritten erasure of document images, convenient automation process, reduces labor costs, and has significant effect when handwritten and printed characters overlap.
Smart Images

Figure CN114708601B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for erasing handwritten characters, in particular to a method for erasing handwritten characters based on deep learning. Background Art
[0002] Taking a photo of a document and erasing the handwritten characters in the document image is a document restoration technology, which has wide applications in fields such as office work and study. For example: (1) protecting the private information written by hand in the document; (2) as a preprocessing to improve the accuracy of OCR technology for identifying printed characters; (3) restoring the original information of the document, such as collecting wrong questions of students for repeated practice, extracting the original form for re-filling information, removing notes or graffiti on the book image, etc. Comparing the images before and after erasing the handwritten characters can also be used for handwritten character extraction, handwritten information recognition, etc.
[0003] Implementing handwritten character erasure can be divided into two steps: First, distinguish handwritten characters and printed content pixel by pixel, and then fill the pixels in the handwritten character area to make it blend with the background. When distinguishing handwritten characters and printed content, their textures and grayscales are similar, and traditional methods for locating handwritten characters such as edge detection and color positioning fail; when filling the handwritten character area, the existing method of randomly sampling the background pixel value for filling cannot completely restore the printed characters in the case of overlap between handwritten characters and printed characters. Obviously, the effect of automatically erasing handwritten characters in document images using existing methods is not satisfactory, and if using image editing software to process handwritten characters pixel by pixel, the labor cost required is too high. Summary of the Invention
[0004] Object of the Invention: The technical problem to be solved by the present invention is to provide a method for erasing handwritten characters based on deep learning in view of the deficiencies of the prior art.
[0005] To solve the above technical problem, the present invention discloses a method for erasing handwritten characters based on deep learning, including the following steps:
[0006] Step 1, making training samples, including an original image containing handwritten characters and printed content, a mask image for classifying handwritten characters and printed content pixel by pixel, and a target image containing only printed content;
[0007] Step 2, establishing a deep learning model;
[0008] Step 3, after preprocessing the training samples, feeding them into the deep learning model for training. The training process includes: inputting the original image in the training samples, outputting a mask generation image and a target generation image; calculating the loss and optimizing the model parameters; repeating the training until the model converges to obtain the trained deep learning model;
[0009] Step 4, obtaining a document image from which handwritten characters need to be removed;
[0010] Step 5: Input the document image with handwritten characters to be removed into the trained deep learning model, and obtain the image after removing the handwritten characters, thus completing the erasure of handwritten characters based on deep learning.
[0011] The method for making training samples described in Step 1 of the present invention includes:
[0012] Step 1-1: Prepare a document containing handwritten characters and printed content, and use a photographing device or a scanning device to obtain a document image to get the original image.
[0013] Step 1-2: Use image editing software to remove and fill in the handwritten characters in the original image pixel by pixel to obtain the target image.
[0014] Step 1-3: Use an algorithm program to calculate the original image and the target image to obtain a mask image.
[0015] The method for obtaining a mask image by calculation described in Step 1-3 of the present invention includes:
[0016] Subtract the target image from the original image to obtain a difference image.
[0017] After taking the absolute value of the difference image, take the average of the values of the three channels to obtain a difference grayscale image.
[0018] Set the pixels in the difference grayscale image greater than the threshold to 0, indicating classification as handwritten characters, and set the other parts to 1, indicating printed content, to obtain a mask image, and the mask image is a binary mask image.
[0019] In Step 2 of the present invention, the deep learning model is a fully convolutional neural network, including: a mask generation module, a first-stage image generation module, and a second-stage image generation module.
[0020] Each module of the deep learning model described in Step 2 of the present invention adopts an encoder-decoder structure. Among them, the mask generation module shares parameters with the encoder of the first-stage image generation module.
[0021] In the fully convolutional neural network described in Step 2 of the present invention, skip connections are added and deployed between the encoder and decoder of each module, as well as between the decoder of the first-stage image generation module and the encoder of the second-stage image generation module.
[0022] The mask generation module adopts an attention mechanism to generate a spatial domain attention matrix with the mask feature map to guide the generation of the target image.
[0023] In the second-stage image generation module, a deformable convolution is used in the convolutional layer at the junction of the encoder and decoder to achieve adaptive feature sampling.
[0024] In Step 3 of the present invention, the preprocessing method includes:
[0025] Normalize the training samples to the same size, and perform data augmentation using random horizontal flipping and random angular rotation.
[0026] In step 3 of the present invention, the method for calculating the loss includes: using the SmoothL1 loss function to calculate the target map loss, using the Dice loss function to calculate the mask map loss, and adding the target map loss and the mask map loss to obtain the total loss loss. The calculation method includes:
[0027]
[0028]
[0029]
[0030] where img origin represents the original image, img mask represents the mask image, img target represents the target image, and the variable with a top line annotation represents the corresponding prediction result. represents the mask image generated by the mask generation module M. represents the target image output by the first-stage image generation module G1 and the second-stage image generation module G2;
[0031] Mask map loss The calculation method is:
[0032]
[0033] Target map loss The calculation method includes:
[0034]
[0035]
[0036] where Y, respectively represent two images with the same resolution, and each image has n pixel values. y i and respectively represent the i-th pixel values in Y and . The smoothl1 function is used to measure the distance between two values, and y and respectively represent the two values to be measured. β takes 0.5.
[0037] In step 4 of the present invention, the method for obtaining the document image includes:
[0038] Arbitrarily add handwritten handwriting on a paper document containing printed content, and use a photographing device or a scanning device to obtain an image of the document.
[0039] Step 5 of the present invention includes:
[0040] After normalizing the size of the document image, input it into the trained deep learning model. For the output image, use the bilinear interpolation method to adjust it to the size of the input document image, and finally complete the handwritten character erasure based on deep learning.
[0041] The present invention uses a fully convolutional neural network with an encoder-decoder structure as the basic module to realize the recognition and filling of the handwritten character area in the document image, and provides an efficient and high-quality method for erasing handwritten characters in document images. Compared with the existing manual erasure method of handwritten characters, the advantage of the present invention is that it can automatically erase handwritten characters by simply inputting the image into the trained model, with convenient operation and less labor cost; compared with the existing automatic erasure method of handwritten characters, the advantage of the present invention is that the quality of the generated image background details is good, and the recognition and erasure effects of handwritten characters are excellent.
[0042] Beneficial effects:
[0043] Through this method, the erasure of handwritten characters in the document image is realized. The convolutional neural network model used in this method introduces skip connections, combines the shallow features of the network with the deep semantic information, and enhances the image detail generation effect; adopts the deformable convolution method to enable the network to adaptively adjust the convolution sampling position and improve the erasure effect of handwritten characters of different shapes and sizes; through the attention mechanism, guides the network to strengthen the feature extraction of the handwritten character area, improves the ability to distinguish handwritten characters from printed content, and weakens the influence of complex backgrounds on the erasure effect. Description of the drawings
[0044] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0045] Figure 1 is the model schematic diagram of the present invention.
[0046] Figure 2 is the training flowchart of the present invention.
[0047] Figure 3 is the inference flowchart.
[0048] Figure 4 is the schematic diagram of the document image containing handwritten characters and printed content.
[0049] Figure 5 is the schematic diagram of the document image after erasing the handwritten characters using the method of the present invention. Specific embodiments
[0050] The present invention proposes a handwritten character erasure method based on deep learning, which includes the following steps:
[0051] Step 1, making training samples, including an original image containing handwritten characters and printed content, a mask image with pixel-by-pixel classification of handwritten characters and printed content, and a target image containing only printed content;
[0052] Step 1-1, preparing a document containing handwritten characters and printed content, and using a photographing device or a scanning device to obtain a document image to get the original image;
[0053] Step 1-2, using image editing software to remove and fill in handwritten characters in the original image pixel by pixel to obtain the target image;
[0054] Step 1-3, using an algorithm program to calculate the original image and the target image to obtain a mask image. The method includes:
[0055] Taking the difference between the original image and the target image to obtain a difference image;
[0056] After taking the absolute value of the difference image, taking the average of the values of the three channels to obtain a difference grayscale image;
[0057] Setting the pixels in the difference grayscale image greater than the threshold to 0, indicating classification as handwritten characters, and setting other parts to 1, indicating printed content, to obtain a mask image, and the mask image is a binary mask image.
[0058] Step 2, establishing a deep learning model;
[0059] The deep learning model is a fully convolutional neural network, including: a mask generation module, a first-stage image generation module, and a second-stage image generation module;
[0060] Each module of the deep learning model adopts an encoder-decoder structure. Among them, the mask generation module shares parameters with the encoder of the first-stage image generation module;
[0061] In the fully convolutional neural network, skip connections are added and deployed between the encoder and decoder of each module and between the decoder of the first-stage image generation module and the encoder of the second-stage image generation module;
[0062] The mask generation module adopts an attention mechanism to generate a spatial domain attention matrix with the mask feature map to guide the generation of the target image;
[0063] In the second-stage image generation module, a deformable convolution is used in the convolutional layer at the junction of the encoder and decoder to achieve adaptive feature sampling.
[0064] Step 3: After preprocessing the training samples, feed them into the deep learning model for training. The training process includes: inputting the original images in the training samples, and outputting the mask generation images and the target generation images; calculating the loss and optimizing the model parameters; repeating the training until the model converges to obtain the trained deep learning model;
[0065] The preprocessing method includes:
[0066] Normalize the training samples to the same size, and perform data augmentation using random horizontal flipping and random angular rotation.
[0067] The method for calculating the loss includes: using the SmoothL1 loss function (Reference: Ren S, He K, Girshick R, et al. Faster r-cnn: Towards real-time object detection with region proposal networks[J]. Advances in neural information processing systems, 2015, 28.) to calculate the target image loss, using the Dice loss function (Reference: Milletari F, Navab N, Ahmadi SA. V-net: Fully convolutional neural networks for volumetric medical image segmentation[C] / / 2016 fourth international conference on 3D vision(3DV). IEEE, 2016:565-571.) to calculate the mask image loss, and adding the target image loss and the mask image loss to obtain the total loss loss. The calculation method includes:
[0068]
[0069]
[0070]
[0071] where img origin represents the original image, img mask represents the mask image, img target represents the target image. Variables with a top line annotation represent the corresponding prediction results. represents the mask image generated by the mask generation module M, represents the target image output by the first-stage image generation module G1 and the second-stage image generation module G2;
[0072] Mask graph loss The calculation method is as follows:
[0073]
[0074] Target graph loss The calculation method includes:
[0075]
[0076]
[0077] Among them, Y, respectively represent two graphs with the same resolution, and each graph has n pixel values. y i and respectively represent the i-th pixel value in Y and The smoothl1 function is used to measure the distance between two values. y and respectively represent the two values to be measured, and β takes 0.5.
[0078] Step 4, obtain the document image from which the handwritten characters need to be removed. The method includes:
[0079] Randomly add handwritten character strokes on the paper document containing printed content, and use a photographing device or a scanning device to obtain an image of the document.
[0080] Step 5, input the document image from which the handwritten characters need to be removed into the trained deep learning model to obtain the image after removing the handwritten characters, and complete the erasure of handwritten characters based on deep learning. After normalizing the size of the document image, input it into the trained deep learning model, and use the bilinear interpolation method for the output image (Reference: Digital Image Processing: 3rd Edition, written by (US) Gonzalez (Gonzalez, R.C.), (US) Woods (Woods, R.E.), translated by Ruan Qiuqi, etc., Beijing: Publishing House of Electronics Industry, June 2011, page 37) to adjust it to the size of the input document image, and finally complete the erasure of handwritten characters based on deep learning.
[0081] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0082] Embodiment
[0083] A method for erasing handwritten characters based on deep learning, as Figure 1 shown, this example is a fully convolutional neural network model, which consists of three modules: a mask generation module, a first-stage image generation module, and a second-stage image generation module. Among them, each module adopts an encoder-decoder structure, and the encoder of the mask generation module shares parameters with the encoder of the first-stage image generation module.
[0084] This embodiment provides a training method for a handwritten character erasure model based on deep learning, and the training process is as follows Figure 2 shown and the process is as follows:
[0085] 1. Making training samples
[0086] Prepare a paper handwritten character document with printed content. After appropriately adding handwritten characters, use a scanner to obtain the document image as the original image; use Photoshop software to perform pixel-by-pixel recognition and filling of the handwritten character area to obtain the target image; use an algorithm to calculate the mask image. Each group of samples consists of an original image, a target image, and a mask image. To ensure the training effect, a certain amount of samples need to be made. For example, 1100 groups of samples are made, with 1000 groups as the training set and 100 groups as the validation set.
[0087] The algorithm for calculating the mask image is as follows: Take the original image and the corresponding target image, subtract them and take the absolute value, then take the average value of the three channels to obtain the difference grayscale image; Set the pixels in the difference grayscale image greater than 25 to 0 and those less than or equal to 25 to 1, that is, the mask image is obtained.
[0088] 2. Establishing a deep learning model
[0089] Construct Figure 1 the fully convolutional neural network model shown, which can be built using existing deep learning frameworks, such as PyTorch.
[0090] 3. Sample preprocessing and model training
[0091] Before each group of samples is input into the model, preprocessing is required. Sample preprocessing includes training sample size normalization and data augmentation. For example, normalize to 512×512 and use random horizontal flipping and random angle rotation for data augmentation. When performing sample preprocessing, it is necessary to ensure that the preprocessing operations performed on each image within the same group of samples are the same, such as the rotation angles of the images need to be consistent.
[0092] Training the model: Input the preprocessed samples into the model, calculate the generator loss based on the output prediction image, and use the Adam optimizer to optimize the model parameters. The loss during model training consists of two parts: the mask image loss and the target image loss. Among them, the mask image loss is calculated using the Dice function, and the target image loss is calculated using the SmoothL1 function, with β taken as 0.5 in the SmoothL1 function. The specific formulas are as follows
[0093] Total loss:
[0094]
[0095]
[0096]
[0097] Mask graph loss:
[0098]
[0099] Target graph loss:
[0100]
[0101]
[0102] where img origin represents the original graph, img mask represents the mask graph, img target represents the target graph. Variables with a top line annotation represent the corresponding prediction results. For example, represents the mask graph generated by the mask generation module M, represents the target graph output by the first-stage image generation module G1 and the second-stage image generation module G2. Y, respectively represent two graphs of the same resolution. Each graph has n pixel values. y i and respectively represent the i-th pixel value in Y and . The smoothl1 function is used to measure the distance between two values. y and respectively represent the two values to be measured, and β takes 0.5.
[0103] 4. Model performance evaluation
[0104] Calculate the score to measure the performance of the model in erasing handwritten characters. The higher the score value, the better the model performance. The score is specifically composed of PSNR (Peak Signal-to-Noise Ratio, which is often used as a measurement method for signal reconstruction quality in fields such as image compression in image processing. References: It is introduced and the calculation method is given on page 31 of "Document and Image compression". The information of this book is as follows: Document and Image compression, Barni, Mauro, 2018, CRC press) and MSSSIM (Multi-Scale Structural Similarity Index. References: Multiscale structural similarity for image quality assessment, Wang, Zhou and Simoncelli, Eero P and Bovik, Alan C, The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003, 2, 1398—1402, 2003, Ieee, https: / / ieeexplore.ieee.org / document / 1292216):
[0105]
[0106] Among them, img target represents the real target image, and represents the target image output by the module;
[0107] The PSNR calculation formula is as follows:
[0108]
[0109] The MSSSIM calculation formula is as follows:
[0110]
[0111]
[0112]
[0113] Among them, X and Y respectively represent two images with the same resolution. L(X, Y) calculates the luminance similarity between the two images, C(X, Y) calculates the contrast similarity, S(X, Y) calculates the structural similarity, and α M , β j and γ j are constants used to adjust the weights of the luminance similarity, contrast similarity, and structural similarity. C1, C2, and C3 are constants used to stabilize the calculation and prevent the denominator from being too small. μ X and μ Y respectively represent the means of X and Y, σ X and σ Y respectively represent the standard deviations of X and Y, and σ XY represents the covariance of X and Y. M is a constant representing the number of levels for calculating the structural similarity. L M (X, Y) means reducing the widths and heights of X and Y by a factor of 2 M-1 and calculating the luminance similarity between the two reduced images. C j (X, Y) and S j (X, Y) respectively mean reducing the widths and heights of X and Y by a factor of 2 j-1 and respectively calculating the contrast similarity and structural similarity between the two reduced images.
[0114] This embodiment provides a method for inferring a handwritten character erasure model based on deep learning. The inference process is as Figure 3 shown, and the process is as follows:
[0115] 1. Obtain the document image of the handwritten characters to be erased
[0116] Prepare a paper handwritten document with printed content. After appropriately adding handwritten characters, use a scanner to obtain the document image, as Figure 4 shown.
[0117] 2. Use the trained convolutional neural network to erase the handwritten characters
[0118] Load the trained parameters into the convolutional neural network, input the document image normalized to a size of 512×512, and obtain the model output image.
[0119] 3. Post-process the document image
[0120] Interpolate and restore the model output image to the original image size, which is the document image after erasing the handwritten characters, as Figure 5 shown.
[0121] The present invention provides an idea and method for a handwritten character erasure method based on deep learning. There are many methods and ways to specifically implement this technical solution. The above description is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using existing technologies.
Claims
1. A handwritten character erasing method based on deep learning, characterized in that, It includes the following steps: Step 1: Prepare training samples, including the original image containing handwritten characters and printed content, the mask image with pixel-by-pixel classification of handwritten characters and printed content, and the target image containing only printed content; Step 2: Establish a deep learning model; Step 3: After preprocessing the training samples, feed them into the deep learning model for training. The training process includes: inputting the original image in the training samples, outputting the mask generation image and the target generation image; calculating the loss and optimizing the model parameters; repeating the training until the model converges to obtain the trained deep learning model; Step 4: Obtain the document image from which handwritten characters need to be removed; Step 5: Input the document image from which handwritten characters need to be removed into the trained deep learning model to obtain the image after removing handwritten characters, and complete the erasure of handwritten characters based on deep learning; Among them, in Step 3, the method for calculating the loss includes: using the SmoothL1 loss function to calculate the target image loss, using the Dice loss function to calculate the mask image loss, and adding the target image loss and the mask image loss to obtain the total loss loss. The calculation method includes: Among them, img origin represents the original image, img mask represents the mask image, img target represents the target image. Variables with top-line annotations represent the corresponding prediction results. represents the mask image generated by the mask generation module M. represents the target image output by the first-stage image generation module G1 and the second-stage image generation module G2; Mask graph loss The calculation method is as follows: Target graph loss The calculation method includes: Among them, Y, respectively represent two images with the same resolution, and each image has n pixel values. y i and respectively represent the i-th pixel values in Y and , the smoothl1 function is used to measure the distance between two values, y and respectively represent the two values to be measured, and β takes 0.
5.
2. The method for erasing handwritten characters based on deep learning according to claim 1, wherein The method for preparing training samples in Step 1 includes: Step 1-1: Prepare a document containing handwritten characters and printed content, and use a photographing device or a scanning device to obtain the document image to get the original image; Step 1-2: Use image editing software to remove and fill in the handwritten characters in the original image pixel by pixel to get the target image; Step 1-3: Use an algorithm program to calculate the original image and the target image to get the mask image.
3. The method for erasing handwritten characters based on deep learning according to claim 2, characterized in that The method for obtaining the mask image in Step 1-3 includes: Taking the difference between the original image and the target image to get the difference image; After taking the absolute value of the difference image, averaging the values of the three channels to get the difference grayscale image; Setting the pixels in the difference grayscale image greater than the threshold to 0, indicating classification as handwritten characters, and setting other parts to 1, indicating printed content, to obtain the mask image, and the mask image is a binary mask image.
4. The method for erasing handwritten characters based on deep learning according to claim 3, characterized in that, In Step 2, the deep learning model is a fully convolutional neural network, including: a mask generation module, a first-stage image generation module, and a second-stage image generation module.
5. The method for erasing handwritten characters based on deep learning according to claim 4, wherein, Each module of the deep learning model in Step 2 adopts an encoder-decoder structure. Among them, the mask generation module shares parameters with the encoder of the first-stage image generation module.
6. A handwritten character erasing method based on deep learning according to claim 5, characterized in that In the fully convolutional neural network in Step 2, skip connections are added and deployed between the encoder and decoder of each module and between the decoder of the first-stage image generation module and the encoder of the second-stage image generation module; The mask generation module adopts an attention mechanism to generate a spatial domain attention matrix with the mask feature map to guide the generation of the target image; In the second-stage image generation module, deformable convolution is used in the convolutional layer at the junction of the encoder and decoder to achieve adaptive feature sampling.
7. A handwritten character erasing method based on deep learning according to claim 6, characterized in that In Step 3, the method for preprocessing the training samples includes: Normalizing the training samples to the same size and performing data augmentation using random horizontal flipping and random angular rotation.
8. A handwritten character erasing method based on deep learning according to claim 7, characterized in that In Step 4, the method for obtaining the document image includes: Arbitrarily add handwritten character strokes on a paper document containing printed content, and use a photographing device or a scanning device to obtain the image of this document.
9. A method for erasing handwritten characters based on deep learning according to claim 8, characterized in that, Step 5 includes: After normalizing the document image size, input it into the trained deep learning model. For the output image, use the bilinear interpolation method to adjust it to the size of the input document image, and finally complete the erasure of handwritten characters based on deep learning.
Citation Information
Patent Citations
Handwritten font removal method based on artificial intelligence
CN113989816A