An Automatic Image Rectification Method for Archival Scans Based on Deep Learning
Through the two-stage deviation correction method of deep learning technology, the problem of insufficient adaptability of full-angle deviation correction and multiple types of scanned images in the existing technology is solved, and the automatic image correction effect with high accuracy and robustness is achieved.
Patent Information
- Application Number
- CN202111508488.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-10
AI Technical Summary
The existing deviation correction algorithms are not accurate and/or not satisfactory and are not robust when correcting deviations at full angle (0 to 360 degrees) and adapting to various types of scanned pictures.
Deep learning technology is used to automatically correct the image through two-stage deviation correction method. The first stage is to perform preliminary correction in four directions through the improved VGG16 network, and the second stage is to achieve fine deviation correction through deep learning combined with pixel projection correction.
It improves the accuracy and robustness of scanned pictures, and can effectively deal with various types of scanned pictures, including text types, text types and table types, etc., to achieve full-angle correction.
Smart Images

Figure CN114358137B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision and pattern recognition, image processing, and image automatic rectification systems, and particularly relates to a method for automatically rectifying scanned archive images based on deep learning. Background Art
[0002] During the normal operation of banks, a large number of paper archives are generated. Paper archives are not convenient for storage and retrieval, so it is necessary to digitize the paper archives. Scanning paper materials is a common method for digitizing paper archives. However, the electronic version pictures of the archives obtained after scanning are inevitably tilted, which affects the beauty and readability of the pictures and even affects subsequent operations such as OCR recognition. Currently, existing rectification software can usually only perform effective rectification within a small range (within 90 degrees). When the angle of the scanned picture is greater than 90 degrees (the skew angle range of real-world scanned pictures is the full angle, i.e., 0 to 360 degrees), the effect is not ideal and it cannot be used practically.
[0003] Currently, there are many studies on algorithms for picture rectification, but most of them can only rectify specific pictures within 0 to 90 degrees and are only implemented using traditional image processing methods. However, algorithms that can rectify pictures within 0 to 360 degrees and can better adapt to text-type pictures, pictures with embedded graphics and text, and pictures with embedded table types are still relatively few or not well studied. The reason may be that the large rectification range and diverse picture types (especially scanned pictures in real-world scenarios) make the design of related algorithms challenging.
[0004] Xiang Jing et al. (Xiang Jing, Zhang Rufeng, Chu Junying, et al. Oblique correction of text images combining horizontal projection and affine transformation [J]. Paper-making Equipment & Materials, 2020, v.49; No.185(02):257-257.) studied the rectification method of text images. The main idea is the pixel projection method. Specifically, first, the obtained scanned text image is grayscale, binarized using the maximum inter-class variance method, and morphological filtering is used to remove noise. Then, according to the characteristics of the text image, the image is rotated within the range of -90 to 90 degrees using affine transformation. The horizontal projection histogram is calculated every 1 degree of rotation, and the number of zero values, that is, the number of values less than a certain threshold in the projection histogram, is calculated. Finally, the angle at which the number of zero values is the largest is the rotation angle required to achieve text image correction. This method uses the maximum inter-class method in the image segmentation stage. When the image does not conform to the bimodal distribution, the segmentation effect is poor. This method also has an obvious defect. When the skew angle of the image is large, for some scanned documents with vertical arrangements and scanned documents with breakthrough pictures embedded, the accuracy is poor based solely on the pixel projection method (without using deep learning technology).
[0005] Ming Delie et al. (Ming Delie, Liu Jian, Hu Jiazhong, et al. A Fast Detection and Correction Method for Small-Angle Tilted Images [J]. Journal of Huazhong University of Science and Technology (Natural Science Edition), 2000, 28(5): 66-68.) Based on the characteristics of images in the character recognition system, a fast algorithm for calculating and correcting the tilt angle of text images with small-angle tilt (within ±5 degrees) is proposed, effectively overcoming the influence of geometric distortion on the character recognition system. By adopting a linear correction method of moving the whole block segment by segment, the execution speed of the correction algorithm is doubled. Although the algorithm is fast, it can only correct the tilt within (±5 degrees), and the applicable range is small.
[0006] Wang Hengyou et al. (Wang Hengyou, Yu Zhan, Zhang Changlun, et al. Batch Rectification of Scanned Document Images Based on Low-Rank Matrix Decomposition. November 22, 2017 [J]. Computer Engineering and Applications, 2017.) Traditional algorithms based on the texture structure of pictures themselves, such as Hough transform and Radon transform, are not only vulnerable to the special structure or noise of the document itself, but also the average time-consuming for rectifying a single image is relatively long. A batch rectification method for scanned document images based on the theory of low-rank matrix decomposition is proposed. This method constructs a larger matrix from a batch of images and appropriately rotates each column through iteration to achieve the purpose of the matrix having a lower rank, thereby realizing the appropriate estimation and rectification of the deflection angle of each image. This method runs relatively fast. However, for scanned images with embedded pictures and text, the rectification effect is not good.
[0007] In summary, although certain achievements have been made in the automatic image rectification method, due to the diversity of scanned image types (such as scanned images of pure text type, scanned images with embedded pictures, scanned images with embedded tables, scanned images with embedded seals, and the background color of scanned images may not necessarily be white, etc.), current algorithms are difficult to be well applicable to various image types. Current methods are basically based on traditional methods for image rectification, and the rectification effect is not ideal. According to our research, there is currently no scanned image rectification method that can utilize the high-precision feature extraction technology of deep learning and organically combine traditional methods. To meet the requirements of practical applications, it is urgent to make further improvements in the accuracy and real-time performance of image rectification. The present invention creatively proposes to use deep learning object recognition technology to detect the main components of the image, thereby avoiding the influence of black and white or white edges around the scanned image on subsequent image rectification; further, it creatively proposes a two-stage rectification method. In the first stage, by improving the current VGG16 network, the image is preliminarily rectified in 4 directions through deep learning image classification technology (the first stage). In the second stage, fine rectification is achieved by combining deep learning image classification in 91 directions and improved projection rectification (the second stage). In addition, in the image preprocessing stage, by using deep learning image classification technology and determining the key parameters of Gamma correction by intelligently predicting the image brightness, the traditional Gamma correction method for image brightness correction is improved, thereby further ensuring the accuracy of subsequent rectification recognition. And through deep learning object detection technology, the main components of the scanned image are identified, the image area to be processed for rectification is reduced, the extraction of candidate areas is completed, and the accuracy of rectification relative to the entire original image is improved. Summary of the Invention
[0008] The purpose of the embodiment of the present invention is to provide an automatic rectification method for archive scanned images based on deep learning, aiming to solve the problems that the existing rectification algorithms do not meet the standard and / or are not satisfactory in accuracy and / or are not robust enough for full-angle (0 to 360 degrees) rectification and adapting to various types of scanned images.
[0009] It is characterized in that, aiming at the imaging characteristics of archive scanned images, it is proposed to adaptively adjust the brightness of the image by using the deep learning brightness classification result, improve the brightness of the scanned image, and detect the main components of the entire scanned image by using deep learning object detection technology to avoid the influence of black or white edges around the image on the subsequent image rectification angle estimation. Design a two-stage rectification strategy (large-direction rectification and small-direction rectification) to complete the rectification of the scanned image, specifically including:
[0010] Step 1, adjust the image brightness by using the deep learning brightness classification result to improve the brightness of the scanned image;
[0011] Step 2: Mark the main components of the scanned document. After obtaining a model through deep learning training, detect the main components of the scanned document.
[0012] Step 3: The deep learning model performs rectification in 4 directions in the first stage (coarse rectification).
[0013] Step 4: Complete the rectification in the second stage through deep learning combined with pixel projection rectification (fine rectification).
[0014] Furthermore, for the automatic rectification method of archival scanned document images based on deep learning according to the present invention, the deep learning brightness classification in Step 1 means collecting 8000 archival scanned document images in real scenes, and then dividing the pictures into 5 brightness categories manually in 5 brightness levels. Use the self - deepened VGG16 deep network for training to obtain a picture brightness classification network model (the traditional VGG16 contains 16 hidden layers (13 convolutional layers and 3 fully - connected layers), and the self - deepened VGG network in this patent has 15 convolutional layers and 5 fully - connected layers. Experiments show that the detection accuracy of the classification network after such deepening has increased by 5%). Then, use this model to perform deep learning model inference on the original input picture to obtain the deep learning brightness classification result of the test picture, denoted as Level (abbreviation: L), and the range of Level is 1 - 5.
[0015] Furthermore, for the automatic rectification method of archival scanned document images based on deep learning according to the present invention, the adjustment of the picture brightness according to the deep learning brightness classification result in Step 1 means using the classification label Level value obtained from the self - deepened VGG deep learning classification network to adjust the brightness of the original scanned document picture. The adjustment rule is to set the exponential value of Gamma correction as an expression based on the result L obtained from deep learning: 1 / (5 - L + 1).
[0016] Furthermore, for the automatic rectification method of archival scanned document images based on deep learning according to the present invention, the marking of the main components of the scanned document in Step 2 and the detection of the main components of the scanned document after obtaining a model through deep learning mean manually marking the image areas (main components) that only need to be retained in 8000 original scanned documents for the subsequent main component detection training data of deep learning. The detection of the main components of the scanned document means using the YOLOv2 deep learning object detection network to train the marked main components to obtain a main component detection model, and performing forward inference of the deep network to obtain the main component image area of the original scanned image.
[0017] Further, for the automatic image skew correction method of archival scans based on deep learning according to the present invention, it is characterized in that the four-direction skew correction (large-direction skew correction) in the first stage by the deep learning model in step three means collecting 8,000 original pictures with a 0-degree skew (if the original pictures are not at 0 degrees, manually correct them to 0 degrees), and then by rotating the pictures clockwise, obtaining picture sets with skews of 90 degrees, 180 degrees, and 270 degrees respectively. Combining with the original picture set, a picture training set for four-direction large-direction skew correction is obtained. Thus, using the self-deepened VGG, a deep learning network model for four-direction classification of the main component images is trained. Furthermore, for each main component image, using this four-direction classification deep learning network to classify it in four directions, obtaining a class label C (the value range is 0, 1, 2, 3), and then by rotating the picture counterclockwise, that is, counterclockwise rotating by C×90 degrees, thus completing the four-direction skew correction (large-direction skew correction) in the first stage.
[0018] Further, for the automatic image skew correction method of archival scans based on deep learning according to the present invention, it is characterized in that the skew correction (small-direction skew correction) by deep learning combined with pixel projection in step four to complete the second-stage skew correction means that the second-stage skew correction is carried out in two branches on the basis of the image obtained by the first-stage skew correction. The first branch is to use the self-deepened VGG deep learning network to implement 91-category small-direction estimation, and the second branch is small-direction estimation based on the Sobel edge projection method. When the result of the first branch differs from the result of the second branch by less than 5 degrees, the result of the first branch is used as the final small skew correction angle estimation, otherwise the result of the second branch is used as the final small skew correction angle estimation, finally completing the selection of the small skew correction result. The implementation of 91-category small-direction estimation by the self-deepened VGG deep learning network means collecting 16,000 original pictures with a 0-degree skew (if the original pictures are not at 0 degrees, manually correct them to 0 degrees), and then by rotating the pictures, obtaining a picture set with skews from 1 to 90 degrees. Combining with the original picture set, a picture training set for 91-category large-direction skew correction is obtained. Thus, using the self-deepened VGG, a deep learning network model for 91-direction classification of the main component images is trained. Furthermore, for each main component image, using this 91-direction classification deep learning network to classify it in 91 directions, obtaining a class label C1 (the range is 0 to 90 of integers), that is, correspondingly obtaining the result of implementing 91-category small-direction estimation by the self-deepened VGG deep learning network. The small-direction estimation based on the Sobel edge projection method means using the self-designed Sobel weighted segmentation algorithm to segment the original image, and then performing angular rotation traversal on the segmented image. Further, performing pixel value projection in the horizontal direction on the pixels of the segmented image, and finding the skew correction direction with the most 0 values in the projection value, which is the skew correction angle obtained by the small-direction estimation based on the Sobel edge projection method.
[0019] An automatic image rectification method for archival scanned images based on deep learning provided by the present invention focuses on proposing a method that can simultaneously correct various types of scanned images. Compared with existing image stain detection and removal technologies, the present invention has the following advantages and effects: In order to correct the brightness of the scanned image and improve the accuracy of subsequent correction, it is proposed to perform deep learning image classification on the scanned image to obtain the key parameters for subsequent Gamma brightness correction, and it is proposed to improve the traditional Gamma brightness correction with the deep learning brightness classification results to improve the brightness correction effect of the scanned image; before rectifying the image, it is proposed to label the main components of the scanned image, and after obtaining the model through deep learning training, detect the main components of the scanned image, overcoming the influence of the surrounding contours, shadows, etc. of the image on subsequent image rectification, so as to complete the selection of the rectification candidate area; when rectifying the image, a two-stage progressive rectification method is designed. In the first stage, 4-classification is implemented based on the improved VGG16 network for preliminary rectification, and in the second stage, 91-classification is implemented based on the improved VGG16 network and comprehensively judged with the traditional projection rectification to complete fine rectification. The method of the present invention can effectively cope with the influence of the brightness difference of scanned images, the diversity of bank scanned image types (text types, graphic and text types, table types, etc.) on scanned image stains, avoid the influence of surrounding contours, shadows, etc. on subsequent image rectification by designing principal component recognition, effectively achieve full-angle rectification of scanned images through a two-stage method from rough to fine, and improve the traditional Sobel segmentation algorithm to make the pixel projection segmentation accuracy higher. Description of the Drawings
[0020] Figure 1 It is a schematic structural diagram of an automatic image rectification system for archival scanned images based on deep learning provided by an embodiment of the present invention;
[0021] Figure 1 In the figure: A is the candidate area selection module, B is the brightness classification offline training module, C is the principal component detection offline training module, D is the deep learning rectification offline training module, and E is the two-stage image rectification module
[0022] Figure 2 It is the main steps included in an automatic image rectification method for archival scanned images based on deep learning provided by an embodiment of the present invention Detailed Embodiment
[0023] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0024] The application principle of the present invention will be further described below in conjunction with the drawings and specific embodiments.
[0025] As Figure 2 shown, an automatic image rectification method for archive scanned images based on deep learning according to an embodiment of the present invention includes the following steps:
[0026] S101, adjusting the brightness of the picture according to the deep learning brightness classification result to improve the brightness of the scanned image;
[0027] S102, marking the main components of the scanned image, after performing deep learning training to obtain a model, detecting the main components of the scanned image;
[0028] S103, performing rectification in 4 directions in the first stage by the deep learning model (large - direction rectification);
[0029] S104, completing the rectification in the second stage by deep learning combined with pixel projection rectification (small - direction rectification).
[0030] The deep learning brightness classification described in step S101 means collecting 8000 archive scanned images in real scenes, and then dividing the pictures into 5 brightness categories manually according to 5 brightness levels. Using the self - deepened VGG16 deep network for training to obtain a picture brightness classification network model (the traditional VGG16 contains 16 hidden layers (13 convolutional layers and 3 fully - connected layers), and the self - deepened VGG network in this patent has 15 convolutional layers and 5 fully - connected layers. Experiments show that the detection accuracy of the classification network is improved by 5% after such deepening). Then, using this model, performing deep learning model inference on the original input picture to obtain the deep learning brightness classification result of the test picture, denoted as Level (abbreviation: L), and the range of Level is 1 - 5.
[0031] The adjusting the brightness of the picture according to the deep learning brightness classification result described in step S101 means using the classification label Level value obtained by the self - deepened VGG deep learning classification network to adjust the brightness of the original scanned image. The adjustment rule is to set the exponential value of Gamma correction as an expression based on the result L obtained by deep learning: 1 / (5 - L + 1).
[0032] The marking the main components of the scanned image, after performing deep learning training to obtain a model, detecting the main components of the scanned image described in step S102 means manually marking the image areas (main components) that only need to be retained in 8000 original scanned images for use as the main - component detection training data for subsequent deep learning. The detecting the main components of the scanned image means using the YOLOv2 deep learning object detection network to train the marked main components to obtain a main - component detection model, and performing forward inference of the deep network to obtain the main - component image area of the original scanned image.
[0033] The first-stage 4-direction correction (large-direction correction) by the deep learning model described in step S103 means collecting 8,000 original images with a 0-degree skew (if the original images are not at 0 degrees, manually correct them to 0 degrees), and then obtaining image sets with skews of 90 degrees, 180 degrees, and 270 degrees by rotating the images clockwise. Combining with the original image set, a 4-category image training set for large-direction correction is obtained. Thus, a self-deepened VGG is used to train a deep learning network model for 4-direction classification of the main component images. Furthermore, for each main component image, it is classified in 4 directions using this 4-direction classification deep learning network to obtain a class label C (the value range is 0, 1, 2, 3). Then, by rotating the image counterclockwise, that is, rotating counterclockwise by C×90 degrees, the first-stage 4-direction correction (large-direction correction) is completed.
[0034] The deep learning combined pixel projection correction (small-direction correction) described in step S104 to complete the second-stage correction means that the second-stage correction is carried out in two branches based on the images corrected in the first stage. The first branch is to use a self-deepened VGG deep learning network to implement 91-category small-direction estimation, and the second branch is small-direction estimation based on the Sobel edge projection method. When the result of the first branch differs from the result of the second branch by less than 5 degrees, the result of the first branch is used as the final small correction angle estimation; otherwise, the result of the second branch is used as the final small correction angle estimation, and finally the selection of the small correction result is completed. The self-deepened VGG deep learning network implementing 91-category small-direction estimation means collecting 16,000 original images with a 0-degree skew (if the original images are not at 0 degrees, manually correct them to 0 degrees), and then obtaining an image set with skews from 1 to 90 degrees by rotating the images. Combining with the original image set, a 91-category image training set for large-direction correction is obtained. Thus, a self-deepened VGG is used to train a deep learning network model for 91-direction classification of the main component images. Furthermore, for each main component image, it is classified in 91 directions using this 91-direction classification deep learning network to obtain a class label C1 (the range is 0 to 90 as an integer), that is, the result of implementing 91-category small-direction estimation using the self-deepened VGG deep learning network is correspondingly obtained. The small-direction estimation based on the Sobel edge projection method means using a self-designed Sobel weighted segmentation algorithm to segment the original image, then performing angular rotation traversal on the segmented image. Further, the pixel values of the segmented image are projected in the horizontal direction, and the correction direction with the most 0 values in the projection is found, which is the correction angle obtained by the small-direction estimation based on the Sobel edge projection method.
[0035] Such as Figure 1As shown in the figure, an automatic image rectification method for archival scanned images based on deep learning in an embodiment of the present invention mainly consists of a candidate region selection module A, a brightness classification offline training module B, a principal component detection offline training module C, a deep learning rectification deviation offline training module D, and a two-stage image rectification module E.
[0036] The candidate region selection module A is used to improve the traditional Gamma brightness correction through deep learning brightness classification and identify candidate regions using a principal component detection model.
[0037] The brightness classifier offline training module B classifies the brightness of scanned images using an improved VGG16 deep learning image classification network.
[0038] The principal component detection offline training module C detects the principal components of scanned images using a YOLOv2 deep learning object detection network.
[0039] The deep learning rectification deviation offline training module D trains two image direction prediction networks using an improved VGG16 deep learning image classification network, one for 4-angle (class) prediction and one for 91-angle (class) prediction.
[0040] The two-stage image rectification module E is connected to modules A and D to achieve two-stage rectification. In the first stage, a deep learning model is used for 4-direction rectification (large-direction rectification). In the second stage, a deep learning model is used for 91-direction prediction and the result of pixel projection is considered simultaneously to complete small-direction rectification.
[0041] Specific embodiments of the present invention:
[0042] The overall process of the method of the present invention is as Figure 1 As shown in the figure, the main body of the method of the present invention includes four parts: 1) Adjust the image brightness using the deep learning brightness classification result to improve the brightness of scanned images; 2) Mark the main components of scanned images, and after deep learning training to obtain a model, detect the main components of scanned images; 3) Perform 4-direction rectification (large-direction rectification) in the first stage using a deep learning model; 4) Complete the rectification in the second stage through deep learning combined with pixel projection rectification (small-direction rectification).
[0043] 1. Adjust the image brightness using the deep learning brightness classification result to improve the brightness of scanned images
[0044] There are various types of bank scan documents (text type, text and image type, table type, etc.), which makes it difficult to develop a robust scan document rectification method. With the rapid development of computer vision and deep learning technologies, it has become possible to apply deep learning technologies to scan document rectification. For scan document images in natural scenes, the brightness problem is particularly serious. To improve the accuracy and robustness of scan document image rectification, it is proposed to first use a self-deepened VGG16 deep learning network to classify the brightness, obtain the brightness level, and use it as the key parameter for subsequent traditional Gamma image brightness correction to improve the current Gamma image brightness correction method. Specifically, use the classification label Level value obtained by the self-deepened VGG deep learning classification network to adjust the brightness of the original scan document image. The specific method is to design an improved traditional Gamma image brightness correction. For the specific implementation method of the improved traditional Gamma image brightness correction, please refer to Section 1.2.
[0045] 1.1 Brightness Classification Offline Training Module and Brightness Online Classification Module
[0046] The offline training is implemented using the self-deepened VGG16 network. Compared with AlexNet, VGG16 uses continuous 3x3 convolutional kernels instead of the larger convolutional kernels (11x11, 7x7, 5x5) in AlexNet. For a given receptive field, using stacked small convolutional kernels achieves better results than using large convolutional kernels. The VGG16 network contains 13 convolutional layers and 3 fully connected layers, which is a relatively large network, with a total of approximately 138 million parameters, and is considered a very large network even by current standards. However, the structure of VGG-16 is not complex, and this network structure is very regular, with several convolutional layers followed by pooling layers that can compress the image size. The pooling layers reduce the height and width of the image. At the same time, there is a certain pattern in the change of the number of filters in the convolutional layers, doubling from 64 to 128, then to 256 and 512. The traditional VGG16 contains 16 hidden layers. The self-deepened VGG network in this patent has more convolutional layers and fully connected layers (capable of extracting more representative features): 15 convolutional layers and 5 fully connected layers. Experiments show that for such a deepened classification network, the detection accuracy has increased by 5%, and the speed has decreased by less than 0.5%. Therefore, the accuracy of such a deepened network has been significantly improved, while the speed decrease is not obvious.
[0047] Table 1 shows the structure of the self-deepened VGG16 network. The same size of convolutional kernel size (3x3) and maximum pooling size (2x2) are used throughout the network. The combination of several small filter (3x3) convolutional layers is better than a large filter (5x5 or 7x7) convolutional layer.
[0048] Table 1 Structure of the Self-Deepened VGG16 Network
[0049]
[0050]
[0051] When training the self - deepened VGG network, first, 8000 scanned document images in real scenes (with their respective brightness levels) are collected. Then, they are manually divided into 5 brightness levels (categories), and the label ID is an integer between [1, 5]. The learning rate is set to 5×10 -3 , and a weight decay of 0.8 is added to the learning rate of all models. The batch size is set to 40, and the Adam algorithm is used to optimize the loss function. After 55 iterations of training, the offline training for brightness classification is completed, and a brightness classification model is obtained. When performing online brightness classification, the original scanned document image is input into the deep - learning model for brightness classification, and 5 brightness level categories L are obtained. L is used as the key parameter for subsequent traditional Gamma brightness correction.
[0052] 1.2 Improvement of traditional Gamma image brightness correction
[0053] The improved correction rule is to set the exponential value of Gamma correction as an expression based on the result L obtained from deep learning: 1 / (5 - L + 1). Assuming the original image is I, the specific steps for correcting a certain pixel I(i, j) in it are as follows:
[0054] 1) Normalization. Convert I(i, j) to a real number between [0, 1] according to formula (1).
[0055] I(i, j) = (I(i, j)+0.5) / 256 (1)
[0056] 2) Pre - compensation. Complete pre - compensation according to the improved formula (2).
[0057] I(i, j) = I(i, j) 1 / (5-L+1) (2)
[0058] 3) Inverse normalization. According to formula (3), inverse - transform the pre - compensated real - valued I(i, j) to an integer value between [0, 255].
[0059] I(i, j) = I(i, j)×256 - 0.5 (3)
[0060] When performing brightness classification, brightness classification based on deep learning is adopted. Specifically, it is implemented using the self - deepened VGG network, including an offline training module for brightness classification and an online classification module for brightness.
[0061] 2. After annotating the main components of the scanned document and performing deep learning training to obtain a model, the main components of the scanned document are detected
[0062] YOLOv2 has made many improvements compared to YOLOv1, which also significantly improves the accuracy of YOLOv2. Moreover, YOLOv2 is still very fast and maintains its advantage as a one-stage object detection method. Therefore, this patent uses YOLOv2 for the detection of main components. The YOLOv2 algorithm uses a single convolutional network model to achieve end-to-end object detection. First, the input image is scaled to 448x448, then fed into the convolutional network, and finally the network prediction results are processed to obtain the detected objects. YOLOv2 uses a convolutional network to extract features and then uses fully connected layers to obtain prediction values. The network structure refers to the GooLeNet model and contains 24 convolutional layers and 2 fully connected layers. For the convolutional layers, 1x1 convolutions are mainly used to reduce the dimensionality of the feature channels, followed by 3x3 convolutions, but the linear activation function is used for the last layer. Batch Normalization is also an important feature of YOLOv2. Batch Normalization can not only improve the model convergence speed but also have a regularization effect, reducing the overfitting situation of the model. In YOLOv2, a batch normalization layer is added after each convolutional layer, and the Dropout layer is no longer used. After using batch normalization, the mAP of YOLOv2 has increased by 2.4%.
[0063] In terms of the training of YOLOv2, a joint training method for detection and classification is proposed. Using this joint training method, the YOLO9000 model is trained on the COCO detection dataset and the ImageNet classification dataset. It can detect more than 9000 types of objects. In this patent, through transfer learning, this deep learning model is used for the detection of 1 type of principal component template. For the annotation benchmark (Ground Truth) in the scanned image, if the center point of a certain annotation box falls within a certain cell, then the bounding boxes corresponding to the 5 prior boxes within that cell are responsible for predicting it. Specifically, which bounding box predicts it needs to be determined during training, that is, it is predicted by the bounding box with the largest IOU with the annotation benchmark, and the remaining 4 bounding boxes do not match this annotation benchmark. YOLOv2 draws on the prior box strategy of the Region Proposal Network (RPN) in Faster R-CNN. The RPN performs convolution on the feature map obtained by the CNN feature extractor to predict the bounding box and confidence (whether there is an object) at each position, and prior boxes of different scales and ratios are set at each position. Therefore, the Region Proposal Network predicts the offset value of the bounding box relative to the prior box. Using prior boxes makes it easier for the model to learn. So YOLOv2 removes the fully connected layer in YOLOv1 and uses convolution and anchor boxes to predict the bounding box. In the present invention, the loss function of YOLOv2 is changed to the cross-entropy loss function. In the present invention, a total of 8000 scanned images' main components are annotated. The attenuation coefficient, momentum parameter, and learning rate are set to 0.0003, 0.85, and 0.001 respectively, and the exponential decay mode is selected to update the learning rate. When the training iteration times reach 15000 and 17000, the learning rate is reduced to 15% and 7% of the initial learning rate respectively, so that the cross-entropy loss function converges further.
[0064] After training the principal component deep learning model, the scanned image enhanced in brightness is input into the principal component deep learning model to perform forward inference. For the detection box with a confidence higher than 0.5, it is considered as the main component of the current scanned image.
[0065] 3. The deep learning model performs deviation correction in 4 directions in the first stage (large-direction deviation correction)
[0066] When correcting the deviation of the image, a two-stage progressive deviation correction method is designed. In the first stage, 4-class classification is implemented based on the improved VGG16 network for preliminary deviation correction. In the second stage, 91-class classification is implemented based on the improved VGG16 network and combined with the traditional projection deviation correction for comprehensive judgment to complete fine deviation correction.
[0067] During the first-stage correction, a self-deepened VGG16 network is used for training. In terms of the production of the training dataset, 8,000 original images with a 0-degree skew are collected (if the original images are not at 0 degrees, they are manually corrected to 0 degrees using Photoshop image editing software). Then, using the rotate function in Python, a Python script is written to rotate the images to obtain image sets skewed at 90 degrees, 180 degrees, and 270 degrees. Combining with the original image set (1 class), a training image set for large-direction correction in 4 categories can be obtained, achieving the effect of semi-automatically constructing the training dataset. Thus, a deep learning network model for classifying the main component images in 4 directions is trained using the self-deepened VGG16 network. Furthermore, for each main component image, it is classified in 4 directions using this deep learning network for 4-direction classification to obtain the class label C (the value range is 0, 1, 2, 3). Then, according to the value of the class label C, the original image is rotated in the opposite direction, that is, counterclockwise by C×90 degrees, thus completing the correction in 4 directions in the first stage (large-direction correction).
[0068] 4. Deep learning combined with pixel projection correction (small-direction correction) to complete the correction in the second stage
[0069] In the second stage, 91-class classification is implemented based on the improved VGG16 network and comprehensively judged with the traditional projection correction to complete the fine correction.
[0070] 4.1 Deep learning model for 91-direction correction in the second stage (small-direction correction / fine correction)
[0071] For the 91-class classification during the second-stage correction, the specific implementation is as follows: A self-deepened VGG16 network is used for training. In terms of the production of the training dataset, 16,000 original images with a 0-degree skew are collected (if the original images are not at 0 degrees, they are manually corrected to 0 degrees using Photoshop image editing software). Then, using the rotate function in Python, a Python script is written to rotate the images clockwise to obtain an image set skewed at [1, 90] degrees. Combining with the original image set (1 class), a training image set for large-direction correction in 91 categories can be obtained, achieving the effect of semi-automatically constructing the training dataset. Thus, a deep learning network model for classifying the main component images in 91 directions is trained using the self-deepened VGG16 network. Furthermore, for each main component image, it is classified in 91 directions using this deep learning network for 91-direction classification to obtain the class label C2 (the range is 0, 1, 2, 3,..., 90). Then, according to the value of the class label C2, the original image is rotated in the opposite direction, that is, counterclockwise by C2 degrees, thus completing the correction in 4 directions in the first stage (large-direction correction).
[0072] 4.2 Improved Traditional Projection Correction (Small Direction Correction / Fine Correction)
[0073] For traditional projection correction, first, the detection graph is traversed for angular rotation, and then the projection of the original pixels in the horizontal direction is performed to find the direction with the most 0 values in the projection, which is considered the direction of image skew. Based on the traditional projection correction algorithm, this patent takes into account the relatively high noise in the original bank scan images and designs an improved Sobel segmentation to avoid projecting on the original pixels. Specifically, the improved Sobel segmentation method is as follows: 1) Grayscale the original RGB color image to obtain a grayscale image R; 2) Apply the 3×3 Sobel operator and the 5×5 Sobel operator respectively to obtain grayscale images R2 and R3, and add the pixels at the same positions of R2 and R3 with weights (the weight of R2 is 0.7 and the weight of R1 is 0.3) to obtain a new weighted grayscale image R4; 3) Perform watershed segmentation on R4 to obtain the final segmented image R5. This patent replaces R5 in the traditional projection correction algorithm to complete the improvement of the traditional projection correction. The improved algorithm can not only avoid being easily interfered by image noise when projecting based on the original pixels but also better adapt to different table and text edge sizes.
[0074] In the second stage, 91-class classification is implemented based on the improved VGG16 network and simultaneously combined with the improved traditional projection correction for comprehensive judgment to complete fine correction. The so-called comprehensive judgment means that it is carried out in two branches. The first branch is to use a self-deepened VGG deep learning network to implement 91-class small direction estimation, and the second branch is to improve the small direction estimation based on the Sobel edge projection method. When the results of the first branch and the second branch differ by less than 5 degrees, the result of the first branch is used as the final small correction angle estimation; otherwise, the result of the second branch is used as the final small correction angle estimation, and finally, the selection of the small correction result is completed. The reason for this is that the deep learning 91-direction correction in this patent is based on the correction result obtained from image classification. If it is accurate, it will be extremely accurate, but if it is inaccurate, the correction results may vary greatly; while the improved small direction estimation based on the Sobel edge projection method is the opposite. Therefore, it is proposed to jointly combine the small correction results obtained by these two methods for comprehensive judgment to obtain the optimal small correction result, thereby completing the correction of bank scan images.
Claims
1. An automatic image rectification method for archival scanned images based on deep learning, characterized in that in view of the imaging characteristics of archival scanned images, the deep learning brightness classification results are used to adaptively adjust the brightness of the images, improve the brightness of the scanned images, and through the deep learning object detection technology, the main components of the entire scanned image are detected to avoid the influence of black or white edges around the image on the subsequent image rectification angle estimation; a two-stage rectification strategy is designed, where the first stage is large-direction rectification and the second stage is small-direction rectification. The rectification of the scanned image is completed, specifically including: Step 1, use the deep learning brightness classification results to adjust the image brightness and improve the brightness of the scanned image; Step 2, label the main components of the scanned image, and after deep learning training to obtain a model, detect the main components of the scanned image; Step 3, use the main components as the input, and the deep learning model performs rectification in 4 directions in the first stage; Step 4, use deep learning combined with pixel projection rectification to complete the rectification in the second stage. The use of deep learning combined with pixel projection rectification to complete the rectification in the second stage means that the rectification in the second stage is carried out in two branches on the basis of the image obtained by the rectification in the first stage. The first branch is to use the self-deepened VGG deep learning classification network to implement 91-class small-direction estimation, and the second branch is small-direction estimation based on the Sobel edge projection method; when the difference between the results of the first branch and the second branch is less than 5 degrees, the result of the first branch is used as the final small rectification angle estimation, otherwise the result of the second branch is used as the final small rectification angle estimation, and finally the selection of the small rectification result is completed.
2. The method according to claim 1, characterized in that the deep learning brightness classification in Step 1 means collecting 8000 archival scanned images in real scenes, and then manually dividing the images into 5 brightness categories according to 5 brightness levels, and using the self-deepened VGG deep learning classification network for training to obtain an image brightness classification network model. The self-deepened VGG deep learning classification network has 15 convolutional layers and 5 fully connected layers, and then using this model, deep learning model inference is performed on the original input image to obtain the deep learning brightness classification result L of the test image, and the range of L is 1 to 5.
3. The method according to claim 2, characterized in that using the deep learning brightness classification result to adjust the image brightness in Step 1 means using the classification label L obtained by the self-deepened VGG deep learning classification network to adjust the brightness of the original scanned image, and the adjustment rule is to set the exponential value of Gamma correction to the formula based on the deep learning result L: 1 / (5 - L + 1).
4. The method according to claim 1, characterized in that After obtaining a model through deep learning training on the main components of the marked scans described in Step 2, detecting the main components of the scans refers to manually marking the image regions that only need to be retained in 8,000 original scans as the training data for subsequent deep learning main component detection; detecting the main components of the scans means using the marked main components to train the YOLOv2 deep learning object detection network to obtain a main component detection model, and performing forward inference of the deep network to obtain the main component image region of the original scanned image.
5. The method according to claim 1, wherein, using the main component as the input, the first-stage 4-direction rectification by the deep learning model means collecting 8,000 original images with a 0-degree skew. If the original image is not at 0 degrees, it is manually rectified to 0 degrees, and then by rotating the image clockwise, an image set with skews of 90 degrees, 180 degrees, and 270 degrees is obtained. Combining with the original image set, a 4-class image training set for large-direction rectification is obtained, and thus a deep learning network model for 4-direction classification of the main component image is trained using the self-deepened VGG deep learning classification network; furthermore, for each main component image, it is classified in 4 directions using this 4-direction classification deep learning network to obtain a class label C, where the value range of C is 0, 1, 2, or 3, and then by rotating the image counterclockwise, that is, counterclockwise by C×90 degrees, the first-stage 4-direction rectification is completed.
6. The method according to claim 1, wherein, The implementation of the self - deepening VGG deep - learning classification network for 91 - class small - direction estimation means collecting 16,000 original images with a 0 - degree skew. If the original image is not at 0 degrees, it is manually corrected to 0 degrees. Then, by rotating the original image, a set of main - component images with skews from 1 to 90 degrees is obtained. Combining with the original image set, a 91 - class image training set for large - direction correction is obtained; thus, using the self - deepening VGG deep - learning classification network, a deep - learning network model for classifying the main - component images in 91 directions is trained; furthermore, for each main - component image, using the self - deepening VGG deep - learning classification network to classify it in 91 directions, a class label C1 is obtained, where the range of C1 is an integer from 0 to 90, that is, the result of implementing 91 - class small - direction estimation using the self - deepening VGG deep - learning classification network is correspondingly obtained; the small - direction estimation based on the Sobel edge projection method means using an improved Sobel weighted segmentation algorithm to segment the main - component image, then performing angle - rotation traversal on the segmented image. Further, performing pixel - value projection in the horizontal direction on the pixels of the segmented image, and finding the correction direction with the most 0 - value projections, which is the correction angle obtained by the small - direction estimation based on the Sobel edge projection method; the improved Sobel weighted segmentation method is as follows: 1) Grayscale the original RGB color image to obtain a grayscale image R; 2) Respectively execute the 3×3 Sobel operator and the 5×5 Sobel operator to obtain grayscale images R2 and R3, and perform weighted addition on the pixels at the same positions of R2 and R3 to obtain a new weighted grayscale image R4; 3) Perform segmentation on R4 using the watershed segmentation algorithm to obtain the final segmented image R5.
Citation Information
Patent Citations
Text image correction method and device, storage medium and equipment
CN108681729A
Text tilt correction method and device, storage medium and computer equipment
CN113128495A