Document image anti-curling method based on deep learning key point positioning

By building lightweight DenseNet and Transformer models, the hardware dependence and complexity problems in the decurl process of document images are solved, and fast and efficient document image recovery is achieved, achieving quality similar to electronic scanning.

CN120340043AActive Publication Date: 2025-07-18CHENGDU HARIT MEDICAL TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510452687.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-18
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The prior art has strong hardware dependence, high algorithm complexity and limited effect in the process of decurling document images, making it difficult to achieve high-quality digital recovery without the help of stereo cameras or efficient 3D reconstruction algorithms.

Method used

A lightweight deep learning network is adopted to generate training data through 3D rendering, a key point positioning model based on DenseNet and Transformer is built, and a comprehensive loss function optimization model is used to realize the decurl of document images.

Benefits of technology

Without additional hardware support, model training and inference are efficient, and can quickly and accurately restore curled document images to similar quality to electronic scanning, improving the readability and digital archiving effect of document images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340043A_ABST
    Figure CN120340043A_ABST
Patent Text Reader

Abstract

The invention discloses a document image anti-curling method based on deep learning key point positioning, and relates to the technical field of digital image restoration, and the method comprises the following steps: S1, data preparation; s2, constructing a model; s3, model training; and S4, performing anti-curling reasoning. According to the method, training data is generated through 3D rendering, a lightweight deep learning network (DenseNet + Transform) is constructed, an anti-curling problem is converted into key point positioning, and a comprehensive loss function is adopted to optimize a model. According to the method, additional hardware is not needed, the model efficiency is high, the digital anti-curling of the document picture can be realized at a relatively high speed and relatively high accuracy under the condition of not using any auxiliary photography hardware and 3D reconstruction algorithm, and the method is suitable for large-scale document digital processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital image restoration, and particularly relates to a method for anti-curling document images through deep learning key point localization, which is particularly applicable to efficiently digitally restoring document images deformed due to curling, folding or wrinkling. Background Art

[0002] With the popularization of digital cameras and mobile devices, it has become a common way to record paper documents by taking pictures. However, the captured document images often have low quality due to physical deformations (such as curling, folding), and there is a large difference in quality from the document images recorded by electronic scanning. Therefore, it is necessary to digitally anti-curl the captured document images, but the existing anti-curling technologies still have the following problems: 1. Strong hardware dependence: Some methods need to rely on stereo cameras or structured light projectors to obtain the 3D information of the document, and then restore the 2D flat image according to the 3D information of the document, which makes the hardware cost high and the operation complex; 2. High consumption of computing resources: Some methods reconstruct 3D information through multi-view photos of the document or use deep learning methods to build an end-to-end image restoration network, and directly output the anti-curled image according to the document image. This method no longer depends on additional hardware, but requires an efficient 3D reconstruction algorithm. The end-to-end model based on deep learning has a large number of parameters, and consumes a large amount of resources during training and inference and has low inference efficiency; 3. Limited effect: Traditional image processing methods perform anti-curling through basic information such as illumination / shadow and text line position in the image, but can only handle simple deformations, and the anti-curling effect is very limited and it is difficult to handle complex scenarios.

[0003] Therefore, it is necessary to develop a lightweight deep learning algorithm that can achieve digital anti-curling of document pictures at a relatively fast speed and high accuracy without the help of any auxiliary photography hardware and 3D reconstruction algorithms. Summary of the Invention

[0004] 1. Technical problems to be solved: Aiming at the problems of strong hardware dependence, high algorithm complexity and limited anti-curling effect in the prior art, the present invention provides a lightweight deep learning anti-curling method, which can efficiently restore curled document images to a quality similar to that of electronically scanned documents.

[0005] 2. Technical solutions: To solve the above problems, the present invention adopts the following technical solutions.

[0006] A method for anti-curling document images based on deep learning key point localization, comprising the following steps: S1. Data Preparation: Transform the flat electronic scanned document image through 3D rendering technology to simulate curling, folding, and wrinkling, generating a distorted document image and its corresponding inverse transformation coordinates; S2. Model Construction: Build a lightweight deep learning network, including the following modules connected in sequence: Convolutional Neural Network (CNN) Feature Extraction Module: Based on the network block unit of DenseNet and the attention mechanism of Transformer, extract multi-scale features of the document image; Transformer Feature Modeling Module: Model the global features through stacked Transformer Encoder; Coordinate Prediction Module: Output a 40×40×2 inverse transformation coordinate matrix through an upsampling layer and a prediction head; S3. Model Training: Optimize the model using a comprehensive loss function including position loss, neighborhood relationship loss, and interval regression loss, where the expression of the loss function is: Total Loss = Localization Loss + 0.1·Neighbor Loss + 0.01·Interval Loss S4. Anti-Curling Inference: Input the document image to be processed, predict the inverse transformation coordinates through the model, and restore the distorted image to a flat image based on grid sampling.

[0007] Further improvement lies in that the CNN feature extraction module includes the following structures connected in sequence: DownSample block, consisting of a convolutional layer, a batch normalization layer, and a Relu activation function layer; DenseBlock, composed of multiple DenseLayers, and the output of each DenseLayer is concatenated with the previous input as the input of the next layer; Finally, a feature map with 256 channels is output.

[0008] Further improvement lies in that the Transformer feature modeling module consists of 8 stacked Transformer Encoders, the input and output dimensions of each Encoder are the same, and global feature modeling is achieved through the multi-head attention mechanism.

[0009] Further improvement lies in that in the coordinate prediction module: The upsampling layer uses the bilinear interpolation algorithm to increase the resolution of the feature map to the original image size; The prediction head includes two convolutional layers, a normalization layer, and a PReLU activation function layer, and outputs a 2-channel coordinate matrix.

[0010] A further improvement is that the calculation formula for the position loss is as follows: ; where SmoothL1 is defined as: ; P i is the predicted coordinate, g i is the ground truth coordinate, and N is the total number of samples.

[0011] A further improvement is that the neighborhood relationship loss includes a horizontal direction loss and a vertical direction loss, and their calculation formulas are respectively: ; .

[0012] A further improvement is that the interval regression loss includes a horizontal interval loss and a vertical interval loss, and their calculation formulas are respectively: ; .

[0013] A further improvement is that the input image size of the model is 320×320×3, and the output is an inverse transformation coordinate matrix of 40×40×2, which is restored to the original image resolution through grid sampling; 40×40×2 represents the grid coordinates of 40 rows × 40 columns, and each coordinate point contains two channel values in the horizontal (x) and vertical (y) directions.

[0014] 3. Beneficial effects: Adopting the technical solution provided by the present invention, compared with the prior art, it has the following beneficial effects: (1) Reduced hardware dependence: In the past, some digital anti-curling methods needed to rely on high-precision hardware such as stereo cameras and structured light projectors, while this patent does not require any auxiliary photography hardware and can achieve anti-curling only relying on the document image itself, reducing the hardware cost and usage threshold.

[0015] (2) Lightweight and efficient model: Traditional deep learning-based anti-curling algorithms usually have a large volume and consume a lot of resources during training and inference. The lightweight deep learning anti-curling algorithm constructed in this patent simplifies the problem to a key point positioning problem, making the model training and inference more efficient and low-cost, and can better meet the needs of processing a large number of document pictures in daily scenarios.

[0016] (3)Improved anti-curling effect: This patent can efficiently perform anti-curling processing on the captured curled, folded, or wrinkled document images, making them achieve an effect similar to electronically scanned documents, effectively bridging the quality gap between photo shooting and electronic scanning, greatly improving the quality of document images and the readability of text, which is of great significance for the digital archiving of document images.

[0017] It should be noted that the structures not introduced in the present invention are the same as the prior art or can be implemented using the prior art since they do not involve the design key points and improvement directions of the present invention, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is the overall architecture diagram of the anti-curling algorithm network of the present invention; Figure 2 is the network structure diagram of the Transformer Encoder; Figure 3 is the schematic diagram of the effect of the anti-curling process. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.

[0020] Embodiment: Please refer to Figures 1-3 , a method for anti-curling of document images based on deep learning key point localization. The technical solution adopted by the present invention to solve its technical problems is: by transforming the problem of anti-curling of digital document images into a key point localization problem, constructing an algorithm network that is lighter than previous anti-curling deep learning models, realizing more efficient and lower-cost model training and inference, and being more suitable for daily scenarios. The steps of the algorithm are as follows: 1. Data preparation: The training of the deep network requires a large amount of data support. Therefore, first, 3D rendering technology is used to transform the originally flat electronically scanned document images to simulate curling, folding, and wrinkling. Since all transformations are carried out by digital means, the inverse transformation coordinates of each document image can be obtained, and the distorted document can be restored to a flat scanned document through the coordinates.

[0021] 2. Establish the optimization objective: For each distorted document image, by predicting its inverse transformation coordinates, the image can be uncurled through the coordinates. Therefore, the optimization objective of the uncurling model proposed in the present invention is to predict a set of inverse transformation coordinates of 40*40*2 for each document image, that is, 1600 sets of horizontal and vertical coordinates, and uncurle the document image through the predicted coordinates. Each set of coordinates is actually a set of key points in the image.

[0022] 3. Model construction: A lightweight uncurling model with a single-tower structure is built through the combination of the Dense Convolutional Neural Network (DenseNet) and the Transformer, a sequence neural network based on the attention mechanism.

[0023] 4. Model training: All datasets are divided into a training set and a validation set in a ratio of 8:2. The training set is used to perform supervised training on the entire network, and the weight parameters with the best performance on the validation set are saved for uncurling inference of images.

[0024] The uncurling algorithm network for digital document images proposed in the present invention is mainly divided into three modules: the Convolutional Neural Network (CNN) feature extraction module, the Transformer.

[0025] Feature modeling module, coordinate prediction module. The size of the input image to the model is H*W*3, where H is the pixel height of the image, W is the pixel width of the image, and 3 is the number of channels of the RGB 3 channels.

[0026] 1. CNN feature extraction module: The CNN feature extraction module is mainly composed of the network block unit DenseBlock and the downsampling block in DenseNet, and its main function is to perform preliminary feature extraction and downsampling on the input image.

[0027] The downsampling block consists of a convolutional layer, a batch normalization layer, and a Relu activation function layer. For convenience, a downsampling block is denoted as DownSample in the following text.

[0028] DenseBlock consists of several DenseLayers. Each DenseLayer consists of a convolutional layer (1*1) for transforming the channel dimension, a conventional convolutional layer (3*3), a batch normalization layer, and a Relu activation function layer. The input of each DenseLayer is the output of the previous DenseLayer concatenated with the input of the previous DenseLayer.

[0029] The network structure and details of the entire CNN feature extraction module are as follows. For convenience, s represents the convolutional stride, k represents the convolutional kernel size, p represents the padding pixel size, i represents the input channels, and o represents the output channels.

[0030] (1)DownSample1: Input channels 3 → output channels 64, i = 3, o = 64, k = 3, s = 2, p = 1; (2)DenseBlock1: Consists of 2 DenseLayers, i = 64, o = 128, s = 1, p = 1; (3)DownSample2: Input channels 128 → output channels 128, i = 128, o = 128, k = 3, s = 2, p = 1; (4)DenseBlock2: Consists of 6 DenseLayers, i = 128, o = 320, s = 1, p = 1; (5)DownSample3: i = 320, o = 320, k = 3, s = 2, p = 1; (6)DenseBlock2 (DenseLayer * 8): i = 320, o = 576, s = 1, p = 1; (7)DownSample3: i = 576, o = 576, k = 3, s = 2, p = 1; (8)DenseBlock2 (DenseLayer * 4): i = 576, o = 704, s = 1, p = 1; (9)Conv2D: i = 704, o = 256, k = 3, s = 1, p = 0.

[0031] Finally, the CNN feature extraction module outputs a feature map with 256 channels, mainly extracting local features of the image through multi-scale convolution.

[0032] 2. Transformer Feature Modeling Module: This module is composed of 8 stacked Transformer Encoders. The structure of the Transformer Encoder is as Figure 1 shown. Each layer contains a multi-head attention mechanism and a feed-forward network. With the unique attention mechanism of the Transformer, this module can model based on the global features of the input image. The input and output dimensions of each Transformer Encoder remain unchanged.

[0033] 3. Coordinate Prediction Module: The coordinate prediction module consists of an upsampling layer and a prediction head. Among them, the upsampling layer uses the bilinear interpolation algorithm to double the resolution of the feature map; the prediction head consists of two convolutional layers, a normalization layer, and a PReLU activation function layer. The parameter settings of the two convolutional layers in the prediction head are as follows: (1)Conv2D_1: i = 256, o = 64, k = 3, s = 1, p = 1; (2)Conv2D_2: i = 64, o = 2, k = 3, s = 1, p = 1; (3)Activation function: PReLU; (4)Output: 40×40×2 inverse transformation coordinate matrix (1600 coordinate points). The input feature map first passes through the first convolutional layer, then through the normalization layer and the PReLU layer, and finally through the second convolutional layer.

[0034] 4. Loss function: The model proposed in the present invention aims to optimize by reducing the gap between the predicted inverse transformation coordinates of the image and the true inverse transformation coordinates of the image. The loss function for model training is defined as three parts.

[0035] (1)Position loss, which is used to measure the position difference between the true coordinates and the predicted coordinates. Its mathematical formula is as follows: ; where N is the total number of samples, P i is the coordinate predicted by the model, and g i is the true coordinate.

[0036] ; (2)Neighborhood relationship loss, which is used to measure the difference between the neighbors of each coordinate point in the horizontal and vertical directions. Its mathematical formula is as follows: is the coordinate difference between the predicted coordinate point and its horizontal neighbor ; ; P i horiz is the coordinate difference between the predicted coordinate point and its horizontal neighbor; g i horiz is the coordinate difference between the true coordinate point and its horizontal neighbor; N is the total number of horizontal neighbor pairs.

[0037] ; P i vert is the coordinate difference between the predicted coordinate point and its vertical neighbor; g i vert is the coordinate difference between the true coordinate point and its vertical neighbor; N is the total number of vertical neighbor pairs.

[0038] (3) The interval regression loss is used to measure the difference in the interval distances between the predicted values and the true values of the coordinate points. Its mathematical formula is as follows: ; ; d i horiz d(p) is the total distance of the horizontal intervals of the predicted coordinate points; d i horiz d(g) is the total distance of the horizontal intervals of the true coordinate points; N is the total number of horizontal interval pairs.

[0039] ; d i vert d(p) is the total distance of the horizontal intervals of the predicted coordinate points; d i vert d(g) is the total distance of the horizontal intervals of the true coordinate points; N is the total number of horizontal interval pairs.

[0040] (4) During model training, the above three loss functions are used comprehensively.

[0041] ; In the specific model training, α is set to 0.1 and β is set to 0.01.

[0042] (5) Training parameters Optimizer: Adam (initial learning rate 1e-4, decay rate 0.9); Number of training epochs: 200 Epoch; Batch size: 32.

[0043] The overall architecture of the anti-curling network model proposed by the present invention is as Figure 2 shown. The input is a document image with a size of 320*320*3 (uniformly scaled to this size). The model will predict a coordinate matrix of 40*40*2 (1600 coordinate points) for it. By performing grid sampling on these 1600 coordinate points (sampling to the original image size), a flat document image can be restored. The overall process effect is as Figure 3 shown.

[0044] The above-described embodiments merely represent certain implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.

Claims

1. A document image anti-curling method based on deep learning key point localization, characterized in that It includes the following steps: S1. Data preparation: Transform the flat electronic scanned document image through 3D rendering technology to simulate curling, folding, and wrinkling, generating a distorted document image and its corresponding inverse transformation coordinates; S2. Model construction: Build a lightweight deep learning network, including the following modules connected in sequence: Convolutional neural network (CNN) feature extraction module, based on the network block unit of DenseNet and the attention mechanism of Transformer, to extract multi-scale features of the document image; Transformer feature modeling module, to model the global features through stacked Transformer Encoder; Coordinate prediction module, to output an inverse transformation coordinate matrix of 40×40×2 through an upsampling layer and a prediction head; S3. Model training: Optimize the model using a comprehensive loss function including position loss, neighborhood relationship loss, and interval regression loss, where the expression of the loss function is: Total Loss = Localization Loss + 0.1·Neighbor Loss + 0.01·Interval Loss S4. Anti-curling inference: Input the document image to be processed, predict the inverse transformation coordinates through the model, and restore the distorted image to a flat image based on grid sampling.

2. The document image anti-curling method based on deep learning key point positioning according to claim 1, characterized in that The CNN feature extraction module includes the following structures connected in sequence: DownSample block, composed of a convolutional layer, a batch normalization layer, and a Relu activation function layer; DenseBlock, composed of multiple DenseLayers, and the output of each DenseLayer is concatenated with the previous input as the input of the next layer; Finally, a feature map with 256 channels is output.

3. A document image anti-curling method based on deep learning key point localization according to claim 1, characterized in that, The Transformer feature modeling module consists of 8 stacked Transformer Encoders, and the input and output dimensions of each Encoder are the same, and global feature modeling is achieved through the multi-head attention mechanism.

4. A document image anti-curling method based on deep learning key point positioning according to claim 1, characterized in that, In the coordinate prediction module: The upsampling layer uses the bilinear interpolation algorithm to increase the resolution of the feature map to the original image size; The prediction head includes two convolutional layers, a normalization layer, and a PReLU activation function layer, and outputs a coordinate matrix with 2 channels.

5. A method for anti-curling of document images based on deep learning key point localization according to claim 1, characterized in that, The calculation formula of the position loss is: ; where SmoothL1 is defined as: 。 6. A document image anti-curling method based on deep learning key point localization according to claim 1, characterized in that, The neighborhood relationship loss includes horizontal direction loss and vertical direction loss, and the calculation formulas are respectively: ; 。 7. A method for anti - curling of document images based on deep - learning key - point localization according to claim 1, characterized in that, The interval regression loss includes horizontal interval loss and vertical interval loss, and the calculation formulas are respectively: ; 。 8. A document image anti-curling method based on deep learning key point positioning according to claim 1, characterized in that, The input image size of the model is 320×320×3, and the output is an inverse transformation coordinate matrix of 40×40×2, where 40×40 represents the number of rows and columns of the grid, and 2 represents the horizontal and vertical components of each coordinate point, and it is restored to the original image resolution through grid sampling.

Citation Information

Patent Citations

  • Document image shape correction method and system based on deep learning

    CN117975469A

  • Filling field character recognition method, device and equipment and readable storage medium

    CN118072318A

  • Document image correction method and device applied to distorted document

    CN118711191A

  • Medical document image correction and identification method and system based on grid points

    CN118918589A

  • Image correction method not restricted by document picture

    CN119251107A