Document image correction and discrimination method based on deep learning

Through the document image correction and discrimination method based on deep learning, combined with four-point detection, text line segmentation and skeleton extraction, efficient correction and discrimination of document images is achieved, and the inefficiency and cost increase caused by document image correction in traditional technology is solved.

CN120088793APending Publication Date: 2025-06-03BEIJING BAIGEFEICHI TECH LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510133667.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In traditional technology, the obtained full document images are uniformly corrected, resulting in low processing efficiency of downstream tasks for text detection and recognition and increasing system operation costs.

Method used

The document image correction and discrimination method based on deep learning is adopted, corner point coordinate information is obtained through four-point detection, background area is removed, text line segmentation and skeleton extraction are performed, text line distortion is judged, and correction and corresponding image correction operations are carried out.

Benefits of technology

It improves the accuracy of document image correction and discrimination, reduces the occupation of system resources, reduces the system operation cost, and improves the processing efficiency of downstream tasks of text detection and recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088793A_ABST
    Figure CN120088793A_ABST
Patent Text Reader

Abstract

The invention provides a document image correction and discrimination method based on deep learning, and belongs to the field of computer vision, and the method comprises the steps: carrying out the four-point detection of an obtained document image, obtaining the coordinate information of angular points, removing a background region in the document image, and obtaining a preprocessed document image; performing text line segmentation on the preprocessed document image based on a text detection segmentation algorithm to obtain a text line mask image; obtaining a maximum bounding rectangle of the text line mask image, performing cutting and skeleton extraction, obtaining a skeleton bounding rectangle, judging a distortion condition of the text line mask image, and obtaining text line correction judgment information; according to the text line correction judgment information corresponding to all the text lines, correction judgment of the preprocessed document image is carried out, and corresponding image correction operation is executed. According to the technical scheme, when the document image is judged as the distorted image, the image correction processing is performed, so that the cost consumption of the document image correction technology is reduced, and the efficient correction processing of the document image is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and particularly to a method for correcting and discriminating document images based on deep learning. Background Art

[0002] The document image correction technology is a key prerequisite for realizing downstream tasks such as text detection and OCR (Optical Character Recognition). In the traditional process of text detection and recognition, after uniformly correcting the obtained document images, the corresponding text detection steps are then executed.

[0003] However, correcting the document images that do not need to be corrected not only reduces the efficiency of text detection and recognition, but also increases the system operation cost. In the traditional technology, there is a lack of a step for image correction discrimination after obtaining the document images, and all document images are uniformly corrected, resulting in the random occupation of system resources and thus affecting the processing efficiency of downstream tasks.

[0004] Therefore, a method for correcting and discriminating document images based on deep learning is proposed. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a method for correcting and discriminating document images based on deep learning, so as to solve the problem that in the traditional technology, correcting all the obtained document images leads to low processing efficiency of downstream tasks such as text detection and waste of system operation cost.

[0006] An embodiment of the present invention provides a method for correcting and discriminating document images based on deep learning, including:

[0007] Obtain a document image;

[0008] Perform four-point detection on the document image to obtain the corner coordinate information of the document image;

[0009] Remove the background area in the document image according to the corner coordinate information to obtain a preprocessed document image;

[0010] Based on a text detection and segmentation algorithm, perform text line segmentation on the preprocessed document image to obtain a text line mask image;

[0011] Obtain the minimum bounding rectangle of the text line mask image, crop the minimum bounding rectangle, and perform skeleton extraction to obtain a skeleton bounding rectangle;

[0012] Judge the distortion situation of the text line mask image according to the height of the skeleton bounding rectangle to obtain text line correction discrimination information;

[0013] Based on the text line correction discrimination information corresponding to all text lines in the preprocessed document image, perform correction discrimination on the preprocessed document image and execute corresponding image correction operations.

[0014] Preferably, the present invention provides a method for correcting and discriminating a document image based on deep learning. The steps include: performing four-point detection on the document image to obtain the corner coordinate information of the document image.

[0015] Construct a corner detection model based on a backbone network layer, a sampling layer, and a feature coordinate calculation layer.

[0016] Input the document image into the corner detection model to obtain the corner coordinate information. Optionally, it includes:

[0017] Perform dimensionality-up convolution operation on the document image through the convolution expansion module in the backbone network layer to obtain a high-dimensional convolution feature image, and process the high-dimensional convolution feature image through the first batch normalization module, and perform non-linear feature extraction through the first activation function to obtain a high-dimensional feature image.

[0018] Perform convolution operation with a fixed number of channels on the high-dimensional feature image through the depth convolution module in the backbone network layer to obtain a depth convolution feature image, and perform non-linear feature extraction through the second batch normalization module and the first activation function to obtain a depth feature image.

[0019] Perform dimensionality-down convolution operation on the depth feature image through the convolution projection module in the backbone network layer to obtain a dimensionality-down feature image, and output a standard dimensionality-down feature image after being processed by the third batch normalization module.

[0020] Perform deconvolution processing on the standard dimensionality-down feature image through the sampling layer and set the number of channels to output a corresponding document feature heat map.

[0021] Perform activation operation on the document feature heat map through the second activation function in the feature coordinate calculation layer to obtain a standard feature heat map.

[0022]

[0023] Wherein, softmax is the second activation function, p ij is the element value at the i-th row and j-th column in the document feature heat map, w is the number of rows of element values in the document feature heat map, h is the number of columns of element values in the document feature heat map, p sv is the element value at the s-th row and v-th column in the document feature heat map.

[0024] Construct an X-axis constant coordinate matrix and a Y-axis constant coordinate matrix respectively according to the document feature heat map.

[0025] Calculate the corner coordinate information according to the standard feature heat map, the X-axis constant coordinate matrix, and the Y-axis constant coordinate matrix;

[0026]

[0027] where, (x a , y a ) is the corner coordinate information, x ij is the coordinate value corresponding to the element value p ij in the X-axis constant coordinate matrix, and y ij is the coordinate value corresponding to the element value p ij in the Y-axis constant coordinate matrix.

[0028] Preferably, the present invention provides a document image correction discrimination method based on deep learning, and the steps: constructing a corner detection model based on a backbone network layer, a sampling layer, and a feature coordinate calculation layer; further including:

[0029] Training the corner detection model with a training set, and when the comprehensive loss function meets the preset threshold requirement, using the corner detection model for the four-point detection of the document image; optionally, including:

[0030] Inputting the sample document image in the training set into the angle detection model for processing to obtain sample corner coordinate information;

[0031] Calculating a coordinate point loss function according to the corresponding standard corner coordinate information and the sample corner coordinate information of the sample document image;

[0032] L 1 =‖C a - C sa ‖ 2

[0033] where, L 1 is the coordinate point loss function, C a is the sample corner coordinate information, C sa is the standard corner coordinate information, and ‖‖ 2 is the Euclidean norm;

[0034] Obtaining the sample feature heat map corresponding to the sample document image through the angle detection model, and calculating a heat map loss function according to the corresponding sample standard feature heat map and the sample feature heat map of the sample document image;

[0035]

[0036] where, L 2 is the heat map loss function, D KL () is the KL divergence calculation formula, Pa is the heat map of sample features, P Sa is the standard feature heat map of the sample;

[0037] Construct a comprehensive loss function according to the coordinate point loss function and the heat map loss function;

[0038] L = L 1 + μL 2

[0039] where L is the comprehensive loss function and μ is a hyperparameter;

[0040] When the comprehensive loss function meets the preset threshold requirement, use the corner detection model for four-point detection of the document image.

[0041] Preferably, the present invention provides a method for correcting and discriminating a document image based on deep learning. The steps include: removing the background area in the document image according to the corner coordinate information to obtain a preprocessed document image;

[0042] Construct a circumscribed rectangle of the document area in the document image according to the corner coordinate information;

[0043] Construct a perspective transformation matrix according to the coordinate parameters of the circumscribed rectangle of the document area;

[0044] Based on the perspective transformation matrix, map the document area to a new canvas to obtain a preprocessed document image and remove the background area in the document image.

[0045] Preferably, the present invention provides a method for correcting and discriminating a document image based on deep learning. The steps include: performing text line segmentation on the preprocessed document image based on a text detection and segmentation algorithm to obtain a text line mask image;

[0046] Construct a text detection and segmentation model according to a feature extraction network, a sampling network, and a differentiable binarization network, and train the text detection and segmentation model through a sample set;

[0047] Perform text line segmentation on the preprocessed document image through the trained text detection and segmentation model to obtain a text line mask image. Optionally, it includes:

[0048] Extract features from the preprocessed document image through the feature extraction network to obtain feature convolution maps of different scales;

[0049] Perform corresponding upsampling operations on the feature convolution maps of different scales through the sampling network, obtain feature submaps of the same size for feature fusion, and generate a feature fusion map;

[0050] Detect the probability value of each pixel point in the feature fusion map belonging to the text box area to generate a feature probability map; calculate the binary threshold of each pixel point in the feature fusion map to obtain a feature threshold map;

[0051] Process the feature threshold map and the feature probability map through a differentiable binarization network to obtain an approximate binary image;

[0052]

[0053] where, P i ′ j is the pixel point at the i-th row and j-th column in the approximate binary image, l is the magnification factor, and u ij is the pixel point at the i-th row and j-th column in the preprocessed document image, and S ij is the pixel point at the i-th row and j-th column in the feature threshold map;

[0054] Obtain the text box area in the preprocessed document image according to the approximate binary image;

[0055] Perform dilation and contraction operations within the text box area for compensation and optimization processing to obtain the boundary of the text box area;

[0056] G = F - t·M 2 ·P ′

[0057] where, G is the text box area, F is the standard binary image obtained by threshold processing according to the feature probability map, t is the empirical coefficient, M is the connected region recognized within the text box area, and P ′ is the approximate binary image;

[0058] Perform text segmentation operations according to the boundary of the text box area recognized in the preprocessed text image to obtain a text line mask image.

[0059] Preferably, the present invention provides a document image correction and discrimination method based on deep learning. The steps are as follows: construct a text detection and segmentation model according to a feature extraction network, a sampling network, and a differentiable binarization network, and train the text detection and segmentation model through a sample set; including:

[0060] Obtain the preprocessed sample images in the sample set;

[0061] Process the preprocessed sample images through the text detection and segmentation model to obtain a sample feature probability map, a sample feature threshold map, and a sample approximate binary image;

[0062] Calculate the model loss function according to the standard feature probability map, the standard feature threshold map, and the standard approximate binary image corresponding to the preprocessed sample images in the sample set;

[0063] L m = L g + θ 1 L b + θ 2 L p

[0064] Wherein, L m is the model loss function, L g is the probability loss function of the standard feature probability map and the sample feature probability map, L b is the threshold loss function of the standard feature threshold map and the sample feature threshold map, L p is the binary loss function of the standard approximate binary image and the sample approximate binary image, θ 1 is the first adjustment parameter, θ 2 is the second adjustment parameter;

[0065]

[0066] Wherein, g Si is the standard feature probability map corresponding to the i-th preprocessed sample image in the sample set K, g i is the sample feature probability map obtained by the text detection and segmentation model according to the i-th preprocessed sample image;

[0067]

[0068] Wherein, b i is the sample feature threshold map obtained by the text detection and segmentation model according to the i-th preprocessed sample image, b Si is the standard feature threshold map corresponding to the i-th preprocessed sample image in the sample set K;

[0069]

[0070] Wherein, p i is the sample approximate binary image obtained by the text detection and segmentation model according to the i-th preprocessed sample image, p Si is the standard approximate binary image corresponding to the i-th preprocessed sample image in the sample set K;

[0071] Train the text detection and segmentation model according to the model loss function, optimize the control parameters of the feature extraction network, the sampling network and the differentiable binarization network, and when the model loss function meets the loss function threshold, use the text detection and segmentation model for text line segmentation of the preprocessed document image.

[0072] Preferably, the present invention provides a method for correcting and discriminating document images based on deep learning. The steps are as follows: judging the distortion of the text line mask image according to the height of the circumscribed rectangle of the skeleton, and obtaining text line correction discrimination information, including:

[0073] Obtaining the rectangular height data of the circumscribed rectangle of the skeleton;

[0074] Performing segmented detection on the circumscribed rectangle of the skeleton in the height direction to obtain differential height data;

[0075] When the height variance data between the rectangular height data and the differential height data is less than the first preset threshold error, marking the text line mask image as not requiring correction processing;

[0076] When the height variance data between the rectangular height data and the differential height data is greater than the second preset threshold error, marking the text line mask image as requiring correction processing.

[0077] Preferably, the present invention provides a method for correcting and discriminating document images based on deep learning. The steps are as follows: performing segmented detection on the circumscribed rectangle of the skeleton in the height direction to obtain differential height data, including:

[0078] Obtaining the rectangular bottom coordinate data of the circumscribed rectangle of the skeleton;

[0079] Taking the rectangular bottom coordinate data as the starting point, and sequentially selecting a number of height coordinate data in the height direction of the circumscribed rectangle of the skeleton;

[0080] Calculating the differential height data of the circumscribed rectangle of the skeleton according to the rectangular bottom coordinate data and the height coordinate data.

[0081] Preferably, the present invention provides a method for correcting and discriminating document images based on deep learning. The steps are as follows: performing correction discrimination on the preprocessed document image according to the text line correction discrimination information corresponding to all text lines in the preprocessed document image, and performing corresponding image correction operations, including:

[0082] Obtaining the text line correction discrimination information corresponding to all text lines in the preprocessed document image;

[0083] Obtaining the text content within the text line and calculating the text line information amount;

[0084] Obtaining all the text content in the preprocessed document image, obtaining the text information amount, and calculating the proportion of the text line information amount in the text information amount to obtain the information amount weight corresponding to the text line;

[0085] Obtaining the line correction parameter of the text line according to the information amount weight and the corresponding text line correction discrimination information;

[0086] Statistically analyze the line correction parameters corresponding to all text lines in the preprocessed document image to obtain text correction parameters. When the correction threshold parameter is exceeded, perform correction processing on the preprocessed document image. When the correction threshold parameter is not exceeded, there is no need to perform correction processing on the preprocessed document image.

[0087] Compared with the traditional technology, the beneficial effects of the present invention are as follows: A method for correcting and discriminating document images based on deep learning realizes the acquisition of text line correction discrimination information for all text lines in the preprocessed document image, and further realizes the acquisition of the correction discrimination result of the document image, solving the problem in the traditional technology that unified correction processing is performed on the acquired full-volume document image, resulting in low processing efficiency for downstream tasks of text detection and recognition. At the same time, it also reduces the situation of random occupation of system resources, reduces the system operation cost. When the document image is judged as a distorted image, image correction processing is performed. When the document image is judged as a flat image, no correction processing is required, reducing the cost consumption of the document image correction technology, and realizing the effective distinction between flat images and distorted images to ensure the efficient correction processing of document images in subsequent steps.

[0088] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification, or be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings.

[0089] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] Figure 1 It is a schematic flowchart of a method for correcting and discriminating document images based on deep learning provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0091] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described here are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0092] Embodiment 1:

[0093] The embodiment of the present invention provides a method for correcting and discriminating document images based on deep learning, referring to Figure 1 , including:

[0094] Obtain a document image;

[0095] Perform four-point detection on the document image to obtain the corner coordinate information of the document image;

[0096] Remove the background region in the document image according to the corner coordinate information to obtain a preprocessed document image;

[0097] Perform text line segmentation on the preprocessed document image based on a text detection and segmentation algorithm to obtain a text line mask image;

[0098] Obtain the minimum bounding rectangle of the text line mask image, crop the minimum bounding rectangle, and perform skeleton extraction to obtain a skeleton bounding rectangle;

[0099] Judge the distortion condition of the text line mask image according to the height of the skeleton bounding rectangle to obtain text line correction discrimination information;

[0100] Perform correction discrimination on the preprocessed document image according to the text line correction discrimination information corresponding to all text lines in the preprocessed document image, and perform corresponding image correction operations.

[0101] In the above embodiments, four-point detection is performed on the obtained document image to obtain the corner coordinate information of the document image, and the background region in the document image is removed to obtain a preprocessed document image;

[0102] Perform text line segmentation on the preprocessed document image based on a text detection and segmentation algorithm to obtain a text line mask image, obtain the minimum bounding rectangle of the text line mask image, crop the minimum bounding rectangle, and perform skeleton extraction to obtain a skeleton bounding rectangle, and judge the distortion condition of the text line mask image according to the height of the skeleton bounding rectangle to obtain text line correction discrimination information;

[0103] Perform correction discrimination on the preprocessed document image according to the text line correction discrimination information corresponding to all text lines in the preprocessed document image, and perform corresponding image correction operations.

[0104] In the above embodiments, the document image correction discrimination method based on deep learning proposed by the present invention is trained on the classification model of YOLOv8n.

[0105] The beneficial effects of the above technology are as follows: By performing four-point detection on the document image, the corner coordinate information of the document image is obtained, the background area in the document image is removed, and the preprocessed document image is obtained; based on the text detection and segmentation algorithm, the preprocessed document image is segmented into text lines, the text line mask image is obtained, and the maximum bounding rectangle is constructed for the text line mask image, and then cropping and skeleton extraction are performed to obtain the skeleton bounding rectangle. The distortion condition of the text line mask image is judged according to the height of the skeleton bounding rectangle, and the acquisition of text line correction discrimination information is realized; by obtaining the text line correction discrimination information corresponding to all text lines in the preprocessed document image, the correction discrimination of the preprocessed document image is realized, and the corresponding image correction operation is performed according to the correction discrimination result; compared with the traditional technology, the document image correction discrimination method based on deep learning proposed by the present invention realizes the acquisition of text line correction discrimination information for all text lines in the preprocessed document image, and then realizes the acquisition of the document image correction discrimination result, solves the problem that the traditional technology performs unified correction processing on the obtained full-volume document image, resulting in low processing efficiency of the downstream tasks of text detection and recognition, and at the same time reduces the situation of random occupation of system resources, reduces the system operation cost, performs image correction processing when the document image is judged to be a distorted image, and does not need to perform correction processing when the document image is judged to be a flat image, reduces the cost consumption of the document image correction technology, and realizes the effective distinction between flat images and distorted images to ensure the efficient correction processing of the document image in the subsequent steps.

[0106] Embodiment 2:

[0107] The embodiment of the present invention provides a document image correction discrimination method based on deep learning. The steps are as follows: Perform four-point detection on the document image to obtain the corner coordinate information of the document image; including:

[0108] Construct a corner detection model based on a backbone network layer, a sampling layer, and a feature coordinate calculation layer;

[0109] Input the document image into the corner detection model to obtain the corner coordinate information; optionally, including:

[0110] Perform a dimensionality-increasing convolution operation on the document image through the convolution expansion module in the backbone network layer to obtain a high-dimensional convolution feature image, and process the high-dimensional convolution feature image through the first batch normalization module, and perform non-linear feature extraction through the first activation function to obtain a high-dimensional feature image;

[0111] Perform a convolution operation with a fixed number of channels on the high-dimensional feature image through the depth convolution module in the backbone network layer to obtain a depth convolution feature image, and perform non-linear feature extraction through the second batch normalization module and the first activation function to obtain a depth feature image;

[0112] Perform a dimensionality reduction convolution operation on the depth feature image through the convolution projection module in the backbone network layer to obtain a dimensionality reduction feature image, and after processing through the third batch normalization module, output a standard dimensionality reduction feature image;

[0113] Perform a transposed convolution operation on the standard dimensionality reduction feature image through the sampling layer, and set the number of channels to output the corresponding document feature heatmap;

[0114] Perform an activation operation on the document feature heatmap through the second activation function in the feature coordinate calculation layer to obtain a standard feature heatmap;

[0115]

[0116] where softmax is the second activation function, and p ij is the element value at the i-th row and j-th column in the document feature heatmap, w is the number of rows of the element values in the document feature heatmap, h is the number of columns of the element values in the document feature heatmap, and p sv is the element value at the s-th row and v-th column in the document feature heatmap;

[0117] Construct an X-axis constant coordinate matrix and a Y-axis constant coordinate matrix respectively according to the document feature heatmap;

[0118] Calculate the corner coordinate information according to the standard feature heatmap, the X-axis constant coordinate matrix, and the Y-axis constant coordinate matrix;

[0119]

[0120] where (x a , y a ) is the corner coordinate information, x ij is the coordinate value corresponding to the element value p ij in the X-axis constant coordinate matrix, and y ij is the coordinate value corresponding to the element value p ij in the Y-axis constant coordinate matrix.

[0121] In the above embodiments, a corner detection model is constructed based on the backbone network layer, the sampling layer, and the feature coordinate calculation layer.

[0122] In the above embodiments, the convolutional expansion module in the backbone network layer performs upsampling convolutional operations on the document image to obtain a high-dimensional convolutional feature image, processes the high-dimensional convolutional feature image through the first batch normalization module, and performs non-linear feature extraction through the first activation function to obtain a high-dimensional feature image. The depth convolutional module performs convolutional operations with a fixed number of channels on the high-dimensional feature image to obtain a depth convolutional feature image, and performs non-linear feature extraction through the second batch normalization module and the first activation function to obtain a depth feature image. The convolutional projection module performs downsampling convolutional operations on the depth feature image to obtain a downsampled feature image, and outputs a standard downsampled feature image after being processed by the third batch normalization module. The standard downsampled feature image is deconvolved through the sampling layer, and the number of channels is set to output the corresponding document feature heat map. The X-axis constant coordinate matrix and the Y-axis constant coordinate matrix are respectively constructed based on the document feature heat map. The corner coordinate information is calculated based on the standard feature heat map, the X-axis constant coordinate matrix, and the Y-axis constant coordinate matrix.

[0123] In the above embodiments, the backbone network layer adopts an inverted residual structure. The convolutional expansion module performs upsampling convolutional operations on the document image to obtain a high-dimensional convolutional feature image for extracting deep feature information. The depth convolutional module performs convolutional operations independently in each channel, which can effectively reduce the computational complexity and the number of learning parameters of the network, realizing the lightweight of the backbone network layer. The first activation function is implemented as ReLU6, making the backbone network layer have stronger robustness.

[0124] In the above embodiments, the feature coordinate calculation layer respectively constructs the X-axis constant coordinate matrix and the Y-axis constant coordinate matrix based on the document feature heat map, activates the maximum and minimum values in the document feature heat map through the second activation function softmax to obtain the standard feature heat map, uses the values at each pixel position in the standard feature heat map as the weights of the corresponding coordinate positions, and calculates the corner coordinate information by taking the expectation with the axis coordinates.

[0125] The beneficial effects of the above technology are as follows: The lightweight backbone network layer is used to extract features from the document image, realizing the acquisition of the standard downsampled feature image. Through the sampling layer and the feature coordinate calculation layer, the corner coordinates of the standard downsampled feature image are predicted, realizing the acquisition of the corner coordinate information by the corner detection model. Compared with the traditional deep learning model, the corner detection model proposed by the present invention adopts a lightweight structure, improving the efficiency of feature extraction from the document image. The corner coordinate information is obtained by the feature coordinate calculation layer based on the maximum and minimum points of the heat map in combination with the X-axis constant coordinate matrix and the Y-axis constant coordinate matrix, improving the accuracy of corner coordinate acquisition.

[0126] Embodiment 3:

[0127] An embodiment of the present invention provides a method for correcting and discriminating document images based on deep learning. The steps are as follows: constructing a corner detection model based on a backbone network layer, a sampling layer, and a feature coordinate calculation layer; further including:

[0128] Training the corner detection model with a training set, and when the comprehensive loss function meets the preset threshold requirement, using the corner detection model for four-point detection of the document image. Optionally, it includes:

[0129] Inputting the sample document image in the training set into an angle detection model for processing to obtain sample corner coordinate information;

[0130] Calculating a coordinate point loss function according to the corresponding standard corner coordinate information and the sample corner coordinate information of the sample document image;

[0131] L 1 =‖C a -C sa ‖ 2

[0132] where L 1 is the coordinate point loss function, C a is the sample corner coordinate information, C sa is the standard corner coordinate information, and ‖‖ 2 is the Euclidean norm;

[0133] Obtaining a sample feature heat map corresponding to the sample document image through the angle detection model, and calculating a heat map loss function according to the corresponding sample standard feature heat map and the sample feature heat map of the sample document image;

[0134]

[0135] where L 2 is the heat map loss function, D KL () is the KL divergence calculation formula, P a is the sample feature heat map, and P Sa is the sample standard feature heat map;

[0136] Constructing a comprehensive loss function according to the coordinate point loss function and the heat map loss function;

[0137] L=L 1 +μL 2

[0138] where L is the comprehensive loss function and μ is a hyperparameter;

[0139] When the comprehensive loss function meets the preset threshold requirement, using the corner detection model for four-point detection of the document image.

[0140] In the above embodiments, the angle detection model is trained using the sample document images in the training set, and when the comprehensive loss function meets the preset threshold requirement, the corner detection model is used for the four-point detection of the document image.

[0141] In the above embodiments, the original sample images are obtained, and the original sample images are converted into fixed sizes through computer vision software tools to obtain the sample document images.

[0142] In the above embodiments, the comprehensive loss function is constructed by the coordinate point loss function and the heatmap loss function. The coordinate point loss function is calculated based on the deviation between the standard corner coordinate information and the sample corner coordinate information; the heatmap loss function measures the similarity between the sample standard feature heatmap and the sample feature heatmap, reflecting the accuracy of the sample feature heatmap prediction.

[0143] In the above embodiments, the heatmap loss function is an asymmetric measure of the difference between two probability distributions of the sample standard feature heatmap and the sample feature heatmap based on the KL divergence calculation formula.

[0144] In the above embodiments, the coordinate point loss function and the heatmap loss function are weighted and constrained by hyperparameters, and the comprehensive loss function can reflect the corner coordinate detection performance of the angle detection model.

[0145] The beneficial effects of the above technology are as follows: The training of the angle detection model is optimized through the training set, and the accuracy of the corner coordinates predicted by the corner detection model is judged by the comprehensive loss function. When the comprehensive loss function meets the preset threshold requirement, that is, when the model prediction ability of the corner detection model reaches the requirement, the corner detection model is used for the four-point detection of the document image.

[0146] Embodiment 4:

[0147] The embodiment of the present invention provides a method for correcting and discriminating a document image based on deep learning. Steps: Remove the background area in the document image according to the corner coordinate information to obtain a preprocessed document image; including:

[0148] Construct an outer rectangle of the document area in the document image according to the corner coordinate information;

[0149] Construct a perspective transformation matrix according to the coordinate parameters of the outer rectangle of the document area;

[0150] Based on the perspective transformation matrix, map the document area to a new canvas to obtain a preprocessed document image and remove the background area in the document image.

[0151] In the above embodiments, according to the corner coordinate information, an external rectangle of the document area is constructed in the document image, and a perspective transformation matrix is constructed based on the coordinate parameters of the external rectangle of the document area. Based on the perspective transformation matrix, the document area is mapped to a new canvas to obtain a preprocessed document image.

[0152] The beneficial effects of the above technology are as follows: By constructing an external rectangle of the document area according to the corner coordinate information and constructing a perspective transformation matrix based on the coordinate parameters of the external rectangle, the mapping transformation of the document area in the document image is realized, and a preprocessed document image is generated to remove the background area in the document image.

[0153] Embodiment 5:

[0154] An embodiment of the present invention provides a method for correcting and discriminating a document image based on deep learning. The steps are as follows: Based on a text detection and segmentation algorithm, the preprocessed document image is segmented into text lines to obtain a text line mask image, including:

[0155] According to a feature extraction network, a sampling network, and a differentiable binarization network, a text detection and segmentation model is constructed, and the text detection and segmentation model is trained through a sample set;

[0156] Through the trained text detection and segmentation model, a text line segmentation operation is performed on the preprocessed document image to obtain a text line mask image. Optionally, it includes:

[0157] The preprocessed document image is subjected to feature extraction through a feature extraction network to obtain feature convolution maps of different scales;

[0158] Through the sampling network, corresponding upsampling operations are performed on the feature convolution maps of different scales, and feature submaps of the same size are obtained for feature fusion to generate a feature fusion map;

[0159] The probability value that each pixel point in the feature fusion map belongs to the text box area is detected to generate a feature probability map; the binarization threshold of each pixel point in the feature fusion map is calculated to obtain a feature threshold map;

[0160] Through the differentiable binarization network, processing is performed according to the feature threshold map and the feature probability map to obtain an approximate binary image;

[0161]

[0162] Among them, P i ′ j is the pixel point at the i-th row and j-th column in the approximate binary image, l is the magnification factor, and u ij is the pixel point at the i-th row and j-th column in the preprocessed document image, and S ij is the pixel point at the i-th row and j-th column in the feature threshold map;

[0163] Obtain the text box area in the preprocessed document image according to the approximate binary image;

[0164] Perform dilation and contraction operations within the text box area, perform compensation and optimization processing, and obtain the boundary of the text box area;

[0165] G = F - t·M 2 ·P ′

[0166] where G is the text box area, F is the standard binary image obtained by threshold processing according to the feature probability map, t is the empirical coefficient, M is the connected area recognized within the text box area, and P ′ is the approximate binary image;

[0167] Perform text segmentation operations according to the boundary of the text box area recognized in the preprocessed text image to obtain the text line mask image.

[0168] In the above embodiments, a text detection and segmentation model is constructed through a feature extraction network, a sampling network, and a differentiable binarization network, and the text detection and segmentation model is trained through a sample set. The trained text detection and segmentation model performs text line segmentation operations on the preprocessed document image to obtain the text line mask image.

[0169] In the above embodiments, the text detection and segmentation model extracts features from the preprocessed document image through the feature extraction network to obtain feature convolution maps of different scales, performs corresponding upsampling operations on the feature convolution maps of different scales through the sampling network, obtains feature submaps of the same size for feature fusion, and generates a feature fusion map; detects the probability value of each pixel point in the feature fusion map belonging to the text box area to generate a feature probability map; calculates the binarization threshold of each pixel point in the feature fusion map to obtain a feature threshold map; processes the feature threshold map and the feature probability map through the differentiable binarization network to obtain an approximate binary image, and further obtains the text box area in the preprocessed document image; perform dilation and contraction operations within the text box area, perform compensation and optimization processing, obtain the boundary of the text box area, and perform text segmentation operations according to the boundary of the text box area recognized in the preprocessed text image to obtain the text line mask image.

[0170] In the above embodiments, the text detection and segmentation model is constructed based on DBNet (Differentiable BinarizationNet).

[0171] In the above embodiments, the differentiable binarization network is constructed based on the Tanh function, which can continuously expand the feature effect during the loop process, make the gradient change faster, reduce the number of iterations, and improve the performance of text detection.

[0172] In the above embodiments, dilation and contraction operations are performed within the text box area to adapt to text box areas of different sizes and shapes, and compensation and optimization processing are carried out, so that the boundaries of the obtained text box areas can be more accurately separated from the text lines, optimizing the text segmentation effect.

[0173] The beneficial effects of the above technology are as follows: Through the feature extraction network, sampling network, and differentiable binarization network in the text detection model, the acquisition of an approximate binary image is achieved, and based on the dilation and contraction operations within the text box area, compensation and optimization processing are carried out, the acquisition of the boundaries of the text box area is achieved, text segmentation operations are performed, and the acquisition of the text line mask image is achieved.

[0174] Embodiment 6:

[0175] The embodiment of the present invention provides a method for correcting and discriminating document images based on deep learning. The steps are as follows: A text detection and segmentation model is constructed according to the feature extraction network, sampling network, and differentiable binarization network, and the text detection and segmentation model is trained through a sample set, including:

[0176] Obtain the preprocessed sample images in the sample set;

[0177] Process the preprocessed sample images through the text detection and segmentation model to obtain the sample feature probability map, sample feature threshold map, and sample approximate binary image;

[0178] Calculate the model loss function according to the standard feature probability map, standard feature threshold map, and standard approximate binary image corresponding to the preprocessed sample images in the sample set;

[0179] L m =L g +θ 1 L b +θ 2 L p

[0180] Among them, L m is the model loss function, L g is the probability loss function between the standard feature probability map and the sample feature probability map, L b is the threshold loss function between the standard feature threshold map and the sample feature threshold map, L p is the binary loss function between the standard approximate binary image and the sample approximate binary image, θ 1 is the first adjustment parameter, θ 2 is the second adjustment parameter;

[0181]

[0182] Among them, g Siis the standard feature probability map corresponding to the i-th preprocessed sample image in the sample set K, g i is the sample feature probability map obtained by the text detection and segmentation model according to the i-th preprocessed sample image;

[0183]

[0184] where b i is the sample feature threshold map obtained by the text detection and segmentation model according to the i-th preprocessed sample image, b Si is the standard feature threshold map corresponding to the i-th preprocessed sample image in the sample set K;

[0185]

[0186] where p i is the sample approximate binary image obtained by the text detection and segmentation model according to the i-th preprocessed sample image, p Si is the standard approximate binary image corresponding to the i-th preprocessed sample image in the sample set K;

[0187] Train the text detection and segmentation model according to the model loss function, optimize the control parameters of the feature extraction network, sampling network and differentiable binarization network. When the model loss function meets the loss function threshold, use the text detection and segmentation model for text line segmentation of the preprocessed document image.

[0188] In the above embodiments, train the text detection and segmentation model through the preprocessed sample images in the sample set, optimize the control parameters of the feature extraction network, sampling network and differentiable binarization network in the text detection and segmentation model. When the model loss function meets the loss function threshold, use the text detection and segmentation model for text line segmentation of the preprocessed document image.

[0189] In the above embodiments, the model loss function is constructed by the probability loss function of the standard feature probability map and the sample feature probability map, the threshold loss function of the standard feature threshold map and the sample feature threshold map, and the binary loss function of the standard approximate binary image and the sample approximate binary image.

[0190] In the above embodiments, the preprocessed sample images in the sample set correspond to a standard feature probability map, a standard feature threshold map and a standard approximate binary image; the text detection and segmentation model outputs a sample feature probability map, a sample feature threshold map and a sample approximate binary image according to the preprocessed sample image.

[0191] The beneficial effects of the above technology are as follows: By training the text detection and segmentation model with a sample set, when the model loss function meets the loss function threshold, the text detection and segmentation model is used for text line segmentation of the preprocessed document image to improve the text line segmentation performance of the model; the model loss function is constructed based on the probability loss function, the threshold loss function, and the binary loss function to improve the training efficiency of the text detection and segmentation model and the optimization effect of the control parameters.

[0192] Embodiment 7:

[0193] The embodiment of the present invention provides a method for correcting and discriminating a document image based on deep learning. The steps are as follows: Judging the distortion situation of the text line mask image according to the height of the outer rectangle of the skeleton, and obtaining the text line correction and discrimination information; including:

[0194] Obtaining the rectangular height data of the outer rectangle of the skeleton;

[0195] Performing segmented detection on the outer rectangle of the skeleton in the height direction to obtain differential height data;

[0196] When the height variance data of the rectangular height data and the differential height data is less than the first preset threshold error, the text line mask image is marked as not requiring correction processing;

[0197] When the height variance data of the rectangular height data and the differential height data is greater than the second preset threshold error, the text line mask image is marked as requiring correction processing.

[0198] In the above embodiments, the rectangular height data and the differential height data of the outer rectangle of the skeleton are obtained, the height variance data is calculated, when the height variance data is less than the first preset threshold error, the text line mask image is marked as not requiring correction processing, and when the height variance data of the rectangular height data and the differential height data is greater than the second preset threshold error, the text line mask image is marked as requiring correction processing.

[0199] In the above embodiments, the text line correction and discrimination information includes a mark for not requiring correction processing or requiring correction processing of the text line mask image.

[0200] In the above embodiments, the first preset threshold error and the second preset threshold error are set by the staff, realizing the real-time adjustment of the text line correction strategy by the staff to meet the correction applications under different requirements.

[0201] The beneficial effects of the above technology are as follows: By calculating the rectangular height data and the differential height data of the outer rectangle of the skeleton corresponding to the text line mask image, based on the first preset threshold error and the second preset threshold error, the judgment of the distortion situation of the text line is realized, and the text line mask image is marked accordingly, realizing the acquisition of the text line correction and discrimination information.

[0202] Example 8:

[0203] An embodiment of the present invention provides a method for correcting and discriminating document images based on deep learning. The steps are as follows: perform segmented detection on the circumscribed rectangle of the skeleton in the height direction to obtain differential height data, including:

[0204] Obtain the rectangular bottom coordinate data of the circumscribed rectangle of the skeleton;

[0205] Taking the rectangular bottom coordinate data as the starting point, based on the height direction of the circumscribed rectangle of the skeleton, successively select a number of height coordinate data;

[0206] According to the rectangular bottom coordinate data and the height coordinate data, calculate the differential height data of the circumscribed rectangle of the skeleton.

[0207] In the above embodiments, obtain the rectangular bottom coordinate data of the circumscribed rectangle of the skeleton. Taking the rectangular bottom coordinate data as the starting point, based on the height direction of the circumscribed rectangle of the skeleton, successively select a number of height coordinate data, and calculate the differential height data of the circumscribed rectangle of the skeleton.

[0208] The beneficial effect of the above technology is that by obtaining the rectangular bottom coordinate data and selecting height coordinate data in the vertical direction, the acquisition of the differential height data of the circumscribed rectangle of the skeleton is realized, which is convenient for the subsequent steps to judge the distortion situation of the text line mask image.

[0209] Example 9:

[0210] An embodiment of the present invention provides a method for correcting and discriminating document images based on deep learning. The steps are as follows: according to the text line correction and discrimination information corresponding to all text lines in the preprocessed document image, perform correction and discrimination on the preprocessed document image, and perform corresponding image correction operations, including:

[0211] Obtain the text line correction and discrimination information corresponding to all text lines in the preprocessed document image;

[0212] Obtain the text content within the text line and calculate the text line information amount;

[0213] Obtain all the text content in the preprocessed document image, obtain the text information amount, calculate the proportion of the text line information amount in the text information amount, and obtain the information amount weight corresponding to the text line;

[0214] According to the information amount weight and the corresponding text line correction and discrimination information, obtain the line correction parameter of the text line;

[0215] Statistically preprocess the line correction parameters corresponding to all text lines in the document image, obtain the text correction parameters, and when the correction threshold parameter is exceeded, perform correction processing on the preprocessed document image. When the preset correction threshold parameter is not exceeded, there is no need to perform correction processing on the preprocessed document image.

[0216] In the above embodiments, by obtaining the text line information amount of the text content within the text line and the text information amount of all text content in the preprocessed document image, the information amount weight corresponding to the text line is obtained, and based on the information amount weight and the corresponding text line correction discrimination information, the line correction parameter is obtained. According to the line correction parameters corresponding to all text lines, the text correction parameter is obtained.

[0217] In the above embodiments, when the text correction parameter exceeds the preset correction threshold parameter, correction processing is performed on the preprocessed document image. When the preset correction threshold parameter is not exceeded, there is no need to perform correction processing on the preprocessed document image.

[0218] The beneficial effects of the above technology are as follows: By the information amount weight of the text line and the text line correction discrimination information, the acquisition of the line correction parameter of the text line is realized, and by statistically preprocessing the line correction parameters of all text lines in the document image, the text correction parameter is obtained. Based on the preset correction threshold parameter, the correction discrimination of the document image is realized. By comprehensively considering the information amount weight of the text line and the corresponding text line correction discrimination information to obtain the line correction parameter, the problem of correction discrimination only based on the text line correction discrimination information is avoided, and the accuracy of the correction discrimination of the document image is improved.

[0219] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.

Claims

1. A document image correction and discrimination method based on deep learning, characterized in that: include: Acquire document images; Performing four-point detection on the document image to obtain corner point coordinate information of the document image; According to the corner point coordinate information, the background area in the document image is removed to obtain a preprocessed document image; Performing text line segmentation on the preprocessed document image based on a text detection and segmentation algorithm to obtain a text line mask image; Obtaining the maximum bounding rectangle of the text line mask image, cropping the maximum bounding rectangle, and performing skeleton extraction to obtain a skeleton bounding rectangle; Determining the distortion of the text line mask image according to the height of the skeleton circumscribed rectangle, and obtaining text line correction discrimination information; According to the text line correction discrimination information corresponding to all text lines in the pre-processed document image, correction discrimination of the pre-processed document image is performed, and a corresponding image correction operation is executed.

2. The document image correction and discrimination method based on deep learning according to claim 1, characterized in that: The step of performing four-point detection on the document image to obtain the corner point coordinate information of the document image comprises: Construct a corner detection model based on the backbone network layer, sampling layer and feature coordinate calculation layer; Inputting the document image into the corner point detection model to obtain the corner point coordinate information; optionally, including: Performing a dimensionality-enhancing convolution operation on the document image through a convolution expansion module in a backbone network layer to obtain a high-dimensional convolution feature image, processing the high-dimensional convolution feature image through a first batch normalization module, performing nonlinear feature extraction through a first activation function to obtain a high-dimensional feature image; Performing a convolution operation with a fixed number of channels on the high-dimensional feature image through a deep convolution module in the backbone network layer to obtain a deep convolution feature image, and performing nonlinear feature extraction through a second batch normalization module and the first activation function to obtain a deep feature image; Performing a dimensionality reduction convolution operation on the deep feature image through a convolution projection module in the backbone network layer to obtain a dimensionality reduction feature image, and outputting a standard dimensionality reduction feature image after processing through a third batch normalization module; Performing deconvolution processing on the standard dimension reduction feature image through a sampling layer, setting the number of channels, and outputting a corresponding document feature heat map; Performing an activation operation on the document feature heat map through a second activation function in a feature coordinate calculation layer to obtain a standard feature heat map; Among them, softmax is the second activation function, p ij is the element value in the i-th row and j-th column of the document feature heat map, w is the number of rows of element values ​​in the document feature heat map, h is the number of columns of element values ​​in the document feature heat map, and p sv is the element value in the sth row and vth column of the document feature heat map; Constructing an X-axis constant coordinate matrix and a Y-axis constant coordinate matrix respectively according to the document feature heat map; Calculating the corner point coordinate information according to the standard feature heat map, the X-axis constant coordinate matrix and the Y-axis constant coordinate matrix; Among them, (x a ,y a ) is the corner point coordinate information, x ij is the element value p in the X-axis constant coordinate matrix ij The corresponding coordinate value, y ij is the element value p in the Y-axis constant coordinate matrix ij The corresponding coordinate values.

3. The document image correction and discrimination method based on deep learning according to claim 2, characterized in that: The step of: building a corner point detection model based on the backbone network layer, the sampling layer and the feature coordinate calculation layer; and also includes: The corner point detection model is trained by a training set, and when the comprehensive loss function meets the preset threshold requirement, the corner point detection model is used for four-point detection of the document image; optionally, including: Inputting the sample document images in the training set into the angle detection model for processing to obtain sample corner point coordinate information; Calculating a coordinate point loss function according to the standard corner point coordinate information corresponding to the sample document image and the sample corner point coordinate information; L1=‖C a -C sa ‖2 Among them, L1 is the coordinate point loss function, C a is the sample corner point coordinate information, C sa is the standard corner point coordinate information, ‖‖2 is the Euclidean normal form; Obtaining a sample feature heat map corresponding to the sample document image through the angle detection model, and calculating a heat map loss function according to the sample standard feature heat map corresponding to the sample document image and the sample feature heat map; Among them, L2 is the heat map loss function, D KL () is the KL divergence calculation formula, P a is the sample feature heat map, P Sa is the sample standard feature heat map; Constructing a comprehensive loss function according to the coordinate point loss function and the heat map loss function; L=L1+μL2 Among them, L is the comprehensive loss function, μ is the hyperparameter; When the comprehensive loss function meets a preset threshold requirement, the corner point detection model is used for four-point detection of the document image.

4. The document image correction and discrimination method based on deep learning according to claim 1, characterized in that: The step of removing the background area in the document image according to the corner point coordinate information to obtain a pre-processed document image comprises: Constructing a circumscribed rectangle of a document area in the document image according to the corner point coordinate information; Construct a perspective transformation matrix based on the coordinate parameters of the circumscribed rectangle of the document area; Based on the perspective transformation matrix, the document area is mapped to a new canvas, a pre-processed document image is obtained, and a background area in the document image is removed.

5. The document image correction and discrimination method based on deep learning according to claim 1, characterized in that: The step: performing text line segmentation on the pre-processed document image based on a text detection and segmentation algorithm to obtain a text line mask image; include: Constructing a text detection and segmentation model according to the feature extraction network, the sampling network and the differentiable binarization network, and training the text detection and segmentation model through the sample set; Performing a text line segmentation operation on the preprocessed document image through the trained text detection and segmentation model to obtain a text line mask image; optionally, including: Extracting features from the preprocessed document image through a feature extraction network to obtain feature convolution graphs of different scales; Through the sampling network, the feature convolution maps of different scales are upsampled accordingly to obtain feature sub-maps of the same size for feature fusion and generate feature fusion maps; Detect the probability value of each pixel point in the feature fusion map belonging to the text box area to generate a feature probability map; calculate the binarization threshold of each pixel point in the feature fusion map to obtain a feature threshold map; Processing is performed according to the feature threshold map and the feature probability map through a differentiable binarization network to obtain an approximate binary image; Among them, P i ′ j is the pixel in the i-th row and j-th column of the approximate binary image, l is the magnification factor, u ij is the pixel at the i-th row and j-th column in the preprocessed document image, S ij is the pixel in the i-th row and j-th column in the feature threshold map; Acquire a text box region in the preprocessed document image according to the approximate binary image; Perform expansion and contraction operations in the text box area, perform compensation optimization processing, and obtain the boundary of the text box area; G=F-t·M 2 ·P ′ Among them, G is the text box area, F is the standard binary image processed by threshold according to the feature probability map, t is the empirical coefficient, M is the connected area identified in the text box area, and P ′ is an approximate binary image; According to the boundary of the text box area identified in the pre-processed text image, a text segmentation operation is performed to obtain a text line mask image.

6. The document image correction and discrimination method based on deep learning according to claim 5, characterized in that: The step of constructing a text detection and segmentation model according to a feature extraction network, a sampling network and a differentiable binarization network, and training the text detection and segmentation model through a sample set includes: Obtain preprocessed sample images in a sample set; The preprocessed sample image is processed by the text detection and segmentation model to obtain a sample feature probability map, a sample feature threshold map and a sample approximate binary image; Calculating a model loss function according to a standard feature probability map, a standard feature threshold map, and a standard approximate binary image corresponding to the preprocessed sample image in the sample set; L m =L g +θ1L b +θ2L p Among them, L m is the model loss function, L g is the probability loss function of the standard feature probability map and the sample feature probability map, L b is the threshold loss function of the standard feature threshold map and the sample feature threshold map, L p is the binary loss function of the standard approximate binary image and the sample approximate binary image, θ1 is the first adjustment parameter, and θ2 is the second adjustment parameter; Among them, g Si is the standard feature probability map corresponding to the i-th preprocessed sample image in the sample set K, g i The sample feature probability map obtained by the text detection and segmentation model based on the i-th preprocessed sample image; Among them, b i The text detection and segmentation model obtains the sample feature threshold map based on the i-th preprocessed sample image, b Si is the standard feature threshold map corresponding to the i-th preprocessed sample image in the sample set K; Among them, p i The text detection segmentation model obtains the sample approximate binary image based on the i-th preprocessed sample image, p Si is the standard approximate binary image corresponding to the i-th preprocessed sample image in the sample set K; The text detection and segmentation model is trained according to the model loss function, and the control parameters of the feature extraction network, the sampling network and the differentiable binarization network are optimized. When the model loss function meets the loss function threshold, the text detection and segmentation model is used to pre-process the text line segmentation of the document image.

7. The document image correction and discrimination method based on deep learning according to claim 1, characterized in that: The step of determining the distortion of the text line mask image according to the height of the skeleton circumscribed rectangle and obtaining text line correction discrimination information comprises: Obtaining the rectangular height data of the circumscribed rectangle of the skeleton; Performing segmented detection on the skeleton circumscribed rectangle based on the height direction to obtain differential height data; When the height variance data between the rectangular height data and the differential height data is less than a first preset threshold error, marking the text line mask image as not requiring correction processing; When the height variance data between the rectangle height data and the differential height data is greater than a second preset threshold error, the text line mask image is marked as to be corrected.

8. The document image correction and discrimination method based on deep learning according to claim 7 is characterized in that: The step of performing segmented detection on the skeleton circumscribed rectangle based on the height direction to obtain differential height data comprises: Obtaining the bottom coordinate data of the rectangle circumscribing the skeleton; Taking the bottom coordinate data of the rectangle as the starting point, based on the height direction of the circumscribed rectangle of the skeleton, select a number of height coordinate data in sequence; The differential height data of the skeleton circumscribed rectangle is calculated according to the rectangle bottom coordinate data and the height coordinate data.

9. The document image correction and discrimination method based on deep learning according to claim 1, characterized in that: The step of: performing correction discrimination of the pre-processed document image according to the text line correction discrimination information corresponding to all text lines in the pre-processed document image, and executing the corresponding image correction operation; comprises: Acquire text line correction discrimination information corresponding to all text lines in the preprocessed document image; Obtaining the text content in the text line and calculating the information volume of the text line; Acquire all text contents in the preprocessed document image, acquire text information, calculate the proportion of text line information in the text information, and acquire the information weight corresponding to the text line; Acquire a line correction parameter of the text line according to the information weight and the corresponding text line correction discrimination information; The line correction parameters corresponding to all text lines in the preprocessed document image are counted to obtain text correction parameters. When the preset correction threshold parameters are exceeded, the preprocessed document image is corrected. When the preset correction threshold parameters are not exceeded, there is no need to correct the preprocessed document image.