A method and system for automatically correcting printed documents based on deep learning
Through improved U-shaped convolutional network, dual-stream Transformer network and improved Hough transformation based on deep learning, the inefficiency of traditional bias correction methods in complex scenarios is solved, and high-quality automatic bias correction and printing effect of printed files are achieved.
Patent Information
- Application Number
- CN202510300515.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-03-14
AI Technical Summary
Traditional printing file deviation correction methods are inefficient and difficult to deal with complex scenarios, such as low contrast, blurred edges or noise interference, resulting in deviation correction failure or accumulation of errors.
Using a deep learning-based approach, page tilt and deformation of printed files are automatically identified and corrected through improved U-shaped convolutional networks, dual-stream Transformer networks and improved Hough transformations.
It realizes high-quality automatic deviation correction of printed files, improves printing effect, and significantly improves the accuracy of content area and edge detection in complex scenarios.
Smart Images

Figure CN119832555B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of printer control technology, and in particular to a method and system for automatically correcting a printed document based on deep learning. Background Art
[0002] In the field of modern office and publishing, the accuracy and aesthetics of printed documents directly affect work efficiency and user experience. However, in actual operation, the printed documents provided by customers often have common problems such as page tilt, which usually require manual document cropping and correction, which not only takes a lot of time, but also easily causes paper waste, thereby increasing printing costs.
[0003] Traditional manual deflection correction methods rely on the operator's experience and technical level, are inefficient, and are difficult to ensure consistency. In addition, with the growing demand for printing, manual processing can no longer meet the requirements of large-scale, high-precision production. In recent years, although some automated tools have attempted to make preliminary adjustments to print files through simple geometric transformations or edge detection algorithms, their effects are often unsatisfactory due to their lack of adaptability to complex scenarios. For example, when the contrast between the content area of the document and the background is low, the edges are blurred, or there is noise interference, traditional methods have difficulty accurately identifying the boundaries of the document, resulting in correction failure or error accumulation. Summary of the invention
[0004] In view of this, the present invention provides a method and system for automatic correction of printed documents based on deep learning, aiming to solve the problem of poor printing effect of traditional printers caused by the tilt of printed document pages.
[0005] A method for automatically correcting the deviation of printed documents based on deep learning, comprising:
[0006] S1: Obtain a PDF file to be printed and preprocess it in the form of an image to obtain a preprocessed page image;
[0007] S2: Based on the preprocessed page image, an improved U-shaped convolutional network is designed to generate a content area probability image; the probability of each pixel in the content area probability image being an edge is calculated to obtain an edge probability image;
[0008] S3: Based on the edge probability image, an improved Hough transform is designed to perform line segment clustering to obtain an accumulator matrix; based on the accumulator matrix, a set of candidate line segments of the file boundary, a main direction angle, a file boundary line segment and a page tilt angle are calculated; based on the page tilt angle, a rotation transformation is performed on the preprocessed page image to obtain a page image after tilt correction;
[0009] S4: Design a two-stream Transformer network, extract fusion features according to the page tilt angle and edge probability image, and judge whether the page image after tilt correction has page deformation according to the fusion features; if there is page deformation, calculate the grid offset, and perform deformation correction on the page image after tilt correction to obtain the page image after deformation correction; if there is no page deformation, directly use the page image after tilt correction as the page image after deformation correction;
[0010] S5: starting a printing task through a printer driver according to the page image after deformation correction, and recording a printing result;
[0011] Furthermore, the step S1 further includes:
[0012] S11: using a PDF parsing tool, reading each page content of the PDF file to be printed, and converting it into a pixel matrix form to obtain an original page image;
[0013] S12: using a weighted average method, converting the original page image into a standardized grayscale format to obtain a grayscale page image;
[0014] S13: using a median filter algorithm to perform denoising on the grayscale page image to obtain a denoised page image;
[0015] S14: For the denoised page image, a histogram equalization method is used to adjust the brightness and contrast, and then a size normalization process is performed to obtain a preprocessed page image.
[0016] Furthermore, the step S2 further includes:
[0017] The improved U-shaped convolutional network includes: introducing a spatial attention mechanism and weighted optimization, dynamically adjusting the decoding layer features through the spatial attention weight matrix to obtain the attention-weighted decoding layer features; channel-joining the attention-weighted decoding layer features with the encoding layer features to achieve feature fusion;
[0018] The edge probability image is obtained by weighted calculation using adaptive weight parameters, specifically, by calculating the edge probability image using the adaptive weight parameters in combination with a sigmoid function, the variance of the content region probability image, and a Sobel gradient field.
[0019] Furthermore, the step S2 further includes:
[0020] S21: Design an improved U-shaped convolutional network to generate a content area probability image based on the preprocessed page image. The calculation method is:
[0021] ;
[0022] in, is the content region probability image, is the sigmoid function, is the encoding layer feature, is the feature concatenation operation, is the spatial attention weight matrix, is element-wise multiplication, is the decoding layer feature, are the convolution and pooling operations in the encoder, is the preprocessed page image, are the convolution and pooling operations in the decoder, is the fully connected layer, is global average pooling;
[0023] S22: Calculate the probability of each pixel in the content area probability image being an edge to obtain an edge probability image. The calculation method is:
[0024] ;
[0025] in, is the edge probability image, is the adaptive weight parameter, is the Sobel gradient field, is the content area probability image at coordinates The two-dimensional derivative at , are the horizontal and vertical coordinates of the content area probability image, respectively. Respectively represent the gradient components of the preprocessed page image in the horizontal and vertical directions, is the Sobel operator operation in the horizontal direction, is the Sobel operator operation in the vertical direction, is the adaptive weight hyperparameter, is the variance, is the gradient strength of the Sobel gradient field.
[0026] It should be noted that the jump connection of the traditional U-shaped convolutional network directly splices the encoder and decoder layer features, which easily introduces redundant information or noise. The improved U-shaped convolutional network introduces the spatial attention weight matrix By weighting the decoding layer features element by element, the ability to focus on key areas is enhanced and irrelevant background information is suppressed; in step S21, , which realizes the effective integration of multi-scale features and reduces redundant information compared with the simple splicing in the traditional U-shaped convolutional network. In the task of correcting the printed document, the improved U-shaped convolutional network can more accurately identify the content area and edge area, especially in complex scenes such as low contrast, blurred edges or noise interference. In addition, the improved U-shaped convolutional network outputs a content area probability image rather than a simple binary mask. This probabilistic output provides a more refined segmentation result for subsequent processing, which facilitates higher quality edge detection and content area positioning.
[0027] In step S22, the present invention uses the Sobel operator to extract the significant edge information of the image, which has strong local sensitivity and the two-dimensional spatial derivative of the content area probability image Combined with the deep learning model's understanding of the content area, it can capture more detailed semantic boundary information; the adaptive weight parameter α dynamically adjusts the contribution of variance and gradient strength through the sigmoid function to ensure that better edge detection results can be obtained in different scenarios (such as high contrast areas or low contrast areas); compared with using the Sobel operator alone, the traditional Sobel operator only relies on the local gradient information of the image and is easily affected by noise interference or low contrast areas, while the improved algorithm of the present invention can more accurately detect the physical boundaries and content area boundaries of the file, especially in complex scenarios such as low contrast, blurred edges or noise interference; in addition, the improved algorithm outputs an edge probability image rather than a simple binary mask, which provides a more refined edge detection result for subsequent processing and facilitates the realization of higher quality boundary line segment extraction and tilt correction.
[0028] Furthermore, the step S3 further includes:
[0029] When the accumulator matrix performs accumulation calculation, the pixel value of the content area probability image at the coordinate is used as the accumulation weight; wherein the polar angle parameter is selected The range is .
[0030] Furthermore, the step S3 further includes:
[0031] S31: According to the edge probability image, an improved Hough transform is designed to perform line segment clustering and obtain the accumulator matrix. The calculation method is:
[0032] ;
[0033] in, is the accumulator matrix, is the content area probability image at coordinates The pixel value at is the impulse function, is the polar diameter parameter;
[0034] S32: performing local peak detection on the accumulator matrix, and using a clustering algorithm on each detected local peak to divide adjacent local peaks with the same direction into the same group, thereby obtaining a set of file boundary candidate line segments;
[0035] S33: in the set of file boundary candidate line segments, according to the polar angle parameter of each file boundary candidate line segment, a weighted average method is used to calculate the main direction angle; the main direction angle is a weighted average of the polar angle parameters of each file boundary candidate line segment, and the weight is the length of the corresponding file boundary candidate line segment;
[0036] S34: in the set of file boundary candidate line segments, retain the straight line segment with the longest length and the smallest deviation from the main direction angle as the file boundary line segment, and output the polar radius parameter and polar angle parameter corresponding to the file boundary line segment;
[0037] S35: Calculate the page tilt angle according to the polar angle parameter of the file boundary line segment, and the calculation method is:
[0038] ;
[0039] in, is the page tilt angle, is the reference direction angle, ; is the polar angle parameter of the file boundary segment;
[0040] S36: Rotate the pre-processed page image according to the page tilt angle to obtain a page image after tilt correction .
[0041] It should be noted that the scanned PDF files or printed files may have blurred edges or low contrast due to factors such as scanning quality and lighting conditions. Traditional Hough transform usually assigns the same weight to all edge points (i.e., the cumulative value is 1), which is easily affected by noise interference or low-contrast areas. The improved Hough transform converts the content area probability image The pixel value is directly used as the accumulated weight, so that the voting value of each edge point is no longer fixed to 1, which makes the contribution of significant edge points greater, thereby improving the signal-to-noise ratio of the accumulator matrix H, which is more suitable for complex scenes such as low contrast, blurred edges or noise interference in the print document correction task; the range of the polar angle parameter θ is limited to , in order to adapt to the common tilt angle range in the print document correction task, it reduces the amount of calculation while avoiding false detection in irrelevant directions, thereby improving detection efficiency and accuracy.
[0042] Furthermore, the step S4 further includes:
[0043] The design of the two-stream Transformer network includes: constructing parallel geometric streams and edge streams, and extracting geometric features and edge features respectively; introducing inter-stream attention weights to fuse information of geometric features and edge features, and updating the geometric features by layer normalization to obtain updated geometric features; designing a two-stream gate to fuse edge features and updated geometric features to obtain fused features.
[0044] Furthermore, the step S4 further includes:
[0045] S41: Design a two-stream Transformer network to extract edge features and geometric features based on the page tilt angle and edge probability image, and then extract fusion features. The calculation method is:
[0046] ;
[0047] in, is the geometric feature, For the Transformer network, is the edge feature, is a convolutional neural network, is the attention weight between streams, is the Softmax function, is the transpose of the edge feature, is the feature dimension, is the updated geometric feature, is layer normalization, For double flow door, is the double-stream gate weight, is the dual-stream gate bias vector, For fusion features;
[0048] S42: judging whether the page image after the tilt correction has page deformation according to the fusion feature, the calculation method is:
[0049] ;
[0050] in, To determine whether there is a page deformation result, When the page is deformed, It is considered that there is no page deformation;
[0051] S43: If there is page deformation, the grid offset is calculated, including the horizontal grid offset and the vertical grid offset, and the calculation method is:
[0052] ;
[0053] in, is the deformation probability image, is the horizontal grid offset, is the hyperbolic tangent function, is the vertical grid offset;
[0054] According to the horizontal grid offset and the vertical grid offset, the page image after tilt correction is deformed by bilinear interpolation to obtain the page image after deformation correction. ;
[0055] S44: If there is no page deformation, directly convert the page image after tilt correction as a distortion-corrected page image;
[0056] It should be noted that the present invention extracts geometric features and edge features respectively through the Transformer network, and dynamically adjusts the contribution of the two by using the inter-stream attention weights; the dual-stream gate K realizes the adaptive fusion of geometric features and edge features by weighting the fused features element by element, further enhancing the model's adaptability to different scenes, so that it can obtain the best feature fusion results under different conditions, such as high-contrast areas or low-contrast areas.
[0057] Furthermore, the step S5 further includes:
[0058] S51: configuring printing task parameters, including paper size, printing resolution, color mode, printing direction, and number of copies;
[0059] S52: starting a printing task for the page image after deformation correction through a printer driver;
[0060] S53: Monitor the execution status of the printing task in real time and record the printing result.
[0061] The present invention also discloses a printing document automatic deviation correction system based on deep learning, comprising:
[0062] The module for acquiring and preprocessing the files to be printed: acquiring the PDF files to be printed and preprocessing them in the form of images to obtain the preprocessed page images;
[0063] Edge probability image calculation module: Based on the preprocessed page image, an improved U-shaped convolutional network is designed to generate a content area probability image and calculate the probability of each pixel in the content area probability image being an edge to obtain an edge probability image;
[0064] Skew correction module: According to the edge probability image, an improved Hough transform is designed for line segment clustering to obtain an accumulator matrix; based on the accumulator matrix, a set of candidate line segments for the file boundary, the main direction angle, and the angle between the file boundary line segment and the page skew are calculated; according to the page skew angle, the preprocessed page image is rotated to obtain the skew-corrected page image;
[0065] Distortion correction module: A dual-stream Transformer network is designed. According to the page skew angle and the edge probability image, fused features are extracted, and it is determined whether there is page distortion in the skew-corrected page image based on the fused features; if there is page distortion, the grid offset is calculated, and the skew-corrected page image is corrected for distortion to obtain the distortion-corrected page image; if there is no page distortion, the skew-corrected page image is directly used as the distortion-corrected page image;
[0066] Printing module: According to the distortion-corrected page image, a printing task is started through the printer driver, and the printing result is recorded.
[0067] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0068] (1) Aiming at the problem of poor printing effect of traditional printers caused by page skew of printed files, the present invention designs an improved U-shaped convolutional network, a dual-stream Transformer network, and an improved Hough transform method, and performs skew correction and distortion correction on the file to be printed. Finally, printing is performed according to the distortion-corrected page image, realizing automatic high-quality rectification of printed files and improving the printing effect.
[0069] (2) Aiming at the problem of insufficient accuracy that traditional U-Net is prone to in scenarios such as low contrast, blurred edges, or noise interference, the present invention improves the U-shaped convolutional network, introduces spatial attention weights and weighted optimization, dynamically adjusts the contribution of the features in the decoding layer, significantly improves the generation accuracy of the content area probability image and the edge probability image, and further improves the calculation accuracy of the page skew angle.
[0070] (3) Aiming at the problem that traditional Hough transform assigns the same weight to all edge points and is easily affected by noise interference or low-contrast regions, the present invention uses the pixel value of the content area probability image as the accumulation weight and restricts the range of the polar angle parameter to adapt to the common skew angle range in the printed file rectification task, improving the accuracy and robustness of page content boundary detection.
[0071] (4) The present invention designs a dual-stream feature extraction and fusion mechanism, combines the inter-stream attention weight and the dual-stream gating mechanism, dynamically adjusts the contribution ratio of geometric features and edge features, captures the overall deformation trend and local deformation details of the page, and improves the accuracy and robustness of page deformation detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 A schematic diagram of a flow chart of a method for automatic deviation correction of printed documents based on deep learning provided by the present invention;
[0073] Figure 2 This is a schematic diagram of the algorithm flow for generating a content region probability image provided by the present invention. DETAILED DESCRIPTION
[0074] The present invention is further described below in conjunction with the accompanying drawings, but the present invention is not limited in any way. Any changes or substitutions made based on the teachings of the present invention belong to the protection scope of the present invention.
[0075] Embodiment 1: A method for automatically correcting the deviation of printed documents based on deep learning, such as Figure 1 and Figure 2 As shown, the following steps are included:
[0076] S1: Obtain a PDF file to be printed and preprocess it in the form of an image to obtain a preprocessed page image;
[0077] S11: using a PDF parsing tool, reading each page content of the PDF file to be printed, and converting it into a pixel matrix form to obtain an original page image;
[0078] S12: using a weighted average method, converting the original page image into a standardized grayscale format to obtain a grayscale page image;
[0079] S13: using a median filter algorithm to perform denoising on the grayscale page image to obtain a denoised page image;
[0080] S14: For the denoised page image, a histogram equalization method is used to adjust the brightness and contrast, and then a size normalization process is performed to obtain a preprocessed page image.
[0081] S2: Based on the preprocessed page image, an improved U-shaped convolutional network is designed to generate a content area probability image; the probability of each pixel in the content area probability image being an edge is calculated to obtain an edge probability image;
[0082] S21: Design an improved U-shaped convolutional network to generate a content area probability image based on the preprocessed page image. The calculation method is:
[0083] ;
[0084] in, is the content region probability image, is the sigmoid function, is the encoding layer feature, is the feature concatenation operation, is the spatial attention weight matrix, is element-wise multiplication, is the decoding layer feature, are the convolution and pooling operations in the encoder, is the preprocessed page image, are the convolution and pooling operations in the decoder, is the fully connected layer, is global average pooling;
[0085] S22: Calculate the probability of each pixel in the content area probability image being an edge to obtain an edge probability image. The calculation method is:
[0086] ;
[0087] in, is the edge probability image, is the adaptive weight parameter, is the Sobel gradient field, is the content area probability image at coordinates The two-dimensional derivative at , are the horizontal and vertical coordinates of the content area probability image, respectively. Respectively represent the gradient components of the preprocessed page image in the horizontal and vertical directions, is the Sobel operator operation in the horizontal direction, is the Sobel operator operation in the vertical direction, is the adaptive weight hyperparameter, is the variance, is the gradient strength of the Sobel gradient field.
[0088] It should be noted that the jump connection of the traditional U-shaped convolutional network directly splices the encoder and decoder layer features, which easily introduces redundant information or noise. The improved U-shaped convolutional network introduces the spatial attention weight matrix By weighting the decoding layer features element by element, the ability to focus on key areas is enhanced and irrelevant background information is suppressed; in step S21, , which realizes the effective integration of multi-scale features and reduces redundant information compared with the simple splicing in the traditional U-shaped convolutional network. In the task of correcting the printed document, the improved U-shaped convolutional network can more accurately identify the content area and edge area, especially in complex scenes such as low contrast, blurred edges or noise interference. In addition, the improved U-shaped convolutional network outputs a content area probability image rather than a simple binary mask. This probabilistic output provides a more refined segmentation result for subsequent processing, which facilitates higher quality edge detection and content area positioning.
[0089] In step S22, the present invention uses the Sobel operator to extract the significant edge information of the image, which has strong local sensitivity and the two-dimensional spatial derivative of the content area probability image Combined with the deep learning model's understanding of the content area, it can capture more detailed semantic boundary information; the adaptive weight parameter α dynamically adjusts the contribution of variance and gradient strength through the sigmoid function to ensure that better edge detection results can be obtained in different scenarios (such as high contrast areas or low contrast areas); compared with using the Sobel operator alone, the traditional Sobel operator only relies on the local gradient information of the image and is easily affected by noise interference or low contrast areas, while the improved algorithm of the present invention can more accurately detect the physical boundaries and content area boundaries of the file, especially in complex scenarios such as low contrast, blurred edges or noise interference; in addition, the improved algorithm outputs an edge probability image rather than a simple binary mask, which provides a more refined edge detection result for subsequent processing and facilitates the realization of higher quality boundary line segment extraction and tilt correction.
[0090] For example, the preprocessed page image contains a rectangular content area (such as a text block) located in the center of the matrix, and the background area is black. The generated content area probability image P is as follows:
[0091] ;
[0092] The values in the background area (outskirts of the matrix) are close to 0, indicating that these areas are highly likely to be background;
[0093] The values of the content area (the center part of the matrix) are close to 1, indicating that these areas are highly likely to be content areas;
[0094] The values of boundary areas (such as rows 2 and 7) are between 0 and 1, indicating that the model is uncertain about the classification of these areas;
[0095] Generated edge probability image As shown below:
[0096] ;
[0097] The values of the background regions (outskirts of the matrix) are close to 0, indicating that these regions are highly unlikely to be edges;
[0098] The values of the borders of the content area (such as rows 2, 3, 6, and 7) are close to 1, indicating that these areas are likely to be edges;
[0099] The values inside the content area (such as rows 4 and 5) are close to 0, indicating that these areas are highly unlikely to be edges;
[0100] The values of fuzzy boundary regions (such as partial pixels in rows 2 and 7) are between 0 and 1, indicating that the model is uncertain about the classification of these regions.
[0101] In particular, in order to solve the problem that the text or table boundaries are blurred due to poor scanning quality or paper bending in the preprocessed page image, the present invention also provides a new method for calculating the spatial attention weight matrix, which is used to enhance the accuracy of generating the content area probability image in step S21. The new method for calculating the spatial attention weight matrix is:
[0102] ;
[0103] in, is the noise suppression parameter, is the noise intensity.
[0104] S3: Based on the edge probability image, an improved Hough transform is designed to perform line segment clustering to obtain an accumulator matrix; based on the accumulator matrix, a set of candidate line segments of the file boundary, a main direction angle, a file boundary line segment and a page tilt angle are calculated; based on the page tilt angle, a rotation transformation is performed on the preprocessed page image to obtain a page image after tilt correction;
[0105] S31: According to the edge probability image, an improved Hough transform is designed to perform line segment clustering and obtain the accumulator matrix. The calculation method is:
[0106] ;
[0107] in, is the accumulator matrix, is the content area probability image at coordinates The pixel value at is the impulse function, is the polar diameter parameter, is the polar angle parameter, ;
[0108] S32: performing local peak detection on the accumulator matrix, and using a clustering algorithm on each detected local peak to divide adjacent local peaks with the same direction into the same group, thereby obtaining a set of file boundary candidate line segments;
[0109] S33: in the set of file boundary candidate line segments, according to the polar angle parameter of each file boundary candidate line segment, a weighted average method is used to calculate the main direction angle; the main direction angle is the weighted average of the polar angle parameters of each file boundary candidate line segment, and the weight is the length of the corresponding file boundary candidate line segment;
[0110] S34: in the set of file boundary candidate line segments, retain the straight line segment with the longest length and the smallest deviation from the main direction angle as the file boundary line segment, and output the polar radius parameter and polar angle parameter corresponding to the file boundary line segment;
[0111] S35: Calculate the page tilt angle according to the polar angle parameter of the file boundary line segment, and the calculation method is:
[0112] ;
[0113] in, is the page tilt angle, is the reference direction angle, ; is the polar angle parameter of the file boundary segment;
[0114] S36: Rotate the pre-processed page image according to the page tilt angle to obtain a page image after tilt correction .
[0115] It should be noted that the scanned PDF files or printed files may have blurred edges or low contrast due to factors such as scanning quality and lighting conditions. Traditional Hough transform usually assigns the same weight to all edge points (i.e., the cumulative value is 1), which is easily affected by noise interference or low-contrast areas. The improved Hough transform converts the content area probability image The pixel value is directly used as the accumulated weight, so that the voting value of each edge point is no longer fixed to 1, which makes the contribution of significant edge points greater, thereby improving the signal-to-noise ratio of the accumulator matrix H, which is more suitable for complex scenes such as low contrast, blurred edges or noise interference in the print document correction task; the range of the polar angle parameter θ is limited to , in order to adapt to the common tilt angle range in the print document correction task, it reduces the amount of calculation while avoiding false detection in irrelevant directions, thereby improving detection efficiency and accuracy.
[0116] S4: Design a two-stream Transformer network, extract fusion features according to the page tilt angle and edge probability image, and judge whether the page image after tilt correction has page deformation according to the fusion features; if there is page deformation, calculate the grid offset, and perform deformation correction on the page image after tilt correction to obtain the page image after deformation correction; if there is no page deformation, use the page image after tilt correction as the page image after deformation correction;
[0117] S41: Design a two-stream Transformer network to extract edge features and geometric features based on the page tilt angle and edge probability image, and then extract fusion features. The calculation method is:
[0118] ;
[0119] in, is the geometric feature, For the Transformer network, is the edge feature, is a convolutional neural network, is the attention weight between streams, is the Softmax function, is the transpose of the edge feature, is the feature dimension, is the updated geometric feature, is layer normalization, For double flow door, is the double-stream gate weight, is the dual-stream gate bias vector, For fusion features;
[0120] S42: judging whether the page image after the tilt correction has page deformation according to the fusion feature, the calculation method is:
[0121] ;
[0122] in, To determine whether there is a page deformation result, When the page is deformed, It is considered that there is no page deformation;
[0123] S43: If there is page deformation, the grid offset is calculated, including the horizontal grid offset and the vertical grid offset, and the calculation method is:
[0124] ;
[0125] in, is the deformation probability image, is the horizontal grid offset, is the hyperbolic tangent function, is the vertical grid offset;
[0126] According to the horizontal grid offset and the vertical grid offset, the page image after tilt correction is deformed by bilinear interpolation to obtain the page image after deformation correction. ;
[0127] S44: If there is no page deformation, directly convert the page image after tilt correction As the page image after distortion correction.
[0128] It should be noted that the present invention extracts geometric features and edge features respectively through the Transformer network, and dynamically adjusts the contribution of the two by using the inter-stream attention weights; the dual-stream gate K realizes the adaptive fusion of geometric features and edge features by weighting the fused features element by element, further enhancing the model's adaptability to different scenes, so that it can obtain the best feature fusion results under different conditions, such as high-contrast areas or low-contrast areas.
[0129] S5: starting a printing task through a printer driver according to the page image after deformation correction, and recording a printing result;
[0130] S51: configuring printing task parameters, including paper size, printing resolution, color mode, printing direction, and number of copies;
[0131] S52: starting a printing task for the page image after deformation correction through a printer driver;
[0132] S53: Monitor the execution status of the printing task in real time and record the printing result.
[0133] Embodiment 2: The present invention also discloses a printing document automatic deviation correction system based on deep learning, comprising:
[0134] The module for acquiring and preprocessing the files to be printed: acquiring the PDF files to be printed and preprocessing them in the form of images to obtain the preprocessed page images;
[0135] Edge probability image calculation module: Based on the preprocessed page image, an improved U-shaped convolutional network is designed to generate a content area probability image and calculate the probability of each pixel in the content area probability image being an edge to obtain an edge probability image;
[0136] Tilt correction module: Based on the edge probability image, an improved Hough transform is designed to perform line segment clustering to obtain an accumulator matrix; based on the accumulator matrix, the set of candidate line segments at the file boundary, the main direction angle, the file boundary line segments and the page tilt angle are calculated; based on the page tilt angle, the preprocessed page image is rotated to obtain a page image after tilt correction;
[0137] Deformation correction module: A two-stream Transformer network is designed to extract fusion features based on the page tilt angle and edge probability image, and determine whether the page image after tilt correction has page deformation based on the fusion features; if there is page deformation, the grid offset is calculated, and the page image after tilt correction is deformed to obtain the page image after deformation correction; if there is no page deformation, the page image after tilt correction is directly used as the page image after deformation correction;
[0138] Printing module: Based on the page image after deformation correction, the printing task is started through the printer driver and the printing result is recorded.
[0139] The present invention is particularly suitable for printing documents of the image scan type. Such documents are usually prone to problems such as low contrast, blurred edges and background noise due to the limitations of the scanning equipment or the influence of paper quality. As a result, traditional methods have insufficient accuracy in content area detection and boundary extraction, thereby affecting the subsequent tilt correction and deformation correction effects. The present invention can effectively cope with these challenges and significantly improve the accuracy of content area and edge detection.
[0140] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. And the terms "including", "comprising" or any other variants thereof in this article are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "including a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0141] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in multiple embodiments of the present invention.
[0142] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for automatic deviation correction of printed documents based on deep learning, characterized in that: The following steps are involved: S1: Obtain a PDF file to be printed and preprocess it in the form of an image to obtain a preprocessed page image; S2: Based on the preprocessed page image, an improved U-shaped convolutional network is designed to generate a content area probability image; The probability of each pixel in the content area probability image being an edge is calculated to obtain an edge probability image; wherein the improved U-shaped convolutional network includes: introducing a spatial attention mechanism and weighted optimization, dynamically adjusting the decoding layer features through the spatial attention weight matrix, and obtaining attention-weighted decoding layer features; channel-joining the attention-weighted decoding layer features with the encoding layer features to achieve feature fusion; The edge probability image is obtained by weighted calculation using an adaptive weight parameter, specifically: calculated by combining the adaptive weight parameter with a sigmoid function, the variance of the content area probability image, and the Sobel gradient field; S3: Based on the edge probability image, an improved Hough transform is designed to perform line segment clustering to obtain an accumulator matrix; based on the accumulator matrix, a set of candidate line segments of the file boundary, a main direction angle, a file boundary line segment and a page tilt angle are calculated; based on the page tilt angle, a rotation transformation is performed on the preprocessed page image to obtain a page image after tilt correction; S4: Design a two-stream Transformer network, extract fusion features according to the page tilt angle and edge probability image, and judge whether the page image after tilt correction has page deformation according to the fusion features; if there is page deformation, calculate the grid offset, and perform deformation correction on the page image after tilt correction to obtain the page image after deformation correction; if there is no page deformation, use the page image after tilt correction as the page image after deformation correction; S5: starting a printing task through a printer driver according to the page image after deformation correction, and recording a printing result.
2. The method for automatic deviation correction of printed documents based on deep learning according to claim 1, characterized in that: The step S1 comprises: S11: using a PDF parsing tool, reading each page content of the PDF file to be printed, and converting it into a pixel matrix form to obtain an original page image; S12: using a weighted average method, converting the original page image into a standardized grayscale format to obtain a grayscale page image; S13: using a median filter algorithm to perform denoising on the grayscale page image to obtain a denoised page image; S14: For the denoised page image, a histogram equalization method is used to adjust the brightness and contrast, and then a size normalization process is performed to obtain a preprocessed page image.
3. The method for automatic correction of printed documents based on deep learning according to claim 1, characterized in that: The step S2 comprises: S21: Design an improved U-shaped convolutional network to generate a content area probability image based on the preprocessed page image. The calculation method is: ; in, is the content region probability image, is the sigmoid function, is the encoding layer feature, is the feature concatenation operation, is the spatial attention weight matrix, is element-wise multiplication, is the decoding layer feature, are the convolution and pooling operations in the encoder, is the preprocessed page image, are the convolution and pooling operations in the decoder, is the fully connected layer, is global average pooling; S22: Calculate the probability of each pixel in the content area probability image being an edge to obtain an edge probability image. The calculation method is: ; in, is the edge probability image, is the adaptive weight parameter, is the Sobel gradient field, is the content area probability image at coordinates The two-dimensional derivative at , are the horizontal and vertical coordinates of the content area probability image, respectively. Respectively represent the gradient components of the preprocessed page image in the horizontal and vertical directions, is the Sobel operator operation in the horizontal direction, is the Sobel operator operation in the vertical direction, is the adaptive weight hyperparameter, is the variance, is the gradient strength of the Sobel gradient field.
4. The method for automatic correction of printed documents based on deep learning according to claim 3, characterized in that: The S3 step comprises: When the accumulator matrix performs accumulation calculation, the pixel value of the content area probability image at the coordinate is used as the accumulation weight; wherein the polar angle parameter is selected The range is .
5. The method for automatic correction of printed documents based on deep learning according to claim 4, characterized in that: The S3 step comprises: S31: According to the edge probability image, an improved Hough transform is designed to perform line segment clustering and obtain the accumulator matrix. The calculation method is: ; in, is the accumulator matrix, is the content area probability image at coordinates The pixel value at is the impulse function, is the polar diameter parameter; S32: performing local peak detection on the accumulator matrix, and using a clustering algorithm on each detected local peak to divide adjacent local peaks with the same direction into the same group, thereby obtaining a set of file boundary candidate line segments; S33: in the set of file boundary candidate line segments, according to the polar angle parameter of each file boundary candidate line segment, a weighted average method is used to calculate the main direction angle; the main direction angle is a weighted average of the polar angle parameters of each file boundary candidate line segment, and the weight is the length of the corresponding file boundary candidate line segment; S34: in the set of file boundary candidate line segments, retain the straight line segment with the longest length and the smallest deviation from the main direction angle as the file boundary line segment, and output the polar radius parameter and polar angle parameter corresponding to the file boundary line segment; S35: Calculate the page tilt angle according to the polar angle parameter of the file boundary line segment, and the calculation method is: ; in, is the page tilt angle, is the reference direction angle, ; is the polar angle parameter of the file boundary segment; S36: Rotate the pre-processed page image according to the page tilt angle to obtain a page image after tilt correction .
6. The method for automatic deviation correction of printed documents based on deep learning according to claim 5, characterized in that: The S4 step comprises: The design of the two-stream Transformer network includes: constructing parallel geometric streams and edge streams, and extracting geometric features and edge features respectively; introducing inter-stream attention weights, fusing information of geometric features and edge features, and updating the geometric features by layer normalization to obtain updated geometric features; designing a two-stream gate to fuse edge features and updated geometric features to obtain fused features.
7. The method for automatic correction of printed documents based on deep learning according to claim 6, characterized in that: The S4 step comprises: S41: Design a two-stream Transformer network to extract edge features and geometric features based on the page tilt angle and edge probability image, and then extract fusion features. The calculation method is: ; in, is the geometric feature, For the Transformer network, is the edge feature, is a convolutional neural network, is the attention weight between streams, is the Softmax function, is the transpose of the edge feature, is the feature dimension, is the updated geometric feature, is layer normalization, For double flow door, is the double-stream gate weight, is the dual-stream gate bias vector, For fusion features; S42: judging whether the page image after the tilt correction has page deformation according to the fusion feature, the calculation method is: ; in, To determine whether there is a page deformation result, When the page is deformed, It is considered that there is no page deformation; S43: If there is page deformation, the grid offset is calculated, including the horizontal grid offset and the vertical grid offset, and the calculation method is: ; in, is the deformation probability image, is the horizontal grid offset, is the hyperbolic tangent function, is the vertical grid offset; According to the horizontal grid offset and the vertical grid offset, the page image after tilt correction is deformed by bilinear interpolation to obtain the page image after deformation correction. ; S44: If there is no page deformation, the page image after tilt correction is As the page image after distortion correction.
8. The method for automatic correction of printed documents based on deep learning according to claim 7, characterized in that: The step S5 comprises: S51: configuring printing task parameters, including paper size, printing resolution, color mode, printing direction, and number of copies; S52: starting a printing task for the page image after deformation correction through a printer driver; S53: Monitor the execution status of the printing task in real time and record the printing result.
9. A printing document automatic correction system based on deep learning, characterized in that: include: The module for acquiring and preprocessing the files to be printed: acquiring the PDF files to be printed and preprocessing them in the form of images to obtain the preprocessed page images; Edge probability image calculation module: Based on the preprocessed page image, an improved U-shaped convolutional network is designed to generate a content area probability image and calculate the probability of each pixel in the content area probability image being an edge to obtain an edge probability image; Tilt correction module: Based on the edge probability image, an improved Hough transform is designed to perform line segment clustering to obtain an accumulator matrix; based on the accumulator matrix, the set of candidate line segments at the file boundary, the main direction angle, the file boundary line segments and the page tilt angle are calculated; based on the page tilt angle, the preprocessed page image is rotated to obtain a page image after tilt correction; Deformation correction module: A two-stream Transformer network is designed to extract fusion features based on the page tilt angle and edge probability image, and determine whether the page image after tilt correction has page deformation based on the fusion features; if there is page deformation, the grid offset is calculated, and the page image after tilt correction is deformed to obtain the page image after deformation correction; if there is no page deformation, the page image after tilt correction is directly used as the page image after deformation correction; Printing module: based on the page image after deformation correction, start the printing task through the printer driver and record the printing result; To realize a method for automatic correction of printed documents based on deep learning as described in any one of claims 1-8.
Citation Information
Patent Citations
Image instance segmentation method, device, apparatus, and storage medium
CN109242869A
Archive digitization method and system based on intelligent image enhancement and automatic classification
CN119049066A