A method and device for correcting the inclination of Uyghur printed pages
By performing noise reduction processing on Uyghur printed pages and scanning projection segmentation to identify the first pixel of the row, combined with Hough transformation and rigid body transformation, the problems of low efficiency and low accuracy of Uyghur text tilt correction are solved, and efficient and accurate correction effect is achieved.
Patent Information
- Application Number
- CN202210621030.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-02
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-06-02
AI Technical Summary
In the prior art, the tilt correction of Uyghur texts is inefficient and has low accuracy.
By determining the target image of the Uyghur printed page to be corrected, the noise reduction algorithm is used for pre-processing, the row head pixel is identified using a scanning projection segmentation method, the document inclination angle is determined based on the Hough transformation, and the correction is performed through the rigid body transformation.
It realizes efficient and accurate tilt correction of Uyghur printed pages, and improves the efficiency and accuracy of correction.
Smart Images

Figure CN114998900B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of printed page skew correction, and particularly relates to a method and device for correcting the skew of Uyghur printed pages. Background Art
[0002] As a way with strong generalization and logical expression ability, characters play a role in preserving and spreading cultural knowledge. Since paper documents are not easy to preserve and manage, with the rapid development of computer science and technology, paper documents can be digitized using Optical Character Recognition (OCR). Uyghur, as the common language of the Uyghur people, has a large number of users and a high usage frequency. Since it is inevitable that document images are skewed during shooting or scanning, this will lead to errors in subsequent text cutting and cause great difficulties in text understanding. It is necessary to use a skew correction method to remove the image skew for convenient subsequent operations.
[0003] There are many existing skew correction methods. For example, performing Fourier transform on an image to obtain a spectrum, and obtaining the text skew angle through the spectrum; obtaining the skew angle through the contour projection method; or performing thinning operation on the text image, and then performing least squares fitting on the thinned skeleton to obtain the skew angle; obtaining the minimum circumscribed rectangle of the document using the document contour to obtain the document skew angle, etc.
[0004] However, the existing methods are not ideal for the skew correction of Uyghur text, and some of these correction methods require a lot of time for calculation. Therefore, how to avoid the low efficiency and low accuracy of the existing Uyghur correction is still an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0005] The purpose of the present invention is to overcome the above technical deficiencies, and provide a method and device for correcting the skew of Uyghur printed pages, so as to solve the technical problems of low efficiency and low accuracy in Uyghur correction in the prior art.
[0006] To achieve the above technical purpose, the technical solution of the present invention provides a method for correcting the skew of Uyghur printed pages, including:
[0007] S1: Determine the target image of the Uyghur printed matter to be corrected;
[0008] S2: Preprocess the target image using a noise reduction algorithm to obtain a denoised image;
[0009] S3: Identify the row start pixels in the denoised image using a scanning projection segmentation method;
[0010] S4: Determine the document skew angle based on the row start pixels through Hough transform;
[0011] S5: Apply a rigid transformation to the target image based on the document tilt angle to obtain a corrected image.
[0012] A Uyghur printed page tilt correction method provided by the present invention further includes:
[0013] S6: If the corrected image meets the preset conditions, the correction is successful; otherwise, return to S4, update the starting pixels of the rows for Hough transform, determine the document tilt angle, and execute S5 - S6 until the corrected image meets the preset conditions.
[0014] For a Uyghur printed page tilt correction method provided by the present invention, the S2 specifically includes:
[0015] Convert the target image into a binary image;
[0016] Then perform connected - component filtering denoising on the binary image to obtain a denoised image.
[0017] For a Uyghur printed page tilt correction method provided by the present invention, the S3 specifically includes:
[0018] Perform a right - half image scan on the denoised image from right to left and from top to bottom. If the position of the first foreground point scanned is in the lower half of the image, it is determined that the document is tilted to the left; if the position of the first foreground point scanned is in the upper half of the image, it is determined that the document is tilted to the right;
[0019] Based on the document tilt direction, in a projection - splitting manner, determine and save the starting pixels of the image text lines in an array.
[0020] For a Uyghur printed page tilt correction method provided by the present invention, the step of determining and saving the starting pixels of the image text lines in an array based on the document tilt direction in a projection - splitting manner specifically includes:
[0021] If the document tilt direction is to the left, determine the starting point of each text line through projection from right to left and from bottom to top, save it in an array, and remove the area below and to the left of the starting point. Use the starting point detected in the previous step as the starting point of the projection, and repeat until the pixel point is less than the first preset threshold away from the end of the document;
[0022] If the document tilt direction is to the left, determine the starting point of each text line through projection from right to left and from top to bottom, save it in an array, and remove the area above and to the left of the starting point. Use the starting point detected in the previous step as the starting point of the projection, and repeat until the pixel point is less than the first preset threshold away from the end of the document.
[0023] A Uyghur printed page skew correction method provided by the present invention, the S4 specifically includes:
[0024] Based on the pixel point straight line formula, use Hough transform to detect the straight line formed by the collinear leading pixel points in the leading pixel points, where the pixel point straight line formula is:
[0025]
[0026] Wherein, is the distance from the origin to the detected straight line, is the inclination angle of the detected straight line;
[0027] For the straight line detected by Hough transform, obtain the document inclination angle .
[0028] A Uyghur printed page skew correction method provided by the present invention, the S6 specifically includes:
[0029] If the corrected image satisfies that the blank gap is greater than the second preset threshold, or satisfies that the blank gap is less than the second preset threshold and the difference between the current document inclination angle and the previous document inclination angle is less than the third preset threshold, then the correction is successful,
[0030] Otherwise, return to S4 and update the leading pixel for Hough transform, determine the document inclination angle, and execute S5 - S6 until the corrected image satisfies that the blank gap is greater than the second preset threshold, or satisfies that the blank gap is less than the second preset threshold and the difference between the current document inclination angle and the previous document inclination angle is less than the third preset threshold.
[0031] A Uyghur printed page skew correction method provided by the present invention, the returning to S4 and updating the leading pixel for Hough transform, determining the document inclination angle, and executing S5 - S6 until the corrected image satisfies that the blank gap is greater than the second preset threshold, or satisfies that the blank gap is less than the second preset threshold and the difference between the current document inclination angle and the previous document inclination angle is less than the third preset threshold, specifically includes:
[0032] Return to S4 to remove the points in the leading pixels whose distance from the straight line is greater than the fourth preset threshold for updating, use the updated leading pixels for Hough transform, determine the document inclination angle, and then execute S5 - S6 until the corrected image satisfies that the blank gap is greater than the second preset threshold, or satisfies that the blank gap is less than the second preset threshold and the difference between the current document inclination angle and the previous document inclination angle is less than the third preset threshold.
[0033] A Uyghur printed page skew correction method provided by the present invention, the S5 specifically includes:
[0034] Convert the target image into a corrected image based on the rigid transformation formula, where the rigid transformation formula is:
[0035]
[0036] where, x, y is the position of any pixel point in the target image, , is the position in the corrected image corresponding to the any pixel point x, y .
[0037] The present invention also provides a Uyghur printed page skew correction device, including:
[0038] A determination unit for determining a target image of the Uyghur printed matter to be corrected;
[0039] A preprocessing unit for preprocessing the target image by using a noise reduction algorithm to obtain a noise-reduced image;
[0040] A scanning and projection unit for identifying the first-pixel of the line in the noise-reduced image by using a scanning and projection segmentation method;
[0041] A transformation unit for determining the document skew angle based on the first-pixel of the line by performing a Hough transform;
[0042] A correction unit for using a rigid body transformation on the target image based on the document skew angle to obtain a corrected image.
[0043] Compared with the prior art, the beneficial effects of the present invention include: providing an efficient and accurate skew correction method to deal with the skew problem of Uyghur printed pages. Description of the Drawings
[0044] Figure 1 It is a schematic flowchart of a Uyghur printed page skew correction method provided by the present invention;
[0045] Figure 2 It is a schematic flowchart of another Uyghur printed page skew correction method provided by the present invention;
[0046] Figure 3 It is a sample pattern of a Uyghur printed page provided by the present invention;
[0047] Figure 4 It is an effect diagram of a Uyghur scanned page after connected component filtering provided by the present invention;
[0048] Figure 5 It is the detection result of the first letter pixel points of a Uyghur scanned page provided by the present invention;
[0049] Figure 6The effect diagram of the inclination correction of the Uyghur scanned page provided by the present invention;
[0050] Figure 7 The structural schematic diagram of a Uyghur printed page inclination correction device provided by the present invention;
[0051] Figure 8 The physical structure schematic diagram of an electronic device provided by the present invention. Specific embodiments
[0052] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following further elaborates on the present invention in conjunction with the attached drawings Figures 1 - 8 , and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0053] Figure 1 A Uyghur printed page inclination correction method provided by the present invention, as Figure 1 shown, the method includes:
[0054] Step 110, determining the target image of the Uyghur printed matter to be corrected.
[0055] Specifically, first, it is necessary to determine the image of the Uyghur printed page to be corrected, ensuring that the text content in the image is Uyghur. The format of the image can be a color image or a grayscale format.
[0056] Step 120, preprocessing the target image by using a noise reduction algorithm to obtain a denoised image.
[0057] Specifically, for the input text image, it is first necessary to determine whether it is a grayscale image or a color image, and convert it into a binary image in a suitable manner respectively. Then, since Uyghur has identifiers, and the positions of the identifiers often exist in the upper or middle and lower parts of the characters, which will cause difficulties in obtaining the first pixel of the text line. Therefore, it is necessary to perform noise reduction preprocessing on the image to remove these identifiers and page stains caused by printing and other problems, and obtain a denoised image.
[0058] Step 130, identifying the first pixel of the line in the denoised image by using a scanning projection segmentation method.
[0059] Specifically, since there are only foreground points and background points in the binary image, the identification of the first pixel of the line only needs to identify the first foreground point that appears in each line, that is, the projection segmentation method can be used to segment out the first foreground point of each line row by row.
[0060] Step 140, determining the document tilt angle based on the first pixel of the line by using the Hough transform.
[0061] Specifically, after obtaining the first pixel points of each line, it is necessary to determine the inclination angle of the straight line formed by connecting these first pixel points of the lines. Here, the Hough detection method is preferably used to detect all collinear first pixel points of the lines, then the straight line is obtained, and the inclination angle of the straight line is used as the document inclination angle.
[0062] Step 150, based on the document inclination angle, perform a rigid body transformation on the target image to obtain a corrected image.
[0063] Specifically, after obtaining the document inclination angle, a corrected image can be obtained by using a rigid body transformation.
[0064] The Uyghur printed page inclination correction method provided by the embodiments of the present invention includes: determining a target image of the Uyghur printed matter to be corrected; preprocessing the target image by using a noise reduction algorithm to obtain a noise-reduced image; identifying the first pixel points in the noise-reduced image by using a scanning projection segmentation method; determining the document inclination angle based on the first pixel points by performing a Hough transform; performing a rigid body transformation on the target image based on the document inclination angle to obtain a corrected image; if the corrected image meets a preset condition, the correction is successful, otherwise, return to S4 and update the first pixel points for performing the Hough transform, determine the document inclination angle, and execute S5-S6 until the corrected image meets the preset condition, and obtain a corrected image after passing the preset condition test. Therefore, the present invention provides an efficient and accurate inclination correction method to deal with the inclination problem of Uyghur printed pages.
[0065] Based on the above embodiments, in this method, it further includes:
[0066] Step 160, if the corrected image meets a preset condition, the correction is successful, otherwise, return to step 140 and update the first pixel points for performing the Hough transform, determine the document inclination angle, and execute steps 150-160 until the corrected image meets the preset condition.
[0067] Specifically, a verification condition, that is, a preset condition, is preset here. If the corrected image meets the verification condition, the inclination correction is successful. If the corrected image does not meet the preset condition, it means that the inclination correction is not successful, and the document inclination angle needs to be recalculated, that is, the first pixel points participating in the calculation of the document inclination angle need to be re-obtained. It should be noted here that the preset condition can be whether the size of the blank gap between lines of the corrected image is greater than a second preset threshold, or whether the difference between the currently calculated document inclination angle and the previously calculated document inclination angle is less than a third preset threshold, or any combination of the two.
[0068] Based on the above embodiments, in the step 120 of this method, it specifically includes:
[0069] Convert the target image into a binary image;
[0070] Then, perform connected component filtering and noise reduction on the binary image to obtain the denoised image.
[0071] Specifically, for the input text image, it is first necessary to determine whether it is a grayscale image. If so, it needs to be binarized to convert it into a binary image. If not, a grayscale conversion operation needs to be added before the binarization step to convert the three-channel color image into a single-channel grayscale image. Here, since there is a lot of text information in the image, the global binarization method, i.e., the bimodal method, is used for binarization. In the grayscale conversion process 、 、 are the red, green, and blue channels of the color image respectively, and the process is as follows:
[0072] In the bimodal method, the value of the threshold T is the gray value corresponding to the trough in the gray histogram of the grayscale image. The binarization value-taking process is implemented through the following formula:
[0073]
[0074] After binarization, the image only has two pixel values, 0 and 255. The binarization operation can effectively weaken the influence caused by creases and slight printing through.
[0075] Since Uyghur text has identifiers, and the positions of the identifiers often exist in the upper or middle and lower parts of the characters, this will cause difficulties in obtaining the pixels at the beginning of the text line. Therefore, it is necessary to perform connected component analysis on the image to remove these identifiers and page stains caused by printing and other problems. Using the seed filling algorithm can effectively retrieve the connected components. By setting an area threshold , the connected components with an area smaller than this threshold are set to the background color to achieve the effect of filtering out the identifiers and provide conditions for subsequent detection of the pixels at the beginning of the text line.
[0076] The connected component filtering algorithm is as follows: Set an area threshold , scan the entire image, find the first foreground pixel, set it as the seed point, and push it onto the stack for storage; then pop the stack, traverse downward from the popped point until a non-foreground point or a boundary point is reached and stop. Judge whether the left and right pixel points of this pixel point are foreground points. If so, save them onto the stack; then fill upward from the stop point until a boundary point or a background point is reached and stop. During this process, judge the left and right pixel points, and the method is the same as above; finally, repeat this process until all connected components are found. When all connected components are found, count the areas of each connected component. If it is less than the threshold then set this connected component to the background color.
[0077] The specific process of the connected component filtering algorithm is as follows: (1) Create three variables. An integer variable AreaPoint is used to store the number of pixel points and is initialized to 0. The other two boolean variables LeftPoint and RightPoint are used as criteria for judging the left and right pixel points of the filling point and are initialized to False. (2) By traversing the entire image, find the first foreground point, use it as the seed point, and create a stack to push this point onto the stack. (3) Perform a pop operation, check the point popped out downward until a boundary point or a non-foreground pixel point is found and stop. (4) Fill the stop point, increment the variable AreaPoint by 1. Then judge whether the left and right pixel points of this point are foreground points. If the left side is, set the variable LeftPoint to True and push the left pixel point onto the stack, otherwise set it to False. The same applies to the right side. (5) Fill upward from the stop point, increment AreaPoint by 1, judge the left and right pixels until a boundary point or a non-foreground pixel point is found and stop. This ends one round. (6) Repeat the process of (3), (4), and (5) until the stack is emptied by popping. (7) Compare AreaPoint with the area threshold If it is less, set all points in this area to the background color, and if it is greater than or equal, keep them. (8) Then repeat the above process until all connected components are found.
[0078] Based on the above embodiments, in this method, step 130 specifically includes:
[0079] Perform a scan of the right half of the denoised image from right to left and from top to bottom. If the position of the first scanned foreground point is in the lower half of the image, it is determined that the document is tilted to the left. If the position of the first scanned foreground point is in the upper half of the image, it is determined that the document is tilted to the right;
[0080] Based on the tilt direction of the document, in a projection splitting manner, determine and save the starting pixels of the image text lines in an array.
[0081] Specifically, perform a scan of the right half of the denoised image from right to left and from top to bottom. Here, an explanation of the right half of the image is given. The right half of the image means cutting the original image in half from left to right and taking the right half of the image. Similarly, the left half of the image, as well as the upper half and lower half of the image to be used later, can be understood. For the first scanned foreground point, judge whether it is in the upper half or the lower half of the image. If it is in the lower half, the document is tilted to the left, and vice versa, it is tilted to the right.
[0082] Then, based on the tilt direction of the document, perform projection splitting on the document to determine the starting pixel points of each line of the image text lines and save them sequentially in an array.
[0083] Based on the above embodiments, in this method, determining and saving the starting pixels of the image text lines in an array in a projection splitting manner based on the document tilt direction specifically includes:
[0084] If the document tilt direction is to the left, determine the starting point of each text line through projection from right to left and from bottom to top, save it in an array, and remove the areas below and to the left of the starting point. Using the starting point detected in the previous step as the starting point of the projection, repeat until the pixel point is less than the first preset threshold away from the end of the document;
[0085] If the document tilt direction is to the left, determine the starting point of each text line through projection from right to left and from top to bottom, save it in an array, and remove the areas above and to the left of the starting point. Using the starting point detected in the previous step as the starting point of the projection, repeat until the pixel point is less than the first preset threshold away from the end of the document.
[0086] Specifically, if the document tilt direction is to the left, determine the starting point of each text line through projection from right to left and from bottom to top, save it in an array, and remove the areas below and to the left of the starting point. Using the starting point detected in the previous step as the starting point of the projection, repeat until the pixel point is less than the first preset threshold away from the end of the document; that is, the upper left corner remains unchanged, and the lower right corner vertex is always modified. Each projection obtains a new rectangle, and then find a new foreground point in the new rectangle until the height of the new rectangle is less than the first preset threshold;
[0087] If the document tilt direction is to the right, determine the starting point of each text line through projection from right to left and from top to bottom, save it in an array, and remove the areas above and to the left of the starting point. Using the starting point detected in the previous step as the starting point of the projection, repeat until the pixel point is less than the first preset threshold away from the end of the document; that is, the lower left corner remains unchanged, and the upper right corner vertex is always modified. Each projection obtains a new rectangle, and then find a new foreground point in the new rectangle until the height of the new rectangle is less than the first preset threshold.
[0088] Based on the above embodiments, in this method, step 140 specifically includes:
[0089] Using the Hough transform to detect the lines formed by the collinear starting pixels among the starting pixels of the text lines based on the pixel point straight line formula, where the pixel point straight line formula is:
[0090]
[0091] Where, is the distance from the origin to the detected line, is the inclination angle of the detected line;
[0092] The inclination angle of the document is obtained from the straight line detected by the Hough transform .
[0093] Specifically, the inclination angle of the document is obtained by using the Hough transform to detect the straight line formed by the starting pixel points of the text lines , and the Hough transform process is as follows, where is the distance from the origin to the detected straight line, is the inclination angle of the detected straight line:
[0094] .
[0095] Based on the above embodiments, in this method, step 160 specifically includes:
[0096] If the corrected image satisfies that the blank gap is greater than the second preset threshold, or satisfies that the blank gap is less than the second preset threshold and the difference between the current document inclination angle and the previous document inclination angle is less than the third preset threshold, then the correction is successful
[0097] Otherwise, return to S4, update the starting pixel of the line for the Hough transform, determine the document inclination angle, and execute S5 - S6 until the corrected image satisfies that the blank gap is greater than the second preset threshold, or satisfies that the blank gap is less than the second preset threshold and the difference between the current document inclination angle and the previous document inclination angle is less than the third preset threshold
[0098] Specifically, perform a horizontal projection on the corrected image. In the preset conditions, it is necessary to judge whether the blank gap of the corrected image is greater than the second preset threshold. If it is greater, then the preset conditions are satisfied, the correction is successful, the corrected image is available, and the loop is exited. Otherwise, it is also necessary to further judge whether the difference between the currently calculated document inclination angle and the previously calculated document inclination angle is less than the third preset threshold. If it is less, then the preset conditions are satisfied, the correction is successful, the corrected image is available, and the loop is exited. Otherwise, re - select the starting pixel points of the lines participating in the Hough transform, continue to execute S5 - S6 until the preset conditions are satisfied, that is, until the corrected image satisfies that the blank gap is greater than the second preset threshold, or satisfies that the blank gap is less than the second preset threshold and the difference between the current document inclination angle and the previous document inclination angle is less than the third preset threshold
[0099] Based on the above embodiments, in this method, the returning to S4, updating the starting pixel of the line for the Hough transform, determining the document inclination angle, and executing S5 - S6 until the corrected image satisfies that the blank gap is greater than the second preset threshold, or satisfies that the blank gap is less than the second preset threshold and the difference between the current document inclination angle and the previous document inclination angle is less than the third preset threshold specifically includes:
[0100] Return to S4 to update by removing the points in the head pixels of the line that are at a distance greater than the fourth preset threshold from the line, perform Hough transform using the updated head pixels of the line to determine the document tilt angle, and then execute S5 - S6 until the corrected image satisfies that the blank gap is greater than the second preset threshold, or satisfies that the blank gap is less than the second preset threshold and the difference between the current document tilt angle and the previous document tilt angle is less than the third preset threshold.
[0101] Specifically, perform horizontal projection on the corrected image to determine whether the blank gap is greater than the preset second preset threshold or determine whether the difference between the newly detected image tilt angle and the previous tilt angle (initialized to 0) is less than the third preset threshold. Preferably, the third preset threshold is set to 10 - 2. If the condition is met, the correction is completed and the program ends; otherwise, remove the points in the array that are at a distance greater than the distance threshold (the fourth preset threshold) from the line, and then perform Hough detection on the array, repeating this process until the condition is met.
[0102] Based on the above embodiments, in this method, step 150 specifically includes:
[0103] Convert the target image into a corrected image based on the rigid transformation formula, where the rigid transformation formula is:
[0104]
[0105] where, x, y is the position of any pixel point in the target image, , is the position in the corrected image corresponding to the any pixel point x, y .
[0106] Specifically, since the image only changes in position during the transformation process and does not change in shape, that is, a rigid body transformation occurs. The coordinate transformation process is as follows, where ( , ) is the coordinate after the rigid body transformation, and ( x , y ) is the coordinate of the tilted document before the rigid body transformation, is the document tilt angle:
[0107]
[0108] Based on the above embodiments, the present invention provides another method for correcting the tilt of a Uyghur printed page. Figure 2 This is a schematic flow diagram of another method for correcting the tilt of a Uyghur printed page provided by the present invention. As Figure 2 shown, this method includes: Step 1, set thresholds , use the seed filling algorithm to perform connected component filtering on the binary image, count the areas of each connected component, and set the connected components with areas smaller than the threshold to the background color; Step 2, project the filtered image from right to left and from top to bottom on the right half of the image to find the first foreground pixel. If the pixel is in the upper (lower) half of the text image, the document shows a right (left) tilt phenomenon; Step 3, determine the process through different tilt directions. If it is a left (right) tilted document, adopt a projection method from right to left and from bottom to top (from top to bottom), find and save the starting pixel points of the text lines, use this pixel point as the lower (upper) boundary point of the new image, and extract this area from the original image; Step 4, repeat Step 3 in the new image until the distance between the pixel point and the boundary is less than the threshold Th dis to stop; Step 5, use the Hough transform to perform line detection on the saved starting pixel points of the text lines. At this time, the tilt angle of the current text line can be obtained, and this angle can be used to approximate the tilt angle of this text subsequently; Step 6, use rigid body transformation to perform tilt correction on the text image; Step 7, perform line projection on the corrected image to determine whether the line spacing meets the threshold , if the condition is met, the algorithm ends. If not, remove the pixel points whose distance from the line is greater than the threshold , and then go back to Step 5 for processing, where the algorithm ends when the difference between the newly detected tilt angle and the previously detected tilt angle (initialized to 0) is less than 10-2. The thresholds set in this solution are all empirical values obtained from multiple experiments.
[0109] Figure 3 is the sample pattern of the Uyghur printed page provided by the present invention, Figure 4 is the effect diagram of the Uyghur scanned page after connected component filtering provided by the present invention, Figure 5 is the detection result of the starting letter pixel points of the Uyghur scanned page provided by the present invention, Figure 6 is the effect diagram of tilt correction of the Uyghur scanned page provided by the present invention.
[0110] The beneficial effects produced by the present invention are as follows: The method of using the starting pixel points of the text lines in the present invention to perform Hough transform to detect the tilt angle of the text realizes the tilt correction operation of the scanned Uyghur printed page. Moreover, the present invention can also perform tilt correction on the document with an image in the center of the text. When the image is in the center of the text, using the method described in the text can avoid the problem that the text appears disconnected due to the embedded arrangement of the image and it is impossible to use the detection baseline to determine the tilt angle.
[0111] Next, the Uyghur printed page tilt correction device provided by the present invention will be described. The three-dimensional model feature extraction device described below can be mutually corresponding and referred to with the three-dimensional model feature extraction method described above.
[0112] Figure 7 The structural schematic diagram of a Uyghur printed page inclination correction device provided by the present invention is as follows. Figure 7 As shown, the device includes a determination unit 101, a preprocessing unit 102, a scanning and projection unit 103, a transformation unit 104, and a correction unit 105. Among them,
[0113] The determination unit 101 is used to determine the target image of the Uyghur printed matter to be corrected;
[0114] The preprocessing unit 102 is used to preprocess the target image by using a noise reduction algorithm to obtain a denoised image;
[0115] The scanning and projection unit 103 is used to identify the first-pixel of each line in the denoised image by using a scanning and projection segmentation method;
[0116] The transformation unit 104 is used to determine the document inclination angle by performing Hough transformation based on the first-pixel of each line;
[0117] The correction unit 105 is used to perform rigid body transformation on the target image based on the document inclination angle to obtain a corrected image.
[0118] The device provided by the present invention determines the target image of the Uyghur printed matter to be corrected; preprocesses the target image by using a noise reduction algorithm to obtain a denoised image; identifies the first-pixel of each line in the denoised image by using a scanning and projection segmentation method; determines the document inclination angle by performing Hough transformation based on the first-pixel of each line; and performs rigid body transformation on the target image based on the document inclination angle to obtain a corrected image. Therefore, the present invention provides an efficient and accurate inclination correction method to deal with the inclination problem of Uyghur printed pages.
[0119] Based on the above embodiments, in this device, there is also a determination unit, which is used for:
[0120] Determine that if the corrected image meets the preset conditions, the correction is successful; otherwise, return to the transformation unit and update the first-pixel of each line for performing Hough transformation, determine the document inclination angle, and execute the process from the correction unit to the determination unit until the corrected image meets the preset conditions.
[0121] Based on the above embodiments, in this device, the preprocessing unit is specifically used for:
[0122] Convert the target image into a binary image;
[0123] Then perform connected component filtering and noise reduction on the binary image to obtain a denoised image.
[0124] Based on the above embodiments, in this device, the scanning and projection unit is specifically used for:
[0125] Perform a right - half image scan on the denoised image from right to left and from top to bottom. If the position of the first foreground point scanned is in the lower half of the image, it is determined that the document is tilted to the left. If the position of the first foreground point scanned is in the upper half of the image, it is determined that the document is tilted to the right;
[0126] Based on the document tilt direction, in a projection - splitting manner, determine and save the starting pixels of the image text lines in an array.
[0127] Based on the above - mentioned embodiments, in this device, based on the document tilt direction, in a projection - splitting manner, determine and save the starting pixels of the image text lines in an array. Specifically, it includes:
[0128] If the document tilt direction is to the left, determine the starting point of each text line through projection from right to left and from bottom to top, save it in an array, and remove the areas below and to the left of the starting point. Taking the starting point detected in the previous step as the starting point of the projection, repeat until the pixel point is less than the first preset threshold away from the end of the document;
[0129] If the document tilt direction is to the left, determine the starting point of each text line through projection from right to left and from top to bottom, save it in an array, and remove the areas above and to the left of the starting point. Taking the starting point detected in the previous step as the starting point of the projection, repeat until the pixel point is less than the first preset threshold away from the end of the document.
[0130] Based on the above - mentioned embodiments, in this device, the transformation unit is specifically used for:
[0131] Based on the pixel - point straight - line formula, use the Hough transform to detect the straight line formed by the collinear starting pixels among the starting pixels of the lines, where the pixel - point straight - line formula is:
[0132]
[0133] where, is the distance from the origin to the detected straight line, is the inclination angle of the detected straight line;
[0134] Obtain the document tilt angle from the straight line detected by the Hough transform .
[0135] Based on the above - mentioned embodiments, in this device, the determination unit is specifically used for:
[0136] If the corrected image satisfies that the blank gap is greater than the second preset threshold, or satisfies that the blank gap is less than the second preset threshold and the difference between the current document tilt angle and the previous document tilt angle is less than the third preset threshold, then the correction is successful,
[0137] Otherwise, return to S4, update the starting pixels of the rows for performing the Hough transform, determine the document tilt angle, and execute S5 - S6 until the corrected image satisfies that the blank gap is greater than the second preset threshold, or satisfies that the blank gap is less than the second preset threshold and the difference between the current document tilt angle and the previous document tilt angle is less than the third preset threshold.
[0138] Based on the above embodiments, in this device, the returning to S4, updating the starting pixels of the rows for performing the Hough transform, determining the document tilt angle, and executing S5 - S6 until the corrected image satisfies that the blank gap is greater than the second preset threshold, or satisfies that the blank gap is less than the second preset threshold and the difference between the current document tilt angle and the previous document tilt angle is less than the third preset threshold specifically includes:
[0139] Return to S4, remove the points in the starting pixels whose distance from the straight line is greater than the fourth preset threshold for updating, perform the Hough transform using the updated starting pixels, determine the document tilt angle, and then execute S5 - S6 until the corrected image satisfies that the blank gap is greater than the second preset threshold, or satisfies that the blank gap is less than the second preset threshold and the difference between the current document tilt angle and the previous document tilt angle is less than the third preset threshold.
[0140] Based on the above embodiments, in this device, the correction unit specifically includes:
[0141] Convert the target image into a corrected image based on the rigid transformation formula, where the rigid transformation formula is:
[0142]
[0143] Where, x, y is the position of any pixel point in the target image, , is the position in the corrected image corresponding to the any pixel point x, y .
[0144] Figure 8 This is a schematic physical structure diagram of an electronic device provided by the present invention, as Figure 8As shown in the figure, the electronic device may include: a processor 1110, a communications interface 1120, a memory 1130, and a communication bus 1140. Among them, the processor 1110, the communications interface 1120, and the memory 1130 complete communication with each other through the communication bus 1140. The processor 1110 may call the logic instructions in the memory 1130 to execute the Uyghur printed page skew correction method, which includes: S1: Determine the target image of the Uyghur print to be corrected; S2: Preprocess the target image using a noise reduction algorithm to obtain a denoised image; S3: Use a scanning projection segmentation method to identify the leading pixels in the denoised image; S4: Perform a Hough transform based on the leading pixels to determine the document skew angle; S5: Use a rigid body transformation on the target image based on the document skew angle to obtain a corrected image.
[0145] In addition, when the logic instructions in the above-mentioned memory 1130 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0146] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the Uyghur printed page skew correction method provided by the above-mentioned methods, which includes: S1: Determine the target image of the Uyghur print to be corrected; S2: Preprocess the target image using a noise reduction algorithm to obtain a denoised image; S3: Use a scanning projection segmentation method to identify the leading pixels in the denoised image; S4: Perform a Hough transform based on the leading pixels to determine the document skew angle; S5: Use a rigid body transformation on the target image based on the document skew angle to obtain a corrected image.
[0147] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the Uyghur printed page skew correction method provided by the above-mentioned various methods. The method includes: S1: determining a target image of the Uyghur print to be corrected; S2: preprocessing the target image by using a noise reduction algorithm to obtain a denoised image; S3: identifying the leading pixel in the denoised image by using a scanning projection segmentation method; S4: determining the document skew angle based on the leading pixel by using Hough transform; S5: using a rigid body transformation on the target image based on the document skew angle to obtain a corrected image.
[0148] The server embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0149] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for correcting the inclination of Uyghur printed pages, characterized in that, it includes: S1: Determine the target image of the Uyghur printed matter to be corrected; S2: Preprocess the target image using a noise reduction algorithm to obtain a denoised image; S3: Identify the starting pixels of each line in the denoised image using a scanning projection segmentation method; S4: Determine the document inclination angle based on the starting pixels through Hough transform; S5: Use rigid body transformation on the target image based on the document inclination angle to obtain a corrected image; S6: If the corrected image meets the preset conditions, the correction is successful; otherwise, return to S4, update the starting pixels for Hough transform, determine the document inclination angle, and execute S5 - S6 until the corrected image meets the preset conditions; The specific content of S3 includes: Based on the document inclination direction, use the projection segmentation method to determine and save the starting pixels of each line of the image text in an array, specifically including: If the document inclination direction is to the left, determine the starting point of each text line through projection from right to left and from bottom to top, save it in an array, and remove the area below and to the left of the starting point. Use the detected starting point as the starting point of the projection and repeat until the pixel point is less than the first preset threshold away from the end of the document; If the document inclination direction is to the right, determine the starting point of each text line through projection from right to left and from top to bottom, save it in an array, and remove the area above and to the left of the starting point. Use the detected starting point as the starting point of the projection and repeat until the pixel point is less than the first preset threshold away from the end of the document; The specific content of S4 includes: Use Hough transform based on the pixel point straight line formula to detect the straight line formed by the collinear starting pixels among the starting pixels of each line, where the pixel point straight line formula is: Among them, is the distance from the origin to the detected straight line, is the inclination angle of the detected straight line; The inclination angle of the document is obtained from the straight line detected by the Hough transform ; The specific content of S6 includes: If the corrected image meets the condition that the blank gap is greater than the second preset threshold, or meets the condition that the blank gap is less than the second preset threshold and the difference between the current document inclination angle and the previous document inclination angle is less than the third preset threshold, then the correction is successful, otherwise, return to S4, update the starting pixels for Hough transform, determine the document inclination angle, and execute S5 - S6 until the corrected image meets the condition that the blank gap is greater than the second preset threshold, or meets the condition that the blank gap is less than the second preset threshold and the difference between the current document inclination angle and the previous document inclination angle is less than the third preset threshold.
2. The method for correcting the inclination of Uyghur printed pages according to claim 1, characterized in that, the specific content of S2 includes: Convert the target image into a binary image; Then perform connected domain filtering and noise reduction on the binary image to obtain a denoised image.
3. The method for correcting the inclination of Uyghur printed pages according to claim 1, characterized in that, the specific content of S3 further includes: Perform a right half - image scan on the denoised image from right to left and from top to bottom. If the position of the first foreground pixel scanned is in the lower half of the image, it is determined that the document is inclined to the left; if the position of the first foreground pixel scanned is in the upper half of the image, it is determined that the document is inclined to the right.
4. The method for correcting the inclination of Uyghur printed pages according to claim 3, characterized in that, Return to S4 and update the starting pixels of the lines for performing the Hough transform, determine the document tilt angle, and execute S5 - S6 until the corrected image satisfies that the blank gap is greater than the second preset threshold, or satisfies that the blank gap is less than the second preset threshold and the difference between the current document tilt angle and the previous document tilt angle is less than the third preset threshold. Specifically, it includes: Return to S4 to remove the points in the starting pixels that are at a distance greater than the fourth preset threshold from the straight line for update, perform the Hough transform using the updated starting pixels to determine the document tilt angle, and then execute S5 - S6 until the corrected image satisfies that the blank gap is greater than the second preset threshold, or satisfies that the blank gap is less than the second preset threshold and the difference between the current document tilt angle and the previous document tilt angle is less than the third preset threshold.
5. The Uyghur printed page tilt correction method according to claim 1, characterized in that, The S5 specifically includes: Convert the target image into a corrected image based on the rigid transformation formula, where the rigid transformation formula is: Among them, x,y is the position of any pixel point in the target image, , is the position in the corrected image corresponding to the any pixel point x,y .
6. A Uyghur printed page tilt correction device for performing the Uyghur printed page tilt correction method according to any one of claims 1 - 5, characterized in that, It includes: A determination unit for determining the target image of the Uyghur printed matter to be corrected; A preprocessing unit for preprocessing the target image using a noise reduction algorithm to obtain a noise-reduced image; A scanning and projection unit for identifying the starting pixels in the noise-reduced image by using a scanning and projection segmentation method; A transformation unit for determining the document tilt angle based on the starting pixels by performing the Hough transform; A correction unit for obtaining a corrected image by using rigid body transformation on the target image based on the document tilt angle.
Citation Information
Patent Citations
Document skew detection method and system
CN102496018A
Morphology and integral projection-based printed Uygur document segmentation method
CN106372639A