Text image correction method and device, electronic equipment and storage medium

By segmenting and correcting text images, the problem of unsatisfactory correction effect of curved text images in the prior art is solved, and higher OCR recognition accuracy is achieved.

CN120126147APending Publication Date: 2025-06-10DUOYI NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510129391.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art has poor correction effects when processing curved text images, especially when curved, and it is difficult to ensure the accuracy of OCR recognition.

Method used

By segmenting the text image to be corrected, the center curve and height of the text line are extracted, the outline boundary points of the text unit are determined, and the text unit is corrected according to the preset outline boundary points, and the corrected text unit is finally spliced ​​into a flat text image.

Benefits of technology

It improves the text image correction effect, enhances the accuracy of subsequent OCR recognition, and is suitable for processing severely curved text images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126147A_ABST
    Figure CN120126147A_ABST
Patent Text Reader

Abstract

The invention relates to a text image correction method and device, electronic equipment and a storage medium, and the method comprises the steps: segmenting each text line in a to-be-corrected first text image, and obtaining a plurality of text units; the text units in the text lines are corrected one by one, and segmented curved surface correction of the first text image is achieved. And then the corrected text units are spliced, so that the bent first text image is corrected into the flat second text image, the text image correction effect is improved, and the accuracy of subsequent OCR recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a method, device, electronic device, and storage medium for correcting text images. Background Art

[0002] Optical Character Recognition (OCR) refers to the technology that enables a computer to recognize and understand text in various types of images. In traditional OCR systems, it is usually required that the text content in the image be horizontal or vertical for easy recognition. However, in many actual application scenarios, scanned or photographed documents may be curved due to various reasons, posing a great challenge to OCR recognition.

[0003] Most of the existing technologies adopt the method of full-image correction, and the correction effect is not ideal. Especially in the case of severe curvature, it is difficult to ensure the accuracy of OCR recognition. Summary of the Invention

[0004] Based on this, the purpose of the present invention is to provide a method, device, electronic device, and storage medium for correcting text images, which have the advantages of improving the correction effect of text images and the accuracy of subsequent OCR recognition.

[0005] According to the first aspect of the embodiments of the present application, a method for correcting a text image is provided, including the following steps:

[0006] Obtain a first text image to be corrected; the first text image includes a plurality of text lines;

[0007] Preprocess the first text image to obtain a mask image of the text line;

[0008] According to the mask image of the text line, determine the text area contour of the text line; extract a preset number of sampling points from the text area contour; and obtain the center curve and height of the text line according to the sampling points;

[0009] Segment the text line to obtain a plurality of text units; determine the contour boundary points of the text unit according to the center curve and height of the text line;

[0010] Correct the text unit according to the contour boundary points of the text unit and the preset contour boundary points to obtain a corrected text unit;

[0011] Stitch the corrected text units to obtain a second text image.

[0012] According to the second aspect of the embodiments of the present application, a device for correcting a text image is provided, including:

[0013] A first text image acquisition module, configured to acquire a first text image to be corrected; the first text image includes a plurality of text lines;

[0014] A mask image acquisition module, configured to preprocess the first text image to obtain a mask image of the text lines;

[0015] A center curve acquisition module, configured to determine the text region contour of the text lines according to the mask image of the text lines; extract a preset number of sampling points from the text region contour; and obtain the center curve of the text lines and the height of the text lines according to the sampling points;

[0016] A contour boundary point determination module, configured to segment the text lines to obtain a plurality of text units; and determine the contour boundary points of the text units according to the center curve of the text lines and the height of the text lines;

[0017] A text unit correction module, configured to correct the text units according to the contour boundary points of the text units and preset contour boundary points to obtain corrected text units;

[0018] A second text image acquisition module, configured to splice the corrected text units to obtain a second text image.

[0019] According to a third aspect of the embodiments of the present application, an electronic device is provided, including: a processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the text image correction method as described in any one of the above.

[0020] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the text image correction method as described in any one of the above is implemented.

[0021] In an embodiment of the present application, a first text image to be corrected is obtained; the first text image includes a plurality of text lines; the first text image is preprocessed to obtain a mask image of the text lines; according to the mask image of the text lines, the text area contour of the text lines is determined; a preset number of sampling points are extracted from the text area contour; according to the sampling points, the center curve of the text lines and the height of the text lines are obtained; the text lines are segmented to obtain a plurality of text units; according to the center curve of the text lines and the height of the text lines, the contour boundary points of the text units are determined; according to the contour boundary points of the text units and the preset contour boundary points, the text units are corrected to obtain the corrected text units; the corrected text units are spliced to obtain a second text image. In the present application, each text line in the first text image to be corrected is segmented to obtain a plurality of text units. Each text unit in the text line is corrected one by one, realizing the segmented surface correction of the first text image. Then, the corrected text units are spliced to realize the correction of the curved first text image into a flat second text image, improving the text image correction effect and the accuracy of subsequent OCR recognition.

[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application.

[0023] For better understanding and implementation, the present invention will be described in detail below with reference to the accompanying drawings. Description of the Drawings

[0024] Figure 1 It is a schematic flowchart of a text image correction method provided by an embodiment of the present application;

[0025] Figure 2 It is a structural block diagram of a text image correction device provided by an embodiment of the present application;

[0026] Figure 3 It is a schematic structural block diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0027] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0028] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0029] The terms used in the embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit the embodiments of the present application. The singular forms "a", "the", and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0030] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects and do not have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0031] In addition, in the description of the present application, unless otherwise specified, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0032] The text image correction method provided in the embodiments of the present application can be executed by a text image correction device, which can be implemented in software and / or hardware. The text image correction device can be composed of two or more physical entities or one physical entity. The text image correction device can be any electronic device installed with image processing software, and the electronic device can be an intelligent device such as a computer, a mobile phone, or a tablet.

[0033] The inventors found in the process of implementing the present invention that in the related art, corner detection, connected component detection, Hough transform and other correction algorithms are used for the entire text image. However, for large-scale surface deformation, the existing correction algorithms are difficult to accurately restore the planar form of the document, resulting in a high error rate in character recognition. Moreover, the algorithm has a high computational complexity and a slow processing speed, making it difficult to be applied to real-time or large-scale document image processing scenarios.

[0034] To this end, the present application segments each text line in the first text image to be corrected to obtain multiple text units. By correcting each text unit in the text line one by one, the segmented surface correction of the first text image is achieved. Then, the corrected text units are spliced to correct the curved first text image into a flat second text image, improving the text image correction effect and the accuracy of subsequent OCR recognition.

[0035] Please refer to Figure 1 , which is a schematic flowchart of a text image correction method provided by an embodiment of the present application. The text image correction method provided by the embodiment of the present application includes the following steps:

[0036] S10: Obtain a first text image to be corrected; the first text image includes several text lines.

[0037] Among them, the first text image to be corrected is an image with a certain degree of curvature in the text content. The first text image to be corrected can be an image obtained by a user taking a picture of the text of a document (such as a book, newspaper, etc.) based on an image acquisition device such as a camera or a scanner.

[0038] Due to reasons such as the deformation of the shooting object, such as the bending of the book or the folding of the page, or due to reasons such as the shooting angle between the image acquisition device and the shooting object not being completely horizontal, the horizontal text lines in the captured text image will be bent, and the bent text lines will cause difficulties in subsequent OCR recognition, resulting in an increase in the text recognition error rate. Therefore, after obtaining the first text image to be corrected in the embodiment of the present application, the subsequent process of correcting the text lines in the first text image is executed to restore the bent text lines in the first text image to horizontal text lines.

[0039] In the embodiment of the present application, the first text image to be corrected input by the user can be obtained, or the first text image to be corrected can be downloaded through the network.

[0040] S20: Preprocess the first text image to obtain a mask image of the text line.

[0041] Among them, the mask image of the text line is used to indicate the text area of the text line. Specifically, the area with text in the text line is the text area, and the blank area without text in the text line is the background area.

[0042] In the embodiment of the present application, the first text image is input into the text line detection network model to obtain the mask images of each text line in the first text image. Among them, the text line detection network model includes but is not limited to deep learning networks such as YOLO, CTPN, and PseNet.

[0043] S30: Determine the contour of the text area of the text line based on the masked image of the text line; extract a preset number of sampling points from the contour of the text area; obtain the center curve of the text line and the height of the text line according to the sampling points.

[0044] Among them, the preset number can be set according to actual needs.

[0045] Among them, the center curve is used to indicate the bending direction of the characters in the text line.

[0046] In the embodiment of the present application, based on the masked image of the text line, the findContours function in Opencv can be used to traverse each text line area to obtain the contour of the text area of each text line.

[0047] The preset number of sampling points can be evenly selected from the upper and lower sides of the contour of the text area. According to the position coordinates of the sampling points, the center curve of the text line and the height of the text line can be calculated.

[0048] S40: Segment the text line to obtain several text units; determine the contour boundary points of the text units according to the center curve of the text line and the height of the text line.

[0049] Among them, considering that the bending degree of each text line is different, when segmenting the text line, if the number of segments is too large, each text unit contains fewer characters and the correction efficiency is not high. If the number of segments is too small, each text unit contains more characters and the change in the bending degree between characters is large, and the correction accuracy is not high. Therefore, the segmentation size of the text unit can be determined according to the width of the text line and the preset number of divided segments. For example, the preset number of divided segments is 5-7 segments.

[0050] Among them, the contour boundary point refers to the vertex of the contour enclosing the text unit.

[0051] In the embodiment of the present application, the text line is evenly segmented according to the preset number of divided segments to obtain a preset number of text units.

[0052] According to the center curve of the text line, the center point position of the text unit can be determined. According to the center point position and the height of the text line, the upper boundary and the lower boundary of the contour of the text unit can be calculated to obtain the contour boundary points of the text unit.

[0053] S50: Correct the text unit according to the contour boundary points of the text unit and the preset contour boundary points to obtain the corrected text unit.

[0054] Among them, the number of preset contour boundary points is the same as the number of contour boundary points of the text unit. The position coordinates of the preset contour boundary points are set according to actual needs.

[0055] In the embodiment of the present application, the preset contour boundary points are connected to form a rectangle. By transforming the contour boundary points of the text unit into the preset contour boundary points, the surface deformation correction of the entire text unit is performed to obtain the corrected text unit.

[0056] S60: Stitch the corrected text units to obtain a second text image.

[0057] In the embodiment of the present application, after correcting each text unit of each text line, the corrected text units are stitched in the order of their positions in the text line to obtain the corrected text line. The corrected text lines are stitched in the order of their positions in the first text image to obtain the second text image.

[0058] Applying the embodiment of the present application, by obtaining a first text image to be corrected; the first text image includes a plurality of text lines; performing preprocessing on the first text image to obtain a mask image of the text line; determining the text area contour of the text line according to the mask image of the text line; extracting a preset number of sampling points from the text area contour; obtaining the central curve and the height of the text line according to the sampling points; segmenting the text line to obtain a plurality of text units; determining the contour boundary points of the text unit according to the central curve and the height of the text line; correcting the text unit according to the contour boundary points of the text unit and the preset contour boundary points to obtain the corrected text unit; stitching the corrected text units to obtain a second text image. The present application segments each text line in the first text image to be corrected to obtain a plurality of text units. Correcting each text unit in the text line one by one realizes the segmented surface correction of the first text image. Then stitching the corrected text units realizes the correction of the curved first text image into a flat second text image, improving the text image correction effect and the accuracy of subsequent OCR recognition.

[0059] In one embodiment, the step of extracting a preset number of sampling points from the text area contour in step S30 includes step S31, which is specifically as follows:

[0060] S31: Sample a first edge point and a corresponding second edge point at a preset distance interval along the horizontal direction from the text area contour to obtain a preset number of sampling points; wherein, the second edge point is the other intersection point of the straight line passing through the first edge point and the text area contour; the direction of the straight line is perpendicular to the horizontal direction.

[0061] Wherein, taking the upper left corner vertex of the first text image as the origin, a coordinate system is established, the X-axis direction of the coordinate system is the horizontal direction of the first text image, and the Y-axis direction of the coordinate system is the vertical direction of the first text image.

[0062] Among them, the preset distance can be set according to actual requirements.

[0063] In the embodiment of the present application, the first edge point can be a point on the upper side of the contour of the text area, and the corresponding second edge point can be a point on the lower side of the contour of the text area. The abscissa of the first edge point is the same as that of the corresponding second edge point. For example, the position coordinates of the first edge points are respectively (x1, t1), (x2, t2),..., (xn, tn), and the position coordinates of the corresponding second edge points are respectively (x1, b1), (x2, b2),..., (xn, bn).

[0064] The steps of obtaining the central curve of the text line and the height of the text line according to the sampling points in step S30 include steps S32 to S35, which are specifically as follows:

[0065] S32: Determine the midpoint between the first edge point and the corresponding second edge point.

[0066] In the embodiment of the present application, the position coordinates of the midpoint between the first edge point and the corresponding second edge point are respectively (x1, (t1 + b1) / 2), (x2, (t2 + b2) / 2),..., (xn, (tn + bn) / 2).

[0067] S33: Use a preset polynomial to perform curve fitting on the midpoints to obtain the central curve of the text line.

[0068] Among them, the preset polynomial can be set according to actual requirements. For example, the preset polynomial is f(x) = ax^3 + bx^2 + cx + d, where a, b, c, and d are coefficients to be fitted.

[0069] In the embodiment of the present application, using the least squares method, through the preset polynomial, perform curve fitting on the position coordinates of all midpoints in the text line, calculate the specific expression of the polynomial, and obtain the central curve of the text line.

[0070] S34: Determine the height difference between the first edge point and the corresponding second edge point.

[0071] In the embodiment of the present application, the height differences between the first edge point and the corresponding second edge point are respectively b1 - t1, b2 - t2,..., bn - tn.

[0072] S35: Calculate the average value of the height differences to obtain the height of the text line.

[0073] In the embodiment of the present application, sum up the height differences, divide the summation result by n, and obtain the height Hn of the text line.

[0074] In the embodiments of the present application, by performing curve fitting on the midpoints, the central curve of the text line can be automatically obtained.

[0075] In one embodiment, the step of segmenting the text line in step S40 to obtain a plurality of text units includes steps S401 to S402, which are specifically as follows:

[0076] S401: Obtain a plurality of equally divided points of the text line according to the preset number of divided segments and the width of the text line.

[0077] Among them, the preset number of divided segments can be set according to actual needs.

[0078] Among them, the width of each text line is the width of the first text image.

[0079] In the embodiments of the present application, taking the width of the text line as IW and the preset number of divided segments as N as an example for illustration, each text line includes N equally divided points, and the abscissas of each equally divided point are respectively denoted as x0, x1, x2,..., xn. Among them, xn = IW / N * n, n = 0, 1, 2,..., N. The ordinates of each equally divided point are the same, and the ordinate of each equally divided point is within the height range of the text line.

[0080] S402: Segment the text line according to the equally divided points to obtain a plurality of text units.

[0081] In the embodiments of the present application, the abscissas of the equally divided points included in the first text unit are x0 and x1 respectively, the abscissas of the equally divided points included in the second text unit are x1 and x2 respectively,..., and the abscissas of the equally divided points included in the Nth text unit are xn-1 and xn respectively.

[0082] The step of determining the contour boundary points of the text unit according to the central curve of the text line and the height of the text line in step S40 includes steps S403 to S405, which are specifically as follows:

[0083] S403: When the text line where the text unit is located is the first text line, determine the contour boundary points of the text unit according to the central curve of the text line and the height of the text line; where the first text line is the text line with the longest text length and the earliest text line order.

[0084] In the embodiments of the present application, the first text image includes a plurality of text lines, and the number of characters included in each text line is different. According to the number of characters, the text line with the longest text length can be determined. Considering that there may be multiple text lines with the longest text length, the text lines are sorted in order from top to bottom, and the text line with the earliest order and the longest text length is selected to obtain the first text line.

[0085] Traverse each text unit in each text line. If the text line where the current text unit is located is the first text line, it means that there is no blank area for the current text unit. The contour boundary points of the text unit can be determined according to the central curve of the text line where the current text unit is located and the height of the text line.

[0086] S404: When the text line where the text unit is located is not the first text line and the text unit is within the text area of the corresponding text line, determine the contour boundary points of the text unit according to the central curve of the text line and the height of the text line.

[0087] In the embodiment of the present application, if the text line where the current text unit is located is not the first text line, it means that there is a blank area in the text line where the current text unit is located. Since the current text unit is within the text area of the corresponding text line, it means that there is no blank area for the current text unit. The contour boundary points of the text unit can be determined according to the central curve of the text line where the current text unit is located and the height of the text line.

[0088] S405: When the text line where the text unit is located is not the first text line and the text unit is not within the text area of the corresponding text line, determine the contour boundary points of the text unit according to the central curve of the text line, the central curve of the first text line, and the height of the text line.

[0089] In the embodiment of the present application, if the text line where the current text unit is located is not the first text line, it means that there is a blank area in the text line where the current text unit is located. Since the current text unit is not within the text area of the corresponding text line, it means that there is a blank area for the current text unit. It is necessary to combine the central curve of the first text line, the central curve of the text line where the current text unit is located, and the height of the text line to determine the contour boundary points of the text unit.

[0090] In the embodiment of the present application, by using the central curve of the first text line to assist in determining the contour boundary points of the text unit with blank areas, the accuracy of the contour boundary points can be improved.

[0091] In one embodiment, the steps of determining the contour boundary points of the text unit according to the central curve of the text line and the height of the text line include steps S411 to S414, specifically as follows:

[0092] S411: Obtain the starting point and the ending point of the text unit; wherein, the starting point and the ending point are two equal division points corresponding to the text unit.

[0093] In the embodiment of the present application, taking the text unit as the first text unit in the text line as an example, the abscissa of the starting point of the text unit is x0, and the abscissa of the ending point of the text unit is x1.

[0094] S412: Use the abscissa of the starting point as the abscissa of the starting center point, and input the abscissa of the starting center point into the center curve of the text line to obtain the ordinate of the starting center point.

[0095] In the embodiment of the present application, taking the first text unit in the text line as an example, the abscissa of the starting center point of the text unit is x0, and the ordinate of the starting center point of the text unit is f(x0).

[0096] S413: Use the abscissa of the ending point as the abscissa of the ending center point, and input the abscissa of the ending center point into the center curve of the text line to obtain the ordinate of the ending center point.

[0097] In the embodiment of the present application, taking the first text unit in the text line as an example, the abscissa of the ending center point of the text unit is x1, and the ordinate of the ending center point of the text unit is f(x1).

[0098] S414: Obtain the contour boundary points of the text unit according to the ordinate of the starting center point, the ordinate of the ending center point, and the height of the text line.

[0099] In the embodiment of the present application, the upper left vertex and the lower left vertex of the contour corresponding to the text unit can be determined according to the ordinate of the starting center point and the height of the text line. The upper right vertex and the lower right vertex of the contour corresponding to the text unit can be determined according to the ordinate of the ending center point and the height of the text line.

[0100] In the embodiment of the present application, the contour of the text unit is determined according to the center curve of the text line and the height of the text line, so that the contour can envelope the text in the text unit.

[0101] In one embodiment, the steps of determining the contour boundary points of the text unit according to the center curve of the text line, the center curve of the first text line, and the height of the text line include steps S41 to S45, which are specifically as follows:

[0102] S41: Obtain a number of first equally divided points and a number of second equally divided points; among them, one first equally divided point corresponds to one second equally divided point. The first equally divided point is an equally divided point located in the text area of the text line, and the second equally divided point corresponding to the first equally divided point is an equally divided point located in the text area of the first text line and having the same abscissa as the first equally divided point.

[0103] In the embodiment of the present application, taking the abscissas of the equally divided points in the text area of the text line as x3, x4, x5 respectively, and the abscissas of the equally divided points in the text area of the first text line as x1, x2, x3, x4, x5, x6 as an example, the abscissas of the first equally divided points are x3, x4, x5 respectively, and the abscissas of the second equally divided points are also x3, x4, x5.

[0104] S42: Use the abscissa of each first equal division point as the abscissa of a first center point, input the abscissa of the first center point into the center curve of the text line to obtain the ordinates of each first center point; use the abscissa of each second equal division point as the abscissa of a second center point, input the abscissa of the second center point into the center curve of the first text line to obtain the ordinates of each second center point.

[0105] In the embodiment of the present application, taking the center curve of the text line as f1(x) and the center curve of the first text line as f2(x) as an example, the abscissas of each first center point are x3, x4, x5, and the ordinates of each first center point are f1(x3), f1(x4), f1(x5). The abscissas of each second center point are x3, x4, x5, and the ordinates of each second center point are f2(x3), f2(x4), f2(x5).

[0106] S43: Calculate the average value of the differences between the ordinates of each first center point and the corresponding ordinates of each second center point.

[0107] In the embodiment of the present application, the average value is: |[f2(x3) - f1(x3)] + [f2(x4) - f1(x4)] + [f2(x5) - f1(x5)]| / 3.

[0108] S44: Obtain the starting point and the ending point of the text unit; wherein, the starting point and the ending point are two equal division points corresponding to the text unit.

[0109] In the embodiment of the present application, taking the first text unit in the text line as an example, the abscissa of the starting point of the text unit is x0, and the abscissa of the ending point of the text unit is x1.

[0110] S45: Determine the contour boundary points of the text unit according to the starting point, the ending point and the average value of the text unit.

[0111] In the embodiment of the present application, determine the upper left vertex and the lower left vertex of the contour corresponding to the text unit according to the ordinate of the starting point and the average value. Determine the upper right vertex and the lower right vertex of the contour corresponding to the text unit according to the ordinate of the ending point and the average value.

[0112] In one embodiment, step S45 includes steps S451 to S453, specifically as follows:

[0113] S451: When the starting point of the text unit is not within the text area of the corresponding text line and the ending point of the text unit is within the text area of the corresponding text line, take the abscissa of the starting point as the abscissa of the starting center point, input the abscissa of the starting center point into the center curve of the text line to obtain the first ordinate; input the abscissa of the starting center point into the center curve of the first text line to obtain the second ordinate; obtain the ordinate of the starting center point according to the first ordinate, the second ordinate and the average value; take the abscissa of the ending point as the abscissa of the ending center point, input the abscissa of the ending center point into the center curve to obtain the ordinate of the ending center point; obtain the contour boundary points of the text unit according to the ordinate of the starting center point, the ordinate of the ending center point and the height of the text line.

[0114] In the embodiment of the present application, when the starting point of the text unit is not within the text area of the corresponding text line and the ending point of the text unit is within the text area of the corresponding text line, it indicates that the left side of the text unit is a blank area and there are characters arranged on the right side. Taking the abscissa of the starting point of the text unit as x0, the abscissa of the ending point of the text unit as x2, the center curve of the text line as f1(x), and the center curve of the first text line as f2(x) as an example, the first ordinate is f1(x0), the second ordinate is f2(x0), and the ordinate of the ending center point is f1(x2).

[0115] According to the second ordinate and the average value, the approximate position of the starting center point in the text line can be obtained. According to the first ordinate and the approximate position of the starting center point, the ordinate of the starting center point is obtained.

[0116] According to the ordinate of the starting center point and the height of the text line, determine the upper left vertex and the lower left vertex of the corresponding contour of the text unit. The upper right vertex and the lower right vertex of the corresponding contour of the text unit can be determined according to the ordinate of the ending center point and the height of the text line.

[0117] S452: When the ending point of the text unit is not within the text area of the corresponding text line and the starting point of the text unit is within the text area of the corresponding text line, take the abscissa of the ending point as the abscissa of the ending center point, input the abscissa of the ending center point into the center curve of the text line to obtain the third ordinate; input the abscissa of the ending center point into the center curve of the first text line to obtain the fourth ordinate; obtain the ordinate of the ending center point according to the third ordinate, the fourth ordinate and the average value; take the abscissa of the starting point as the abscissa of the starting center point, input the abscissa of the starting center point into the center curve to obtain the ordinate of the starting center point; obtain the contour boundary points of the text unit according to the ordinate of the starting center point, the ordinate of the ending center point and the height of the text line.

[0118] In an embodiment of the present application, when the end point of the text unit is not within the text area of the corresponding text line and the start point of the text unit is within the text area of the corresponding text line, it indicates that the right side of the text unit is a blank area and there are characters arranged on the left side. Taking the abscissa of the start point of the text unit as x0, the abscissa of the end point of the text unit as x2, the central curve of the text line as f1(x), and the central curve of the first text line as f2(x) as an example, the third ordinate is f1(x2), the fourth ordinate is f2(x2), and the ordinate of the start center point is f1(x0).

[0119] Based on the fourth ordinate and the average value, the approximate position of the end center point in the text line can be obtained. Based on the third ordinate and the approximate position of the end center point, the ordinate of the end center point is obtained.

[0120] Based on the ordinate of the start center point and the height of the text line, the upper left vertex and the lower left vertex of the corresponding contour of the text unit are determined. The upper right vertex and the lower right vertex of the corresponding contour of the text unit can be determined based on the ordinate of the end center point and the height of the text line.

[0121] S453: When the start point of the text unit is not within the text area of the corresponding text line and the end point of the text unit is not within the text area of the corresponding text line, take the abscissa of the start point as the abscissa of the start center point, input the abscissa of the start center point into the central curve of the text line to obtain the first ordinate; input the abscissa of the start center point into the central curve of the first text line to obtain the second ordinate; based on the first ordinate, the second ordinate and the average value, obtain the ordinate of the start center point; take the abscissa of the end point as the abscissa of the end center point, input the abscissa of the end center point into the central curve of the text line to obtain the third ordinate; input the abscissa of the end center point into the central curve of the first text line to obtain the fourth ordinate; based on the third ordinate, the fourth ordinate and the average value, obtain the ordinate of the end center point; based on the ordinate of the start center point, the ordinate of the end center point and the height of the text line, obtain the contour boundary points of the text unit.

[0122] In an embodiment of the present application, when the start point of the text unit is not within the text area of the corresponding text line and the end point of the text unit is not within the text area of the corresponding text line, it indicates that the entire text unit is a blank area without characters. Taking the abscissa of the start point of the text unit as x0, the abscissa of the end point of the text unit as x2, the central curve of the text line as f1(x), and the central curve of the first text line as f2(x) as an example, the first ordinate is f1(x0), the second ordinate is f2(x0), the third ordinate is f1(x2), and the fourth ordinate is f2(x2).

[0123] Based on the second ordinate and the average value, the approximate position of the starting center point in the text line can be obtained. Based on the first ordinate and the approximate position of the starting center point, the ordinate of the starting center point is obtained.

[0124] Based on the fourth ordinate and the average value, the approximate position of the ending center point in the text line can be obtained. Based on the third ordinate and the approximate position of the ending center point, the ordinate of the ending center point is obtained.

[0125] Based on the ordinate of the starting center point and the height of the text line, the upper left vertex and the lower left vertex of the contour corresponding to the text unit are determined. The upper right vertex and the lower right vertex of the contour corresponding to the text unit can be determined according to the ordinate of the ending center point and the height of the text line.

[0126] In one embodiment, the step of obtaining the ordinate of the starting center point according to the first ordinate, the second ordinate and the average value in step S451 includes steps S4511 to S4512, which are specifically as follows:

[0127] S4511: When the text line is above the first text line, subtract the average value from the second ordinate to obtain a difference value; perform a weighted sum of the first ordinate and the difference value to obtain the ordinate of the starting center point.

[0128] In the embodiment of the present application, taking the abscissa of the starting point of this unit as x0, the abscissa of the ending point of the text unit as x2, the center curve of the text line as f1(x), and the center curve of the first text line as f2(x) as an example, when the text line is above the first text line, the ordinate of the starting center point is: p*f1(x0)+q*[f2(x0)-z], where z is the average value, and p and q are weighting coefficients, and p + q = 1.

[0129] Similarly, the ordinate of the ending center point in the aforementioned step S452 is: p*f1(x2)+q*[f2(x2)-z].

[0130] S4512: When the text line is below the first text line, add the average value to the second ordinate to obtain a sum value; perform a weighted sum of the first ordinate and the sum value to obtain the ordinate of the starting center point.

[0131] In the embodiment of the present application, taking the abscissa of the starting point of this unit as x0, the abscissa of the ending point of the text unit as x2, the center curve of the text line as f1(x), and the center curve of the first text line as f2(x) as an example, when the text line is below the first text line, the ordinate F1 of the starting center point is: p*f1(x0)+q*[f2(x0)+z], where z is the average value, and p and q are weighting coefficients, and p + q = 1.

[0132] Similarly, the ordinate F2 of the termination center point in the aforementioned step S452 is: p*f1(x2) + q*[f2(x2) + z].

[0133] In one embodiment, the contour boundary points of the text unit include the upper left corner boundary point, the lower left corner boundary point, the upper right corner boundary point, and the lower right corner boundary point. The steps of obtaining the contour boundary points of the text unit according to the ordinate of the starting center point, the ordinate of the termination center point, and the height of the text line include steps S421 to S422, which are specifically as follows:

[0134] S421: Subtract half of the height of the text line from the ordinate of the starting center point to obtain the upper left corner boundary point; add half of the height of the text line to the ordinate of the starting center point to obtain the lower left corner boundary point.

[0135] In the embodiment of the present application, taking the abscissa of the starting center point as X1, the ordinate of the starting center point as F1, the abscissa of the termination center point as X2, the ordinate of the termination center point as F2, and the height of the text line as Hn as an example, the coordinates of the upper left corner boundary point are (X1, F1 - Hn / 2), and the coordinates of the lower left corner boundary point are (X1, F1 + Hn / 2).

[0136] S422: Take the abscissa of the termination point as the abscissa of the termination center point, input the abscissa of the termination center point into the center curve to obtain the ordinate of the termination center point; subtract half of the height of the text line from the ordinate of the termination center point to obtain the upper right corner boundary point; add half of the height of the text line to the ordinate of the termination center point to obtain the lower right corner boundary point.

[0137] In the embodiment of the present application, the coordinates of the upper right corner boundary point are (X2, F2 - Hn / 2), and the coordinates of the lower right corner boundary point are (X2, F2 + Hn / 2).

[0138] In one embodiment, step S50 includes steps S51 to S52, which are specifically as follows:

[0139] S51: Obtain the first affine transformation matrix according to the contour boundary points of the text unit and the preset contour boundary points.

[0140] In the embodiment of the present application, taking the contour boundary points of the text unit as (X1, F1 - Hn / 2), (X1, F1 + Hn / 2), (X2, F2 - Hn / 2), (X2, F2 + Hn / 2) as an example for illustration, the preset contour boundary points are (X1, 0), (X1, h), (X2, 0), (X2, h), where h is the preset height value.

[0141] By calculating the first affine transformation matrix, the contour boundary points of the text unit are transformed into corresponding preset contour boundary points. Specifically, (X1, F1 - Hn / 2) is transformed into (X1, 0), (X1, F1 + Hn / 2) is transformed into (X1, h), (X2, F2 - Hn / 2) is transformed into (X2, 0), and (X2, F2 + Hn / 2) is transformed into (X2, h).

[0142] S52: According to the first affine transformation matrix, correct the text unit to obtain the corrected text unit.

[0143] In the embodiment of the present application, after obtaining the first affine transformation matrix, correct the text unit so that the positions of the characters in the text unit are adjusted, presenting the effect of straight arrangement of the characters.

[0144] In one embodiment, step S60 includes steps S61 to S66, specifically as follows:

[0145] S61: splice the corrected text units to obtain each spliced text line.

[0146] In the embodiment of the present application, the corrected text units are spliced in the order of their positions in the text line to obtain the spliced text line, and the spliced text line is a straight text line.

[0147] S62: determine a number of second text lines; the second text line is the text line with the longest text length among the spliced text lines.

[0148] In the embodiment of the present application, the number of characters in each spliced text line can be traversed, and the text line with the most characters is determined as the second text line.

[0149] S63: obtain the starting position and the ending position of the text in the second text line.

[0150] In the embodiment of the present application, the starting position of the text in the second text line can be the position of the first character in the second text line, and the ending position of the text in the second text line can be the position of the last character in the second text line.

[0151] S64: perform linear fitting on the starting position to obtain a first straight line; perform linear fitting on the ending position to obtain a second straight line.

[0152] In the embodiment of the present application, by using the least squares method to perform linear fitting on the starting position and the ending position of the text in each second text line, a first straight line and a second straight line can be obtained.

[0153] S65: obtain a second affine transformation matrix according to the first straight line, the second straight line and a preset straight line.

[0154] Among them, the preset straight line is a straight line in the vertical direction. For example, the straight line equation of the preset straight line is: X = 1.

[0155] In the embodiment of the present application, taking the straight line equation of the first straight line as: A1x + B1y + C1 = 0, the straight line equation of the second straight line as: A2x + B2y + C2 = 0, and the straight line equation of the preset straight line as: X = 1 as an example, calculate the second affine transformation matrix so that the straight line equation A1x + B1y + C1 = 0 of the first straight line and the straight line equation A2x + B2y + C2 = 0 of the second straight line both become X = 1.

[0156] S66: According to the second affine transformation matrix, correct each text line after splicing to obtain a second text image.

[0157] In the embodiment of the present application, after obtaining the second affine transformation matrix, the image where each text line is located after splicing is corrected as a whole to keep the image straight in the vertical direction, and a second text image is obtained.

[0158] The following is an embodiment of the device of the present application, which can be used to execute the content of the method in the present application. For the details not disclosed in the embodiment of the device of the present application, please refer to the content of the method in the embodiment of the present application.

[0159] Please refer to Figure 2 , which shows a schematic structural diagram of a text image correction device provided in an embodiment of the present application. The text image correction device 7 provided in the embodiment of the present application includes:

[0160] A first text image acquisition module 71, configured to acquire a first text image to be corrected; the first text image includes a plurality of text lines;

[0161] A mask image acquisition module 72, configured to preprocess the first text image to obtain a mask image of the text line;

[0162] A central curve acquisition module 73, configured to determine the text region contour of the text line according to the mask image of the text line; extract a preset number of sampling points from the text region contour; and obtain the central curve of the text line and the height of the text line according to the sampling points;

[0163] A contour boundary point determination module 74, configured to segment the text line to obtain a plurality of text units; and determine the contour boundary points of the text units according to the central curve of the text line and the height of the text line;

[0164] A text unit correction module 75, configured to correct the text unit according to the contour boundary points of the text unit and the preset contour boundary points to obtain a corrected text unit;

[0165] The second text image obtaining module 76 is configured to splice the corrected text units to obtain a second text image.

[0166] Applying the embodiments of the present application, by obtaining a first text image to be corrected; the first text image includes a plurality of text lines; preprocessing the first text image to obtain a mask image of the text lines; determining the text region contour of the text lines according to the mask image of the text lines; extracting a preset number of sampling points from the text region contour; obtaining the center curve of the text lines and the height of the text lines according to the sampling points; segmenting the text lines to obtain a plurality of text units; determining the contour boundary points of the text units according to the center curve of the text lines and the height of the text lines; correcting the text units according to the contour boundary points of the text units and the preset contour boundary points to obtain corrected text units; splicing the corrected text units to obtain a second text image. The present application segments each text line in the first text image to be corrected to obtain a plurality of text units. Correcting each text unit in the text line realizes the segmented surface correction of the first text image. Then splicing the corrected text units realizes the correction of the curved first text image into a flat second text image, improves the text image correction effect, and improves the accuracy of subsequent OCR recognition.

[0167] The following is an embodiment of the device of the present application, which can be used to execute the content of the method in the embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the content of the method embodiments of the present application.

[0168] Please refer to Figure 3 , the present application further provides an electronic device 300. The electronic device can specifically be a computer, a mobile phone, a tablet computer, a text image correction device, etc. In an exemplary embodiment of the present application, the electronic device 300 is a text image correction device. The text image correction device may include: at least one processor 301, at least one memory 302, at least one display, at least one network interface 303, a user interface 304, and at least one communication bus 305.

[0169] Among them, the user interface 304 is mainly used to provide an input interface for the user to obtain data input by the user. Optionally, the user interface may further include a standard wired interface and a wireless interface.

[0170] Among them, the network interface 303 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0171] Among them, the communication bus 305 is used to implement connection communication between these components.

[0172] Among them, the processor 301 may include one or more processing cores. The processor connects various parts within the entire electronic device through various interfaces and lines, and executes various functions of the electronic device and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory, and by calling the data stored in the memory. Optionally, the processor may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed in the display layer; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor and may be implemented separately by a single chip.

[0173] Among them, the memory 302 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, code, code sets, or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store the data involved in the above-mentioned method embodiments. Optionally, the memory may also be at least one storage device located far from the aforementioned processor. As Figure 3 shown, the memory as a computer storage medium may include an operating system, a network communication module, a user interface module, and an operation application program.

[0174] The processor can be used to call the application program of the text image correction method stored in the memory and specifically execute the method steps of the above-mentioned embodiments. The specific execution process can refer to the specific description shown in the embodiments and will not be elaborated here.

[0175] The present application also provides a computer-readable storage medium, on which a computer program is stored. The instructions are suitable for being loaded and executed by a processor to perform the method steps of the above-described embodiments. For the specific execution process, reference may be made to the specific descriptions shown in the embodiments, which will not be elaborated herein. The device where the storage medium is located may be an electronic device such as a personal computer, a laptop computer, a smart phone, a tablet computer, etc.

[0176] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can refer to the partial descriptions of the method embodiments. The device embodiments described above are only illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present application. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0177] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0178] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the selected functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the selected functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0179] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for implementing selected functions in a process Figure 1 one process or multiple processes and / or boxes Figure 1 steps of boxes or multiple boxes.

[0180] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0181] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0182] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical memory, magnetic tape, magnetic disk or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0183] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0184] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A text image correction method, characterized in that: The steps include: Acquire a first text image to be corrected; the first text image includes a plurality of text lines; Preprocessing the first text image to obtain a mask image of the text line; Determine a text area contour of the text line according to the mask image of the text line; extract a preset number of sampling points from the text area contour; obtain a center curve of the text line and a height of the text line according to the sampling points; Segmenting the text line to obtain a plurality of text units; determining the contour boundary points of the text unit according to the center curve of the text line and the height of the text line; Correcting the text unit according to the contour boundary points of the text unit and the preset contour boundary points to obtain a corrected text unit; The corrected text units are spliced ​​to obtain a second text image.

2. The text image correction method according to claim 1, characterized in that: The step of extracting a preset number of sampling points from the text area contour comprises: Sampling a first edge point and a corresponding second edge point from the text area contour at every preset interval along the horizontal direction to obtain the preset number of sampling points; wherein the second edge point is another intersection point of a straight line passing through the first edge point and the text area contour; the direction of the straight line is perpendicular to the horizontal direction; The step of obtaining the center curve of the text line and the height of the text line according to the sampling points comprises: Determine a midpoint between the first edge point and the corresponding second edge point; Using a preset polynomial, curve fitting is performed on the midpoint to obtain a central curve of the text line; determining a height difference between the first edge point and the corresponding second edge point; The average value of the height differences is calculated to obtain the height of the text line.

3. The text image correction method according to claim 1, characterized in that: The step of segmenting the text line to obtain a plurality of text units includes: According to the preset number of segment divisions and the width of the text line, a plurality of equally divided points of the text line are obtained; Segmenting the text line according to the equal division points to obtain a plurality of text units; The step of determining the outline boundary points of the text unit according to the center curve of the text line and the height of the text line comprises: When the text line where the text unit is located is the first text line, determining the outline boundary point of the text unit according to the central curve of the text line and the height of the text line; wherein the first text line is the text line with the longest text length and the frontmost text line sequence; When the text line where the text unit is located is not the first text line, and the text unit is within the text area of ​​the corresponding text line, determining the outline boundary points of the text unit according to the center curve of the text line and the height of the text line; When the text line where the text unit is located is not the first text line, and the text unit is not in the text area of ​​the corresponding text line, the contour boundary points of the text unit are determined based on the center curve of the text line, the center curve of the first text line and the height of the text line.

4. The text image correction method according to claim 3, characterized in that: The step of determining the outline boundary points of the text unit according to the center curve of the text line and the height of the text line comprises: Obtaining a starting point and an ending point of the text unit; wherein the starting point and the ending point are two equally divided points corresponding to the text unit; Using the abscissa of the starting point as the abscissa of the starting center point, inputting the abscissa of the starting center point into the center curve of the text line, and obtaining the ordinate of the starting center point; Using the abscissa of the end point as the abscissa of the end center point, inputting the abscissa of the end center point into the center curve of the text line, and obtaining the ordinate of the end center point; The contour boundary points of the text unit are obtained according to the ordinate of the starting center point, the ordinate of the ending center point, and the height of the text line.

5. The text image correction method according to claim 3, characterized in that: The step of determining the outline boundary points of the text unit according to the center curve of the text line, the center curve of the first text line, and the height of the text line comprises: Acquire a plurality of first equally divided points and a plurality of second equally divided points; wherein one first equally divided point corresponds to one second equally divided point, the first equally divided point is an equally divided point located in the text area of ​​the text line, and the second equally divided point corresponding to the first equally divided point is an equally divided point located in the text area of ​​the first text line and having the same horizontal coordinate as the first equally divided point; Taking the abscissa of each of the first equally divided points as the abscissa of a first center point, inputting the abscissa of the first center point into the center curve of the text line, and obtaining the ordinate of each of the first center points; taking the abscissa of each of the second equally divided points as the abscissa of a second center point, inputting the abscissa of the second center point into the center curve of the first text line, and obtaining the ordinate of each of the second center points; Calculate the average value of the difference between the ordinate of each of the first center points and the ordinate of each of the corresponding second center points; Obtaining a starting point and an ending point of the text unit; wherein the starting point and the ending point are two equally divided points corresponding to the text unit; The contour boundary points of the text unit are determined according to the starting point, the ending point and the average value of the text unit.

6. The text image correction method according to claim 5, characterized in that: The step of determining the contour boundary points of the text unit according to the starting point, the ending point and the average value of the text unit comprises: When the starting point of the text unit is not in the text area of ​​the corresponding text line, and the ending point of the text unit is in the text area of ​​the corresponding text line, the horizontal coordinate of the starting point is used as the horizontal coordinate of the starting center point, and the horizontal coordinate of the starting center point is input into the center curve of the text line to obtain a first vertical coordinate; the horizontal coordinate of the starting center point is input into the center curve of the first text line to obtain a second vertical coordinate; the vertical coordinate of the starting center point is obtained according to the first vertical coordinate, the second vertical coordinate and the average value; the horizontal coordinate of the ending point is used as the horizontal coordinate of the ending center point, and the horizontal coordinate of the ending center point is input into the center curve to obtain the vertical coordinate of the ending center point; the contour boundary point of the text unit is obtained according to the vertical coordinate of the starting center point, the vertical coordinate of the ending center point and the height of the text line; When the end point of the text unit is not in the text area of ​​the corresponding text line, and the starting point of the text unit is in the text area of ​​the corresponding text line, the horizontal coordinate of the end point is used as the horizontal coordinate of the end center point, and the horizontal coordinate of the end center point is input into the center curve of the text line to obtain a third vertical coordinate; the horizontal coordinate of the end center point is input into the center curve of the first text line to obtain a fourth vertical coordinate; the vertical coordinate of the end center point is obtained according to the third vertical coordinate, the fourth vertical coordinate and the average value; the horizontal coordinate of the starting point is used as the horizontal coordinate of the starting center point, and the horizontal coordinate of the starting center point is input into the center curve to obtain the vertical coordinate of the starting center point; the contour boundary point of the text unit is obtained according to the vertical coordinate of the starting center point, the vertical coordinate of the end center point and the height of the text line; When the starting point of the text unit is not in the text area of ​​the corresponding text line, and the ending point of the text unit is not in the text area of ​​the corresponding text line, the horizontal coordinate of the starting point is used as the horizontal coordinate of the starting center point, and the horizontal coordinate of the starting center point is input into the center curve of the text line to obtain a first vertical coordinate; the horizontal coordinate of the starting center point is input into the center curve of the first text line to obtain a second vertical coordinate; the vertical coordinate of the starting center point is obtained according to the first vertical coordinate, the second vertical coordinate and the average value; the horizontal coordinate of the ending point is used as the horizontal coordinate of the ending center point, and the horizontal coordinate of the ending center point is input into the center curve of the text line to obtain a third vertical coordinate; the horizontal coordinate of the ending center point is input into the center curve of the first text line to obtain a fourth vertical coordinate; the vertical coordinate of the ending center point is obtained according to the third vertical coordinate, the fourth vertical coordinate and the average value; the contour boundary point of the text unit is obtained according to the vertical coordinate of the starting center point, the vertical coordinate of the ending center point and the height of the text line.

7. The text image correction method according to claim 6, characterized in that: The step of obtaining the ordinate of the starting center point according to the first ordinate, the second ordinate and the average value comprises: When the text line is located above the first text line, subtract the second ordinate from the average value to obtain a difference; perform weighted summation on the first ordinate and the difference to obtain the ordinate of the starting center point; When the text line is located below the first text line, the second ordinate is added to the average value to obtain a sum value; and the first ordinate and the sum value are weightedly summed to obtain the ordinate of the starting center point.

8. The text image correction method according to any one of claims 4 to 7, characterized in that: The outline boundary points of the text unit include an upper left corner boundary point, a lower left corner boundary point, an upper right corner boundary point and a lower right corner boundary point; The step of obtaining the outline boundary point of the text unit according to the ordinate of the starting center point, the ordinate of the ending center point and the height of the text line comprises: Subtract the ordinate of the starting center point from half the height of the text line to obtain the upper left corner boundary point; add the ordinate of the starting center point to half the height of the text line to obtain the lower left corner boundary point; The horizontal coordinate of the termination point is used as the horizontal coordinate of the termination center point, and the horizontal coordinate of the termination center point is input into the center curve to obtain the vertical coordinate of the termination center point; the vertical coordinate of the termination center point is subtracted from half the height of the text line to obtain the upper right corner boundary point; the vertical coordinate of the termination center point is added to half the height of the text line to obtain the lower right corner boundary point.

9. The text image correction method according to claim 1, characterized in that: The step of correcting the text unit according to the contour boundary points of the text unit and the preset contour boundary points to obtain the corrected text unit includes: Obtaining a first affine transformation matrix according to the outline boundary points of the text unit and preset outline boundary points; The text unit is corrected according to the first affine transformation matrix to obtain a corrected text unit.

10. The text image correction method according to claim 1, characterized in that: The step of splicing the corrected text units to obtain a second text image includes: splicing the corrected text units to obtain spliced ​​text lines; Determine a plurality of second text lines; the second text line is the text line with the longest text length among the spliced ​​text lines; Obtaining the starting point and ending point of the text in the second text line; Performing straight line fitting on the starting point position to obtain a first straight line; performing straight line fitting on the end point position to obtain a second straight line; Obtaining a second affine transformation matrix according to the first straight line, the second straight line and a preset straight line; According to the second affine transformation matrix, each of the spliced ​​text lines is corrected to obtain a second text image.

11. A text image correction device, characterized in that: include: A first text image acquisition module, used to acquire a first text image to be corrected; the first text image includes a plurality of text lines; A mask image obtaining module, used for preprocessing the first text image to obtain a mask image of the text line; A center curve obtaining module, used to determine the text area contour of the text line according to the mask image of the text line; extract a preset number of sampling points from the text area contour; and obtain the center curve of the text line and the height of the text line according to the sampling points; A contour boundary point determination module, used to segment the text line to obtain a plurality of text units; determine the contour boundary points of the text unit according to the center curve of the text line and the height of the text line; A text unit correction module, used to correct the text unit according to the contour boundary points of the text unit and the preset contour boundary points to obtain a corrected text unit; The second text image obtaining module is used to splice the corrected text units to obtain a second text image.

12. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the text image correction method as described in any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the text image correction method as described in any one of claims 1 to 10 is implemented.