A text recognition method, device, and storage medium based on deep learning
Through the text recognition method based on deep learning, the feature extraction and Bezier curve correction technology of MobileNeXt network and PAN/PSP modules are used to solve the problems of curve text adhesion and slow speed, and high-precision text recognition is achieved, which is suitable for document electronicization and vehicle tire number detection.
Patent Information
- Application Number
- CN202111244912.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-25
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-10-25
Smart Images

Figure CN113971809B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a text recognition method, device and storage medium based on deep learning, belonging to the technical field of text recognition. Background Art
[0002] With the rapid development of the global economy, the industrial community has paid increasing attention to multi-scene, multi-language, and high-precision text detection and recognition. The needs in aspects such as scene understanding, product recognition, autonomous driving, target geographical positioning, and document digitization are also becoming more and more urgent. In recent years, with the continuous development of AI technology, there have been more and more problems in text detection and recognition. Therefore, the industrial community and academia have been exploring text detection and recognition more and more deeply.
[0003] Existing methods can generally be divided into four categories: detectors based on quadrilateral bounding boxes, character-based methods, segmentation-based methods, and parameter-structured methods. Among them, most existing detectors based on quadrilateral bounding boxes are difficult to locate text of arbitrary shapes and are difficult to be well enclosed in rectangles; most segmentation-based methods may not separate text instances that are very close to each other; character-based and parameter-structured methods require costly annotation information.
[0004] In industrial applications, segmentation-based methods are very popular in scene text detection because the segmentation results can more accurately describe scene text of various shapes, such as curved text; at the same time, they can achieve a good balance in terms of speed, accuracy, and annotation cost. Currently, the general algorithms based on segmentation are PseNet and DBNet, but both have their own disadvantages; the post-processing of PseNet is time-consuming; DBnet achieves a good balance between speed and accuracy, but there are often problems of adhesion of adjacent text, and at the same time, there is also a problem that the detected curved text seriously reduces the text recognition accuracy.
[0005] In view of the above difficult and urgent problems to be solved, the present invention proposes a text recognition method based on deep learning. Summary of the Invention
[0006] In order to overcome the deficiencies in the prior art, the present invention provides a text recognition method, device and storage medium based on deep learning to solve the problems of text adhesion, slow inference speed, and poor correction effect of curved text in the existing text detection and recognition process. The method of the present invention can be widely applied in document digitization and vehicle tire number detection.
[0007] The present invention specifically adopts the following technical solutions to solve the above technical problems:
[0008] A text recognition method based on deep learning, comprising the following steps:
[0009] Step 1: Produce a data set in a specified format;
[0010] Step 2: Construct a text detection network model and a loss function;
[0011] Step 3: Use the produced data set to train the constructed text detection network model and loss function to obtain a trained text detection network model;
[0012] Step 4: Obtain a picture of a certain scene;
[0013] Step 5: Use an open-source image processing operation library to perform fixed-size scaling and normalization processing on the obtained picture;
[0014] Step 6: Use the trained text detection network model to perform inference and prediction on the picture processed in Step 5, and extract the text area in the picture;
[0015] Step 7: Use a Bessel curve to correct the text area in the picture extracted in Step 6 to obtain a corrected text area;
[0016] Step 8: Preprocess the picture of the corrected text area, and then use the CRNN text recognition algorithm to extract text information from the text area in the preprocessed picture.
[0017] Further, as a preferred technical solution of the present invention, the Step 1 of producing a data set in a specified format specifically includes:
[0018] Step 1-1: Collect picture data of a certain scene;
[0019] Step 1-2: Perform data annotation on the collected picture data, respectively mark the four vertices of each text box in the picture, and the four vertices are in clockwise order, and each picture obtains one or more marked text boxes;
[0020] Step 1-3: Make a data set according to the text boxes of the obtained pictures in the data format of PASCAL VOC;
[0021] Further, as a preferred technical solution of the present invention, the Step 2 of constructing a text detection network model based on the MobileNeXt network specifically includes:
[0022] Input a picture, use the MobileNeXt network to extract features from the picture. During the feature extraction process, five downsamplings are performed. Each downsampling outputs a feature map of a certain scale. The width and height of each feature map are 1 / 2 of the width and height of the previous layer's feature map. The feature map of the last layer is 1 / 32 of the original picture;
[0023] The feature map obtained by upsampling the feature map formed by operating the last layer feature map output by the MobileNeXt network through the pyramid scene parsing module is merged with the fourth layer feature map to obtain the merged feature map, and so on for merging until the size of the merged feature map is 1 / 4 of the original image; then, the feature map with a size of 1 / 4 of the original image obtained by merging is downsampled three times, and the feature map of each layer is saved respectively. Then, the pyramid scene parsing module is used to aggregate the last layer feature map of feature extraction, and finally, the feature maps of each layer are merged respectively to output a feature map with a quantity of 6 and a size of 1 / 4 of the original image.
[0024] Furthermore, as a preferred technical solution of the present invention, the loss function constructed in step 2 is specifically:
[0025]
[0026] where D is the calculation function of the dice coefficient; S i is the set of the i-th predicted region, G i is the set of the i-th ground truth region, S i,x,y is the value of the pixel point (x, y) in the i-th predicted region, G i,x,y is the value of the pixel point (x, y) in the i-th ground truth region;
[0027] And, define L c as the text region classification loss, L s as the shrunk text region loss, and the calculation method is as follows:
[0028] L c = 1 - D(S n * M, G n * M)
[0029]
[0030] where M is the mask of the ground truth region during the training process, S n is the set of pixel points in the predicted region, G n is the set of pixel points in the ground truth region; W is the mask of a single text region in S n , and S n,x,y represents the pixel value of (x, y) in S n .
[0031] Furthermore, as a preferred technical solution of the present invention, in step 3, the constructed text detection network model and loss function are trained using the stochastic gradient descent algorithm.
[0032] Further, as a preferred technical solution of the present invention, in step 5, an open-source image processing library is used to perform fixed-size scaling and normalization processing on the acquired pictures, which specifically includes:
[0033] Step 5-1: Perform size scaling processing on the acquired pictures according to the set picture width and height;
[0034] Step 5-2: Perform normalization processing on the scaled pictures: Use an image operation library to read the pictures to make the pictures into operable arrays, then divide each number in the array by 255 for normalization operation, subtract a fixed mean value from each channel of the pictures and divide by a fixed variance; and, if the pictures are read by an open-source image processing library, adjust the channel order of the pictures to make it into the rgb channel order.
[0035] Further, as a preferred technical solution of the present invention, in step 7, a Bezier curve is used to correct the text area in the extracted pictures, which specifically includes:
[0036] Step 7-1: Obtain the upper and lower boundaries of the text area: For curved text, use a text detection network model to detect the curved text area in the pictures and calculate the circumscribed rectangle of the curved text area to obtain the circumscribed rectangle; calculate the angle of the circumscribed rectangle, and calculate the starting points of the upper and lower boundaries of the curved text area according to the border of the circumscribed rectangle;
[0037] Step 7-2: Take 8 points from each of the upper and lower boundaries according to requirements;
[0038] Step 7-3: Use the taken upper and lower boundary points to fit two Bezier curves for the upper and lower boundaries;
[0039] Step 7-4: Use the two Bezier curves for the upper and lower boundaries obtained by fitting to correct the text area in the pictures to obtain the corrected text area.
[0040] Further, as a preferred technical solution of the present invention, in step 8, preprocessing is performed on the pictures of the corrected text area, which specifically includes: Use an open-source image processing library to grayscale the pictures of the corrected text area, and then scale the pictures.
[0041] The present invention also provides an electronic device, including: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the steps in the text recognition method by executing the computer instructions.
[0042] The present invention also provides a computer-readable storage medium storing computer instructions for causing a computer to execute the steps in the text recognition method.
[0043] The present invention adopts the above technical solutions and can achieve the following technical effects:
[0044] In the text recognition method based on deep learning of the present invention, through the redesign of the network model structure, loss function and post-processing, the overall architecture of the network uses the lightweight MobileNeXt network as the backbone network for feature extraction, which can accelerate the inference speed of the network without sacrificing accuracy. At the same time, the neck of the network adopts the PAN structure (Pixel Aggregation Network), and the aggregation of the features of the last layer of the backbone network and the context information interaction use the Pyramid Scene Parsing Module (PSP Moudle), so that while the text detection algorithm achieves high accuracy and speed, it can further reduce the problem of text adhesion and correct curved text.
[0045] Moreover, for the device and storage medium proposed according to the text recognition method of the present invention, the processor executes the steps in the text recognition method by executing the computer instructions, and the computer-readable storage medium stores the computer instructions, so that the device and storage medium have the text recognition function. Therefore, the method of the present invention can effectively solve the problem of text adhesion, accurately correct curved text, effectively improve the text recognition accuracy, and can be widely applied in document digitization and vehicle tire number detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a schematic structural diagram of the text detection network model constructed in the method of the present invention.
[0047] Figure 2 It is a schematic diagram of correction using a Bessel curve in the present invention.
[0048] Figure 3 It is a schematic diagram of the picture input in the embodiment of the present invention.
[0049] Figure 4 It is a schematic diagram of the picture after adopting the method in the embodiment of the present invention.
[0050] Figure 5 It is a schematic diagram of the curved text in the embodiment of the present invention.
[0051] Figure 6 It is a schematic diagram of the text extracted after adopting the method in the embodiment of the present invention.
[0052] Figure 7 This is the schematic diagram of the method of the present invention. Detailed implementation manners
[0053] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with embodiments and the accompanying drawings. The content mentioned in the implementation manners does not limit the present invention.
[0054] Refer to Figure 7 As shown, the present invention relates to a text recognition method based on deep learning, and the method specifically includes the following steps:
[0055] Step 1: Make a data set in a specified format, specifically as follows:
[0056] Step 1-1: Collect picture data of a certain scene.
[0057] Step 1-2: Perform data annotation on the collected picture data, and respectively mark the four vertices of each text box in the picture. The four vertices are in clockwise order, and one or more marked text boxes can be obtained for each picture.
[0058] Step 1-3: Make a data set according to the text boxes of the obtained pictures in the data format of PASCAL VOC.
[0059] Step 2: Construct a text detection network model and a loss function based on the MobileNeXt network, specifically as follows:
[0060] In this method, the overall architecture of the text detection network model uses the lightweight MobileNeXt network as the backbone network for feature extraction, which can accelerate the inference speed of the network without losing accuracy. At the same time, the neck of the network adopts the PAN (Pixel Aggregation Network) structure. The aggregation of the features of the last layer of the backbone network and the context information interaction use the Pyramid Scene Parsing Module (PSP Moudle). Finally, the head of the network uses two-dimensional convolution operations to output 6 branches: S1, S2, ··· S6. S1 is the smallest segmentation result, and S6 is the largest segmentation result.
[0061] Step 2-2: As Figure 1 shown, it is the network architecture of the text detection network model, and its specific construction process is:
[0062] Input an image, and use the MobileNeXt network to extract features from the image. During the feature extraction process, five downsamplings are performed. Each downsampling outputs a feature map of a certain scale. The width and height of each feature map are 1 / 2 of the width and height of the previous layer's feature map. The feature map of the last layer is 1 / 32 of the original image;
[0063] The feature map obtained after the operation of the pyramid scene parsing module PSP Moudle on the last layer feature map output by the MobileNeXt network is upsampled and then merged with the fourth layer feature map to obtain a merged feature map. And so on for merging until the size of the merged feature map is 1 / 4 of the original image; then, three downsamplings are performed on the feature map of 1 / 4 of the original image obtained at this time, and the feature maps of each layer are saved respectively. The pyramid scene parsing module PSP Moudle is used to aggregate the last layer feature map of feature extraction, and finally the feature maps of each layer are merged respectively to output 6 feature maps with the size of 1 / 4 of the original image.
[0064] Step 2-3, design a loss function: The loss function is calculated using the metric function dicecoefficient that evaluates the similarity between two samples. After adjusting the dice coefficient loss function, the phenomenon of text adhesion is reduced; the specific improvement is as follows:
[0065] Make the numerator of the loss function the intersection of the predicted region set and the ground truth region set minus the part in the intersection that does not belong to the ground truth region, where the predicted region is defined as the result of model inference, and the ground truth region is the label region, and the denominator becomes the pixel set of the ground truth region; the new loss function is as follows:
[0066]
[0067] where, S i is the set of the i-th predicted region, G i is the set of the i-th ground truth region, S i,x,y is the value of the pixel point (x, y) in the i-th predicted region, G i,x,y is the value of the pixel point (x, y) in the i-th ground truth region;
[0068] And, define L c as the text region classification loss, L s as the shrinking text region loss, and the calculation method is as follows:
[0069] L c = 1 - D(S n * M, G n * M) (2)
[0070]
[0071] Among them, D is the calculation expression of the dice coefficient, M is the mask of the real region during the training process, S n is the set of pixel points in the predicted region, G n is the set of pixel points in the real region; (Considering that the shrunk text region is surrounded by the original text region, the pixels in the non-text region of the segmentation result S n are ignored to avoid pixel redundancy; ) W is the mask of a single text region in S n and S n,x,y represents S n the pixel value of (x, y) in it.
[0072] Step 3: Use the made dataset to train the constructed text detection network model and loss function to obtain the trained text detection network model, specifically as follows:
[0073] Step 3-1: Data augmentation: Randomly scale the input image to scales {0.5, 1.0, 2.0, 3.0}, and perform horizontal mirroring and random rotation between [-10 degrees, 10 degrees]. Randomly crop a 640*640-sized image from the transformed image, normalize the image using the color mean and variance. For the quadrilateral text dataset, use the minimum bounding box as the final prediction result of the bounding box; for the curved text dataset, use the Ramer-Douglas-Peucker algorithm to generate the bounding box for text regions of arbitrary shapes.
[0074] Step 3-2: Parameter tuning and iterative training, output the optimal model: Use the collected and preprocessed dataset as the training data to train the model. The optimization method during training is the stochastic gradient descent algorithm, and optimize and train the constructed text detection network model and loss function; set the batch size to 16-64, and train for 100-300 epochs. Set the initial learning rate to 10e-3, and decrease it by 1 / 10 at 100 and 200 epochs respectively; set the weight decay rate to 5*10e-4, and set the momentum to 0.99. Keep the model with the highest accuracy at the end as the optimal model.
[0075] Step 4: Obtain the picture of a certain scene;
[0076] Step 5: Use the open-source image operation library to perform fixed-size scaling and normalization processing on the obtained picture, specifically as follows:
[0077] Step 5-1: Resize the image according to the set width and height requirements of the image: Judge the input image. When the longest side is greater than 640, scale the shorter side at the ratio of processing 640 for the longest side, while maintaining the ratio of the original image.
[0078] Step 5-2: Normalize the resized image: Use an open-source image processing library to read the image, convert the image into an operable array, and then divide each number in the array by 255 for normalization. Subtract the fixed mean (0.485, 0.456, 0.406) and divide by the fixed variance (0.229, 0.224, 0.225) for each channel of the image; in addition, if the image is read by the opencv open-source image processing library, the channel order needs to be adjusted to make it the rgb channel order.
[0079] Step 6: Use the trained text detection network model to perform inference and prediction on the image processed in Step 5, and extract the text regions in the image.
[0080] Step 7: Use Bezier curves to correct the text regions in the image extracted in Step 6 to obtain the corrected text regions, specifically as follows:
[0081] Step 7-1: Obtain the upper and lower boundaries of the text region: For curved text, use the text detection network model to detect the curved text region in the image and calculate the circumscribed rectangle of the curved text region to obtain the circumscribed rectangle; calculate the angle of the circumscribed rectangle, calculate the angle in the counterclockwise direction of the long side, and calculate the starting points of the upper and lower boundaries of the curved text region according to the border of the circumscribed rectangle; among them, quadrilateral text is a special form of curved text, and use the minimum bounding box as its circumscribed rectangle.
[0082] Step 7-2: Take 8 points for each of the upper and lower boundaries according to requirements;
[0083] Step 7-3: Use the upper and lower boundary points to fit two Bezier curves for the upper and lower boundaries.
[0084] Step 7-4: Use the two Bezier curves for the upper and lower boundaries obtained by fitting to correct the text region in the image to obtain the corrected text region.
[0085] The fitted Bezier curve can be described by a series of pivot points bi and the following parametric equations about t:
[0086]
[0087] Among them, n' is the order of the Bessel curve. Since the index of the fulcrum bi starts from 0, the number of fulcrums = n + 1. Among them, c(t) represents the value of the curve at time t. The evolution of the parameter t from 0 to 1 forms the entire curve. For any point c(t) on the curve, the coordinates of this point can be regarded as the weighted average of the coordinates of all fulcrums, and the weights are Bi in the above equation. The specific operation process includes the following steps, as Figure 2 shown below:
[0088] (1) For any grid point in the recognition window, such as Figure 2 a point in the right square box in , first calculate the ratio t of the distance from it to the left side of the window to the width of the entire window;
[0089] (2) For Figure 2 the original target box with a left bend in , find the positions of the parameters corresponding to the Bessel curve parameter equations of its upper and lower boundaries with parameter value t, that is, tp and bp, such as Figure 2 the upper and lower hollow points in the original target box with a left bend in ; Figure 2 The solid point in the original target box with a left bend in corresponds to the solid point in the right square box, where w out and hout are the width and height of the corresponding output horizontal shape in the right square box respectively, and g iw and g ih are the width and height of the solid point in the right square box respectively, and op is the coordinate of the left solid point.
[0090] (3) Calculate Figure 2 the ratio of the distance from the grid point in the right square box in to the bottom of the window to the height of the entire window;
[0091] (4) Divide the line segment from bp to tp according to the ratio obtained in step (3) above to obtain the final corresponding point. After obtaining the corresponding point, the eigenvalue at this place can be solved by two-dimensional interpolation.
[0092]
[0093] Step 8: Preprocess the image of the corrected text area, and then use the CRNN text recognition algorithm to extract the text information from the text area in the preprocessed image, specifically as follows:
[0094] Step 8-1: Preprocess the corrected image; use an open-source image operation library to grayscale the image of the corrected text area, and then scale the grayscaled image. Scale the height of the image to 32, and scale the width of the image proportionally according to the ratio of the height scaled to 32. The maximum width is 1024. If it exceeds the maximum value, perform an interception operation; when performing batch recognition, it is necessary to perform a complement operation on images with a width less than 1024, and the complement value is equal to 0.
[0095] Step 8-2: Use the CRNN text recognition algorithm to recognize the text area in the preprocessed picture, and extract the text information corresponding to the text area.
[0096] In order to verify that the method of the present invention can achieve high precision and fast correction, and can further reduce the problem of text adhesion and correct curved text, an embodiment is listed for illustration.
[0097] As Figure 3 shown, it is a schematic diagram of a certain picture input in the embodiment of the present invention, and there is a problem of text adhesion in the picture text detection. As Figure 4 shown, it is a schematic diagram after adopting the method of the present invention. By comparison, it can be seen that the method of the present invention solves the problem of text adhesion in the picture.
[0098] And, as Figure 5 shown, it is a schematic diagram of the curved text in the input picture of the present invention. As Figure 6 shown, it is a schematic diagram of the extracted horizontal text after adopting the method of the present invention. By comparison, it can be seen that the present invention can effectively correct the curved text quickly and can quickly extract the horizontal text.
[0099] According to the above text recognition method based on deep learning, the present invention also provides an electronic device, including: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the steps in the above text recognition method, so that the electronic device has the text recognition function.
[0100] And, the present invention also provides a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and the computer instructions are used to make the computer execute the steps in the text recognition method, so that the computer-readable storage medium stores the text recognition method.
[0101] Therefore, the method, device and storage medium of the present invention can make the text detection algorithm achieve high precision and fast speed, and can further reduce the problem of text adhesion and correct curved text. It can effectively solve the problem of text adhesion, can accurately correct the curved text, effectively improve the text recognition accuracy, and can be widely applied in document digitization and vehicle tire number detection.
[0102] Although the above embodiments of the present invention have been described, the present invention is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Under the inspiration of this specification, those of ordinary skill in the art can also make many forms without departing from the scope protected by the claims of the present invention, and these all fall within the scope of protection of the present invention.
Claims
1. A text recognition method based on deep learning, characterized in that, Including the following steps: Step 1: Produce a data set in accordance with the specified format; Step 2: Construct a text detection network model based on the MobileNeXt network as follows, as well as a loss function; Input an image, and use the MobileNeXt network to extract features from the image. During the feature extraction process, five downsamplings are performed. Each downsampling outputs a feature map of a certain scale. The width and height of each feature map are 1 / 2 of the width and height of the previous layer's feature map. The feature map of the last layer is 1 / 32 of the original image; The feature map obtained after upsampling the feature map formed by operating on the feature map output by the MobileNeXt network through the Pyramid Scene Parsing module is merged with the fourth layer feature map to obtain a merged feature map. And so on for merging until the size of the merged feature map is 1 / 4 of the original image; then perform three downsamplings on the feature map with a size of 1 / 4 of the original image obtained after merging, save the feature map of each layer respectively, then use the Pyramid Scene Parsing module to aggregate the feature map of the last layer of feature extraction, and finally merge the feature map of each layer respectively to output 6 feature maps with a size of 1 / 4 of the original image; Step 3: Use the produced data set to train the constructed text detection network model and loss function to obtain a trained text detection network model; Step 4: Obtain an image of a certain scene; Step 5: Use an open-source image operation library to perform fixed-size scaling and normalization processing on the obtained image; Step 6: Use the trained text detection network model to perform inference and prediction on the image processed in Step 5, and extract the text region in the image; Step 7: Use Bezier curves to correct the text region in the image extracted in Step 6 to obtain a corrected text region; Step 8: Preprocess the image of the corrected text region, and then use the CRNN text recognition algorithm to extract text information from the text region in the preprocessed image.
2. The text recognition method based on deep learning according to claim 1, wherein The above Step 1 produces a data set in accordance with the specified format, specifically including: Step 1-1: Collect image data of a certain scene; Step 1-2: Perform data annotation on the collected image data above, and mark the four vertices of each text box in the image respectively. And the four vertices are in clockwise order, and each image obtains one or more marked text boxes; Step 1-3: Make a data set according to the obtained text boxes of the image in the data format of PASCAL VOC.
3. The text recognition method based on deep learning according to claim 1, characterized in that The loss function constructed in the above Step 2 is specifically: Among them, D is the calculation function of the dice coefficient; S i is the set of the i-th predicted region, G i is the set of the i-th ground truth region, S i,x,y is the value of the pixel (x, y) in the i-th predicted region, G i,x,y is the value of the pixel (x, y) in the i-th ground truth region; And, define L c as the text region classification loss, and L s as the shrunk text region loss, and their calculation methods are as follows: L c = 1 - D(S n *M,G n *M) Among them, M is the mask of the real area during the training process, and S n is the set of pixel points in the predicted area, and G n is the set of pixel points in the real area; W is the mask of a single text area in S n and S n,x,y represents S n The pixel value of (x, y) in 4. The text recognition method based on deep learning according to claim 1, wherein, In the above Step 3, the stochastic gradient descent algorithm is used to optimize and train the constructed text detection network model and loss function.
5. The text recognition method based on deep learning according to claim 1, wherein The above Step 5 uses an open-source image processing library to perform fixed-size scaling and normalization processing on the obtained image, specifically including: Step 5-1: Perform size scaling processing on the obtained image according to the set width and height of the image; Step 5-2: Normalize the scaled image. Use an image operation library to read the image, convert the image into an operable array, then divide each number in the array by 255 for normalization. Subtract a fixed mean and divide by a fixed variance for each channel of the image. And if the image is read by an open-source image operation library, adjust the channel order of the image to make it the rgb channel order.
6. The text recognition method based on deep learning according to claim 1, characterized in that Step 7 corrects the text area in the extracted image using a B-spline curve, specifically including: Step 7-1: Obtain the upper and lower boundaries of the text area. For curved text, use a text detection network model to detect the curved text area in the image and calculate the bounding rectangle of the curved text area to obtain the bounding rectangle. Calculate the angle of the bounding rectangle and calculate the starting points of the upper and lower boundaries of the curved text area based on the border of the bounding rectangle. Step 7-2: Take 8 points from each of the upper and lower boundaries as required. Step 7-3: Use the taken upper and lower boundary points to fit two B-spline curves for the upper and lower boundaries. Step 7-4: Use the two B-spline curves of the fitted upper and lower boundaries to correct the text area in the image to obtain the corrected text area.
7. The text recognition method based on deep learning according to claim 1, wherein Step 8 preprocesses the image of the corrected text area, specifically including: using an open-source image operation library to grayscale the image of the corrected text area, and then scale the image.
8. An electronic device, characterized in that, including: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the text recognition method according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the text recognition method according to any one of claims 1-7.
Citation Information
Patent Citations
Text recognition method and device, electronic equipment and readable storage medium
CN113537187A