A text recognition method

Through the text recognition method of Bessel network and end-to-end training, the problem of insufficient efficiency and accuracy of irregular text area recognition in the prior art is solved, efficient and accurate text recognition is achieved, and suitable for low-performance devices.

CN116110052BActive Publication Date: 2025-07-11JIANGSU AEROSPACE DAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211589738.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-12
Publication Date
2025-07-11
Estimated Expiration
2042-12-12

AI Technical Summary

Technical Problem

Existing end-to-end text recognition technology is not ideal in terms of efficiency and accuracy, especially when dealing with irregular text areas.

Method used

Feature extraction is performed using Bezier network, combined with detection module, area attention module and recognition module for end-to-end training, use Bezier curve to fit character areas, and perform sequence modeling through bidirectional LSTM, and finally use lightweight model parameters to be applied to low-performance devices.

Benefits of technology

It improves the accuracy and efficiency of identifying irregular text areas, and is suitable for character area detection in various shapes, meeting the operation requirements of low-performance edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116110052B_ABST
    Figure CN116110052B_ABST
Patent Text Reader

Abstract

The present application discloses a text recognition method, which relates to the field of text recognition. This method uses a Bezier network to extract features from the input image to be recognized. The text region image obtained by feature extraction is processed by a detection module, a text correction module, a regional attention module, and a recognition module, and finally a text recognition result is obtained. Each module forms a text recognition model through end-to-end training. The model is small and fast. Moreover, the Bezier network can detect character regions of various shapes, and combined with the regional attention module, the recognition accuracy of the text region image fitted by the Bezier network is improved. Therefore, the method of the present application can accurately and efficiently recognize irregular text regions, has a wide range of applications, and has high recognition efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of character recognition, and in particular, to a character recognition method. Background Art

[0002] Character recognition is an important field in computer vision research. In people's daily production and life, a large amount of text information needs to be processed. Character recognition can reduce labor costs and improve efficiency. Character recognition technology has broad prospects in fields such as smart cities, information automation, and industrial automation.

[0003] The existing mainstream character recognition technologies mainly include two types: two-stage character recognition and end-to-end character recognition. Among them, two-stage character recognition realizes text region detection through feature extraction in the first stage, and then uses feature extraction in the second stage for the detected text region to realize character recognition. The end-to-end character recognition based on deep learning has a simple, efficient and unified structure, can complete the two steps of detection and recognition simultaneously, and has the advantages of a smaller model and faster speed. It has gradually replaced the two-stage character recognition method that trains the text region detection and character recognition in stages and then stitches them together, and has become one of the mainstream research directions for natural scene text detection and recognition. However, the current efficiency and accuracy of end-to-end character recognition are still not ideal. Summary of the Invention

[0004] In view of the above problems and technical requirements, the applicant of this application proposes a character recognition method. The technical solution of this application is as follows:

[0005] A character recognition method, the character recognition method includes:

[0006] Performing feature extraction on the input image to be recognized by using a Bessel network to obtain a text region image containing text information;

[0007] Performing text prediction on the text region image by using a detection module to obtain a text prediction result;

[0008] Performing splicing on the text region image, the text prediction result, and the result after the text prediction result is processed by a text correction module by using a region attention module to obtain an attention feature;

[0009] Performing recognition on the attention feature by using a recognition module to obtain a character recognition result, and the region attention module is further used to return the recognition loss of the recognition module to the Bessel network;

[0010] Wherein, the Bessel network, the detection module, the region attention module, the text correction module, and the recognition module are modules in a character recognition model obtained through end-to-end training.

[0011] A further technical solution is that the Bessel network includes a convolution module, a Bessel curve module, and a Bessel alignment module. The method for using the Bessel network to extract features from the input image to be recognized includes:

[0012] Using the convolution module to process the image to be recognized to obtain a character region containing text information in the image to be recognized;

[0013] Using the Bessel curve module to perform Bessel curve fitting on the character region;

[0014] Using the Bessel alignment module to transform the result of the Bessel curve fitting to obtain a text region image.

[0015] A further technical solution is that the text prediction result output by the detection module includes a character region map and a character link map, and the text prediction result is output to the text correction module and the region attention module at the same time.

[0016] A further technical solution is that the method for using the recognition module to recognize the attention features to obtain a text recognition result includes:

[0017] After feature extraction of the text recognition result, bidirectional LSTM is used for sequence modeling to obtain a sequence result, and a decoder based on attention is used to decode the sequence result to obtain text information to obtain the text recognition result.

[0018] A further technical solution is that bidirectional LSTM is used for multiple sequence modelings, and the results of the multiple sequence modelings are averaged to obtain a sequence result.

[0019] A further technical solution is that the method further includes:

[0020] After determining the model parameters through end-to-end training of the model structure constructed based on the Bessel network, the detection module, the region attention module, the text correction module, and the recognition module, the trained model parameters are mapped to parameters with a low bit width within the accuracy range to obtain a text recognition model.

[0021] The beneficial technical effects of this application are:

[0022] This application discloses a text recognition method. This method uses a Bessel network to extract features from the input image to be recognized, so as to realize the detection of character regions of various shapes. Combining with the region attention module improves the recognition accuracy of the text region image fitted by the Bessel network, so that the method of this application can accurately and efficiently realize the recognition of irregular text regions, with a wide range of applications, high recognition efficiency and accuracy.

[0023] In addition, the end-to-end model obtained through lightweight training can meet the operating requirements of more low-performance edge devices. Brief Description of the Drawings

[0024] Figure 1 is a schematic flowchart of a character recognition method in an embodiment of the present application. Detailed Embodiments

[0025] The following further describes the detailed embodiments of the present application with reference to the accompanying drawings.

[0026] The present application discloses a character recognition method. Please refer to Figure 1 , and the character recognition method includes the following processes:

[0027] Use a Bezier network to perform feature extraction on the input image to be recognized to obtain a text region image containing text information. In one embodiment, the Bezier network includes a convolution module, a Bezier curve module, and a Bezier alignment module. Then, use the convolution module to process the input image to be recognized to obtain a character region containing text information in the image to be recognized. Then, use the Bezier curve module to perform Bezier curve fitting on the character region. Finally, use the Bezier alignment module to transform the result of the Bezier curve fitting to obtain the text region image.

[0028] In conventional end-to-end character recognition applications, VGG16 or ResNet50 is generally used for feature extraction, while the present application uses a Bezier network to replace VGG16 and ResNet50 for feature extraction. Using the Bezier curve module to perform Bezier curve fitting on the character region, compared with the conventional method of using rectangular border detection to only detect character regions with rectangular structures, using Bezier curve fitting can detect character regions of various shapes. In addition, the Bezier alignment module BezierAlign is different from RoIAlign, and the shape of its sampling grid is not rectangular, which can accurately calculate the convolutional features of curved character regions and achieve high recognition accuracy with relatively small computational overhead.

[0029] Then, use a detection module to perform text prediction on the text region image to obtain a text prediction result. The text prediction result output by the detection module includes a character region map and a character connection map. The detection module can use a character position-aware text detection algorithm CRAFT designed based on a convolutional neural network CNN for text localization. The convolutional neural network generates a character region score and a mutual relationship score. The region score is used to locate individual characters in the image, while the mutual relationship score is used to group each character into an instance. During inference, a character region map of any shape can be output.

[0030] The text prediction results are output to the text correction module and the regional attention module at the same time. The text correction module processes the text prediction results and outputs them to the regional attention module. The text correction module uses thin plate spline interpolation (TPS) to correct text area images of arbitrary shapes. By updating the control points around the text area image, the curved geometric shape of the text in the text area image can be improved.

[0031] The regional attention module is used to stitch the text region image, the text prediction result, and the text prediction result after being processed by the text correction module to obtain the attention feature. Then the recognition module is used to recognize the attention feature to obtain the text recognition result. After the recognition module extracts the features of the text recognition result, a bidirectional LSTM is used to perform sequence modeling to obtain the sequence result, and an attention-based decoder is used to decode the text information of the sequence result to obtain the text recognition result. In one embodiment, a bidirectional LSTM is used to perform multiple sequence modeling, and the results of multiple sequence modeling are averaged to obtain the sequence result, thereby reducing randomness. In addition, the regional attention module also returns the recognition loss loss of the recognition module to the Bessel network.

[0032] The model structure constructed based on the Bessel network, the detection module, the regional attention module, the text correction module and the recognition module is end-to-end trained to obtain a text recognition model. The Bessel network, the detection module, the regional attention module, the text correction module and the recognition module are modules in the text recognition model obtained through end-to-end training. In actual use, the text recognition process can be realized by inputting the image to be recognized into the text recognition model.

[0033] After end-to-end training of the model structure constructed based on the Bessel network, detection module, regional attention module, text correction module and recognition module, it can be directly used as a text recognition model. Or in another embodiment, after determining the model parameters, the trained model parameters are mapped to low-bit width parameters within the accuracy range to obtain a text recognition model. That is, the high-precision model parameters are converted into low-precision model parameters, thereby reducing the accuracy of the text recognition model as much as possible within the accuracy range, and improving the performance of the text recognition model on the basis of meeting the accuracy requirements, so that it can be applied to some low-performance edge devices.

[0034] The above is only a preferred embodiment of the present application, and the present application is not limited to the above embodiments. It is understood that other improvements and changes directly derived or associated by those skilled in the art without departing from the spirit and concept of the present application should be considered to be included in the protection scope of the present application.

Claims

1. A character recognition method, characterized in that, The described text recognition method includes: Using a Bezier network to extract features from the input image to be recognized to obtain a text region image containing text information; Using a detection module to perform text prediction on the text region image to obtain a text prediction result; Using a region attention module to splice the text region image, the text prediction result, and the result after the text prediction result is processed by a text correction module to obtain an attention feature; Using an identification module to identify the attention feature to obtain a text recognition result, and the region attention module is also used to return the recognition loss of the identification module to the Bezier network; Wherein, the Bezier network, the detection module, the region attention module, the text correction module, and the identification module are modules in a text recognition model obtained through end-to-end training.

2. The method according to claim 1, wherein The Bezier network includes a convolution module, a Bezier curve module, and a Bezier alignment module. The method for using the Bezier network to extract features from the input image to be recognized includes: Using the convolution module to process the image to be recognized to obtain a character region containing text information in the image to be recognized; Using the Bezier curve module to perform Bezier curve fitting on the character region; Using the Bezier alignment module to transform the result of the Bezier curve fitting to obtain the text region image.

3. The method according to claim 1, wherein The text prediction result output by the detection module includes a character region map and a character link map, and the text prediction result is simultaneously output to the text correction module and the region attention module.

4. The method according to claim 1, wherein The method for using an identification module to identify the attention feature to obtain a text recognition result includes: After extracting features from the text recognition result, using a bidirectional LSTM for sequence modeling to obtain a sequence result, and using an attention-based decoder to decode the text information of the sequence result to obtain a text recognition result.

5. The method according to claim 4, characterized in that, Performing multiple sequence modelings using a bidirectional LSTM, and averaging the results of the multiple sequence modelings to obtain the sequence result.

6. The method according to claim 1, wherein The method further includes: After determining the model parameters through end-to-end training of the model structure constructed based on the Bezier network, the detection module, the region attention module, the text correction module, and the identification module, mapping the trained model parameters to low-bit-width parameters within the accuracy range to obtain the text recognition model.

Citation Information

Patent Citations

  • Character recognition method and device based on deep learning, equipment and storage medium

    CN114155540A

  • Complex background seal identification method based on deep learning

    CN115187978A