Scene text recognition method based on recursive learning
By building a scene text recognition network based on recursive learning, and recursive units of the convolution and Transformer modules replace the stacked network, the problems of large amount of parameters and long inference time are solved, and efficient feature extraction and recognition accuracy are improved.
Patent Information
- Application Number
- CN202310175764.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-02-28
AI Technical Summary
Traditional scene text recognition methods have problems such as large number of parameters and weak feature expression capabilities, and recursive learning methods have linear increases and performance degradation in inference time.
A scene text recognition network based on recursive learning is constructed, and a Transformer module with the convolution and self-attention mechanism that does not change the size of the feature map is connected in series, forming a recursive unit through residual connections, replacing the traditional stacked network, enlarging the receptive field and reducing parameters.
While reducing parameters, it improves feature robustness and recognition accuracy, realizes a lightweight and compact network structure, which can effectively deal with the variability and complexity of scene text, and balances the recognition accuracy and reasoning time through recursive distillation strategies.
Smart Images

Figure CN116259058B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image and graphics processing, and particularly relates to a scene text recognition method based on recursive learning. Background Art
[0002] As a visual form of language, text carries rich and precise information and plays an important role in human daily life. Recognizing scene text has a wide range of application scenarios, such as image search, instant translation, robot navigation, etc. Therefore, scene text recognition technology (STR) aimed at recognizing text in natural images has been widely studied and achieved some success. However, due to the diverse visual appearances of scene text, such as complex backgrounds, low resolution, various font styles, etc., it is difficult for traditional methods, such as those using hand-designed features by Neumann et al. in the literature "L. Neumann and J. Matas. Real-Time Scene Text Localization and Recognition. IEEE Conference on Computer Vision and Pattern Recognition, 2012, pp. 3538-3545.", to extract and describe complex scene text features.
[0003] How to effectively extract feature information from natural scene text images to address the above problems is the key to improving text recognition accuracy. Deep learning has achieved certain success in the field of computer vision and significantly enhanced the performance of STR with its powerful feature extraction capabilities. Meanwhile, in practice, researchers have found that recognition performance can be further improved in two aspects: deeper network structures and attention-based mechanisms. Based on these two aspects, He et al. proposed residual learning in the literature "K. He, X. Zhang, S. Ren, et al. Deep Residual Learning for Image Recognition. IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770 - 778.", which overcomes the problem of gradient vanishing / exploding and thus enables the construction of very deep network structures. Vaswani et al. constructed an encoder-decoder network structure completely based on the attention mechanism in the literature "A. Vaswani, N. Shazeer, N. Parmar. Attention is All you Need. Neural Information Processing Systems, 2017, pp. 5998 - 6008." to fully capture long-range dependencies. However, these structures build deep networks by stacking parameter layers, which means that a very deep network will bring a large number of parameters. The large number of parameters limits the use of these structures and STR models based on them in storage-constrained devices. Therefore, a new lightweight and efficient visual feature extraction backbone is necessary.
[0004] To reduce parameters while maintaining or improving performance, one of the most common and simplest methods is weight sharing or recursive learning. For example, Lee et al. proposed replacing the stacked convolutional layers in ResNet with recursive convolutional layers stage by stage in the literature "C. Lee, and S. Osindero. Recursive Recurrent Nets with Attention Modeling for OCR in the Wild. IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 2231 - 2239." to improve the performance of scene text recognition and reduce the number of parameters. However, the following two problems limit its further application: the inference time increases linearly with the number of recursions, and a high number of recursions leads to a high inference time; the design of the recursive unit, an inappropriate recursive unit will lead to a significant performance degradation. Summary of the Invention
[0005] To overcome the deficiencies of the prior art, the present invention provides a scene text recognition method based on recursive learning. A scene text recognition network including a visual feature extraction module and a text sequence decoding module is constructed. In the visual feature extraction module, two convolutions that do not change the size of the feature map and a Transformer module based on the self-attention mechanism are connected in series, and then the input and output are short-circuited through a residual connection to form a recursive unit. By performing a certain number of recursive iterations instead of constructing a traditional stacked network, it is possible to significantly reduce the model parameters while increasing the receptive field and refining the features. The present invention can solve the problems of large number of parameters and weak feature expression ability of the traditional stacked scene text recognition visual feature network, and has the advantages of high parameter efficiency, lightweight and compact network, and robust features.
[0006] A scene text recognition method based on recursive learning, characterized by the following steps:
[0007] Step 1: Combine the publicly available MJSynth and SynthText English data sets into a training data set. The synthesis process includes: font rendering, border and shadow rendering, background coloring, synthesis of font, border and background, application of projection distortion, mixing with real-world images, and adding noise;
[0008] Step 2: Use the training data set to train the scene text recognition network to obtain a trained network. Among them, the scene text recognition network includes a feature map pre-extraction module, a visual feature extraction module and a text sequence decoding module. The specific parameter settings during training are as follows: Use the AdaDelta optimizer, the decay rate β = 0.95, the batch size is 192, the number of iterations is 300000 times, and an evaluation is performed every 2000 iterations. The gradient clipping threshold is 5, and all parameters are initialized using the He's method;
[0009] The specific processing process of the feature map pre-extraction module is as follows: Input the picture, perform feature extraction through a convolutional layer and a pooling layer to obtain a feature map. The feature map is fused with the spatial transformation parameter matrix to obtain an affine-transformed feature map, and the thin plate spline interpolation transformation is performed on the affine-transformed feature map. The transformed feature map is used as the input of the next module;
[0010] The described visual feature extraction module includes three convolutional layers and a recursive residual module. The specific processing process is as follows: Input the feature map, pass it through two convolutional layers with a convolutional kernel size of 3×3, a stride of 2, and a padding of 1, which halves the width and height of the feature map and increases the number of channels from 3 to 128 and then to 256, obtaining the feature map V1; The recursive residual module consists of a convolutional layer, a Transformer module, and a residual connection module. The feature map V1 passes through two convolutional layers with a convolutional kernel size of 3×3, a stride of 1, and a padding of 1 and an 8-head 1-layer Transformer module to obtain the feature map V′2, and the feature map V′2 is added to V1 through the residual connection to obtain the feature map Replace the feature map Input into the recursive residual module instead of the feature map V1, and repeat the above process until the set number of recursions is reached; The feature map finally obtained by the recursive module Pass through a convolutional layer with a convolutional kernel size of 3×3, a stride of 2, and a padding of 1, which halves the width and height of the feature map and increases the number of channels to 512, obtaining the feature map V3; Among them, each convolutional layer includes a ReLU activation function and a batch normalization function;
[0011] The described text sequence decoding module adopts a sequence decoding method based on a bidirectional long short-term memory language model and an attention mechanism. The specific decoding process is as follows: The output feature map V3 of the visual feature extraction network passes through an LSTM unit to obtain the hidden state at the current moment. Use the hidden state at the current moment and the output features of the visual feature extraction network to calculate the Attention weights, and use the Attention weights to weight all features to obtain the context feature at the current moment. Input the hidden state and context feature at the current moment into a fully connected layer to calculate the character probability distribution at the current position, and splice the character probability distributions at all positions to obtain the probability distribution of the entire text sequence;
[0012] Step 3: Input the scene text image to be recognized into the network trained in Step 2, and take the category with the highest output probability as the recognition result of the text image. Further, when training the network in Step 2, first set the number of recursions to 10, pre-train the model until convergence, retain the network weights at this time to obtain a network model with a high number of recursions; Then, set the number of recursions to 1, train again until convergence, and retain the network weights at this time to obtain the final network model; When applying, select the network model with a high number of recursions or the final network model according to the task scenario and actual requirements.
[0013] The beneficial effects of the present invention are as follows: Since the visual feature extraction module includes a recursive module composed of a convolutional layer, a Transformer module, and a residual connection structure, this module replaces the traditional method of stacking modules to build a deep network through recursive iteration, which can effectively increase the receptive field of the module without increasing parameters and extract better high-level semantic visual feature expressions; due to the complementary advantages of the local feature extraction ability of the convolutional neural network and the long-distance dependence capture ability of the Transformer in the recursive module, it can further enhance the network feature extraction ability and improve feature robustness, thereby improving the network's ability to cope with the variability and complexity challenges of scene text and effectively improving the accuracy of scene text recognition; in addition, through recursive distillation, a balance can be achieved between the model recognition accuracy and the inference time, and models with different recursive times can be flexibly selected according to the scene characteristics and task requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a schematic diagram of the network structure of the visual feature extraction module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The present invention will be further described below in conjunction with the drawings and embodiments. The present invention includes but is not limited to the following embodiments.
[0016] The present invention provides a method for scene text recognition based on recursive learning. The core is to construct a lightweight feature extraction network, which replaces the traditional method of stacking modules to build a deep model through recursive learning, thereby significantly reducing the number of model parameters; secondly, a recursive distillation strategy is used to balance the recognition accuracy and the model inference time, and solve the problem of long inference time of recursive learning, and finally achieve optimization in three aspects of the existing natural scene text recognition method: model size, recognition accuracy, and inference time. The specific implementation process of the present invention is as follows:
[0017] Step 1: Select two large-scale synthetic English datasets, the publicly available mainstream MJSynth (MJ) and SynthText (ST), as the training datasets. The MJ dataset was proposed by Jaderberg et al. in the literature "M. Jaderberg, K. Simonyan, A. Vedaldi, et al. Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition. arXiv:1406.2227, 2014.", and the ST dataset was proposed by Gupta et al. in the literature "A. Gupta, A. Vedaldi and A. Zisserman. Synthetic Data for Text Localisation in Natural Images. IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 2315 - 2324.". The synthesis process of the two datasets is as follows: 1) Font rendering, 2) Border and shadow rendering, 3) Background coloring, 4) Synthesis of font, border and background, 5) Application of projection distortion, 6) Mixing with real-world images, 7) Adding noise. Among them, ST was originally designed for scene text detection. The pictures are cropped by word boxes and non-alphanumeric characters are filtered out for scene text recognition training.
[0018] Step 2: Use the training dataset to train the scene text recognition network to obtain a trained network. Among them, the training parameters set for the model are as follows: The optimizer uses the AdaDelta optimizer, and its decay rate is set to ρ = 0.95. The training batch size is 192, the number of training iterations is 300000 times, an evaluation is performed every 2000 iterations, the gradient clipping threshold is 5, and all parameters are initialized using the He's method.
[0019] The scene text recognition network constructed by the present invention includes a feature map pre-extraction module, a visual feature extraction module, and a text sequence decoding module.
[0020] For the input picture, first obtain the feature map through the feature map pre-extraction module. The specific processing process is as follows: Feature extraction is performed through the convolutional layer and the pooling layer to obtain the feature map. The feature map is fused with the spatial transformation parameter matrix to obtain the affine-transformed feature map. The thin plate spline interpolation transformation is performed on the affine-transformed feature map, and the transformed feature map is used as the input for the next module.
[0021] The visual feature extraction module is the core of the present invention, such as Figure 1As shown, it includes three convolutional layers and a recursive residual module. The specific processing process is as follows: The input feature map passes through two convolutional layers with a convolutional kernel size of 3×3, a stride of 2, and a padding of 1, reducing the width and height of the feature map by half respectively and increasing the number of channels from 3 to 128 and then to 256, obtaining the feature map V1; The recursive residual module consists of a convolutional layer, a Transformer module, and a residual connection module. The feature map V1 passes through two convolutional layers with a convolutional kernel size of 3×3, a stride of 1, and a padding of 1 and an 8-head 1-layer Transformer module to obtain the feature map V′2, and the feature map V′2 is added to V1 through a residual connection to obtain the feature map Replace the feature map with the feature map V1 and input it into the recursive residual module, repeating the above process until the set number of recursions is reached; The feature map finally obtained by the recursive module passes through a convolutional layer with a convolutional kernel size of 3×3, a stride of 2, and a padding of 1, reducing the width and height of the feature map by half respectively and increasing the number of channels to 512, obtaining the feature map V3. Among them, each convolutional layer contains a ReLU activation function and a batch normalization function to achieve non-linear transformation and accelerate the convergence speed of the deep network respectively; The Transformer module is a sequence-to-sequence model based on self-attention, used to better extract the long-range dependencies between features, complementary to the local feature extraction ability of the convolutional layer. Each layer in the encoder and decoder of the Transformer consists of a multi-head attention sub-layer and a fully connected feed-forward sub-layer. The recursive residual module uses the method of recursive learning to replace the traditional stacking of multiple convolutional layers to build a deep network, increasing the receptive field and extracting robust visual features, which can maintain the feature extraction ability while significantly reducing the huge parameter problem caused by stacking before.
[0022] When training the network, a high number of recursions can be used first, that is, the number of recursions of the recursive residual module is set to 10, and the pre-trained model is trained until convergence, and then the number of recursions is set to 1 and trained again until convergence. The weights obtained from the two trainings are saved separately, and according to the task scenario and actual needs, the high-recursion model or the final-recursion model is selected.
[0023] The text sequence decoding module is based on a bidirectional long short-term memory language model and an attention mechanism. The specific decoding process is as follows: The output feature sequence map V3 of the visual feature extraction network passes through an LSTM unit. The LSTM unit calculates the hidden state at the current moment based on the feature vector at the current moment and the hidden state at the previous moment. Then, the attention weights are calculated using the hidden state at the current moment and the feature vectors at all moments. The attention weights are used to weighted average the feature vectors at all moments to obtain the context vector at the current moment. Next, using the hidden state and the context vector at the current moment as inputs, a fully connected layer is used to calculate the character probability distribution at the current position. Finally, the character probability distributions at all positions are concatenated to obtain the probability distribution of the entire text sequence.
[0024] Step 3: Input the scene text image to be recognized into the network trained in Step 2, and use the category with the highest probability as the recognition result of the text image.
[0025] To verify the effectiveness of the method of the present invention, a simulation experiment is carried out using python in the system environment of Linux version 5.0.0-23-generic (buildd@lgw01-amd64-030) (gcc version 7.4.0 (Ubuntu 7.4.0-1ubuntu1~18.04.1)). Four mainstream STR benchmark datasets are selected as the test sets for the experiment: the SVT dataset, the IC13 dataset, the IC15 dataset, and the SVTP dataset. Among them, the first two datasets are regular scene text datasets, and the last two are irregular scene text datasets. Specifically: SVT (Street View Text) is outdoor street images collected from Google Street View. Some images have high noise, are blurred or have low resolution, and it contains 647 evaluation images; IC13 contains 1095 images for evaluation. After filtering out words with non-alphanumeric characters, 1015 images can be obtained; IC15 is created for the 2015 ICDAR competition, which contains 2077 test set images. These photos are taken by Google Glass during the natural movements of the wearer. Therefore, many images are noisy, blurred and rotated, and some are of low resolution. SVTP is collected from Google Street View and contains 645 images for evaluation. Many images contain perspective projections.
[0026] In the experiment, the existing TRBA algorithm, RARM algorithm, Rosetta algorithm, and STAR-Net algorithm were selected as comparison algorithms. The recognition accuracy, average accuracy, and the number of algorithm model parameters of different algorithms on the above four datasets are shown in Table 1. It can be seen that the method of the present invention has a better recognition rate than other methods on the 4 scenario text datasets, and the number of parameters of the model is also the least, indicating the beneficial effects of the method of the present invention in terms of efficient parameters, robust features, and accurate recognition. At the same time, it can be seen from the ablation experiment data in Table 2 that although the recognition accuracy of the distilled model is lower than that of the model with a high number of recursions, it is higher than that of the model with a low number of recursions. At the same time, it has the advantage of the shorter inference time of the model with a low number of recursions, indicating the beneficial effect that the recursive distillation of the present invention can achieve a balance between the recognition accuracy and the inference time of the model.
[0027] Table 1
[0028]
[0029] Table 2
[0030]
Claims
1. A scene text recognition method based on recursive learning, characterized in that The steps are as follows: Step 1: Combine the publicly available English data sets of MJSynth and SynthText into a training data set. The synthesis process includes: font rendering, border and shadow rendering, background coloring, synthesis of font, border and background, application of projection distortion, mixing with real-world images, and adding noise; Step 2: Use the training data set to train the scene text recognition network to obtain a trained network. Among them, the scene text recognition network includes a feature map pre-extraction module, a visual feature extraction module, and a text sequence decoding module. The specific parameter settings during training are: use the AdaDelta optimizer, the decay rate β = 0.955, the batch size is 152, the number of iterations is 3,000 times, evaluate once every 200 times of iteration, use a gradient clipping threshold of 5, and all parameters are initialized using the He's method; The specific processing process of the feature map pre-extraction module is: input the picture, perform feature extraction through the convolutional layer and the pooling layer to obtain the feature map, fuse the feature map with the spatial transformation parameter matrix to obtain the affine-transformed feature map, perform thin plate spline interpolation transformation on the affine-transformed feature map, and use the transformed feature map as the input of the next module; The described visual feature extraction module includes three convolutional layers and a recursive residual module. The specific processing process is as follows: Input the feature map, pass it through two convolutional layers with a convolutional kernel size of 3×3, a stride of 2, and a padding of 1, which halves the width and height of the feature map and increases the number of channels from 3 to 128 and then to 256, obtaining the feature map V1; The recursive residual module consists of a convolutional layer, a Transformer module, and a residual connection module. The feature map V1 passes through two convolutional layers with a convolutional kernel size of 3×3, a stride of 1, and a padding of 1 and an 8-head 1-layer Transformer module to obtain the feature map V2′. The feature map V2′ is added to V1 through a residual connection to obtain the feature map Replace the feature map Input it into the recursive residual module instead of the feature map V1, and repeat the above process until the set number of recursions is reached; The feature map finally obtained by the recursive module Pass through a convolutional layer with a convolutional kernel size of 3×3, a stride of 2, and a padding of 1, which halves the width and height of the feature map and increases the number of channels to 512, obtaining the feature map V3; Among them, each convolutional layer includes a ReLU activation function and a batch normalization function; The text sequence decoding module adopts a sequence decoding method based on a bidirectional long short-term memory language model and an attention mechanism. The specific decoding process is: the output feature sequence map V3 of the visual feature extraction network passes through the LSTM unit to obtain the hidden state at the current moment, calculate the Attention weight using the hidden state at the current moment and the output features of the visual feature extraction network, use the Attention weight to weight all features to obtain the context feature at the current moment, input the hidden state and context feature at the current moment into a fully connected layer, calculate the character probability distribution at the current position, and splice the character probability distributions at all positions to obtain the probability distribution of the entire text sequence; Step 3: Input the scene text picture to be recognized into the network trained in Step 2, and use the category with the highest probability as the recognition result of the text picture.
2. The method for scene text recognition based on recursive learning according to claim 1, wherein: When performing network training in Step 2, first set the number of recursions to 1, pre-train the model until convergence, retain the network weights at this time to obtain a network model with a high number of recursions; then, set the number of recursions to 1 and train again until convergence, retain the network weights at this time to obtain the final network model; when applying, select the network model with a high number of recursions or the final network model according to the task scenario and actual requirements.
Citation Information
Patent Citations
A Chinese scene text line identification method based on residual convolution and a recurrent neural network
CN109948714A
Text recognition method and system based on decoupling attention mechanism
CN111967470A