A Real-time Irregular Text Recognition Method in Complex Scenarios
By using lightweight convolutional networks and decoupled attention modules in irregular text recognition algorithms, the problem of difficult to take into account recognition accuracy and efficiency in complex scenarios is solved, and efficient and accurate text recognition performance is achieved.
Patent Information
- Application Number
- CN202111452587.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-01
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-12-01
AI Technical Summary
Existing irregular text recognition algorithms are difficult to take into account both recognition accuracy and efficiency in complex scenarios, especially when text curvature is high and background noise is high.
A lightweight convolutional network is used to extract image features, and combined with the decoupled attention module and self-attention module, we will learn local and global text features respectively to improve global attention to local attention to improve efficiency.
Implement high-precision irregular text recognition in complex scenarios, and support the recognition of regular horizontal printed text, improving the real-time inference performance of the algorithm on different hardware platforms.
Smart Images

Figure CN114495119B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for recognizing text images in the field of text recognition, specifically a method for real-time irregular text recognition in complex scenarios. Background Art
[0002] Optical Character Recognition (OCR) refers to converting the text in an image into editable text information by a computer, which mainly includes two stages: text detection and text recognition. As an important application branch of computer vision, text recognition has a wide range of implementation scenarios and research significance, such as the entry of bill and card information, the recognition of handwritten text in educational scenarios, the content information review of network pictures, and so on. The text recognition for printed text and network text is becoming increasingly mature, while the text recognition algorithm in natural complex scenarios has problems due to the overly wide data distribution, such as a large curvature or slope of the text distribution shape, a small proportion of the text part in the picture, and a large background noise. The current mainstream text recognition methods use an additional correction network or a computationally intensive attention mechanism to handle the text in these scenarios, which makes it difficult for the recognition algorithm to balance accuracy and efficiency. Therefore, for the irregular text in many real-world scenarios, many current algorithms are difficult to meet the industrial application level in terms of both speed and performance.
[0003] Currently, many mainstream text detection algorithms based on segmentation can cut out irregular text and cooperate with a recognition algorithm based on a correction network for text recognition. However, once the scenario is complex and the text curvature is too large, the correction ability of the correction network is difficult to meet the recognition requirements. Summary of the Invention
[0004] In order to overcome the difficulty of balancing the speed and accuracy of model inference in existing irregular text recognition algorithms, the present invention provides a method for real-time irregular text recognition in complex scenarios. The present invention aims to handle the recognition of irregular text in various complex scenarios. By using a lightweight convolutional network to extract image features, as well as an attention module for decoupling and extracting the features of the text part and a self-attention module for context semantic modeling, each module learns its own information, achieving the ability to balance the recognition of text arranged horizontally at a conventional level and also having good recognition performance in complex scenarios. At the same time, a lightweight backbone network is selected and the global attention is improved to local attention, greatly improving the running efficiency of the algorithm, so that the algorithm can achieve the performance of real-time inference on different hardware platforms.
[0005] The technical solution adopted by the present invention is as follows:
[0006] The present invention includes the following steps:
[0007] 1) After preprocessing the natural complex scene image containing text, a preprocessed text block image is obtained;
[0008] 2) After inputting the preprocessed text block image into a backbone convolutional neural network for feature extraction and feature slicing, multiple sliced local text features are obtained and the serial numbers of each sliced local text feature are recorded;
[0009] 3) The multiple sliced local text features are respectively input into multiple local self-attention modules according to the serial numbers corresponding to the sliced local text features. After each local self-attention module performs local text feature aggregation in parallel, local text feature aggregation vectors are respectively output;
[0010] 4) The multiple local text feature aggregation vectors are simultaneously input into a global self-attention module for splicing, downsampling, and global text feature extraction, and then a global text feature vector is output;
[0011] 5) The global text feature vector is input into a character decoding module for decoding, and then the recognized text string is output.
[0012] The specific content of step 1) is as follows:
[0013] First, after using a text detection algorithm to intercept text from the natural complex scene image containing text, an original text block image is obtained. Then, the height of the original text block image is interpolated to a preset height H, and the width of the original text block image after adjusting the height is scaled according to the aspect ratio of the original text block image to obtain a standard text block image. Finally, a normalization operation with a mean of 0 and a variance of 1 is performed on the standard text block image to obtain a preprocessed text block image.
[0014] The backbone convolutional neural network in step 2) is mainly composed of a convolutional downsampling module connected to a feature slicing module after passing through a 6-layer depth convolutional module, a first separable convolutional downsampling module, a 12-layer depth convolutional module, and a second separable convolutional downsampling module in sequence. The preprocessed text block image is input into the convolutional downsampling module, and the feature slicing module outputs multiple sliced local text features and records the serial numbers of each sliced local text feature.
[0015] The input of the feature slicing module is the visual feature map output by the second separable convolutional downsampling module. The feature slicing module determines the size of the sliding window according to the height of the visual feature map, slides pixels in the visual feature map using the sliding window, and after sliding, copies and slices the visual feature map covered by the sliding window as a sliced local text feature and outputs it. Continuous sliding slicing is performed to output multiple sliced local text features and record the serial numbers of each sliced local text feature.
[0016] In step 3), the structures of multiple local self-attention modules are the same, specifically:
[0017] The sliced local text features are formed by splicing multiple column vectors in the width dimension. Each local self-attention module copies the middle column vector of the input sliced local text features and linearly transforms it to obtain a local text query vector. Then, after cutting the column vectors of the input sliced local text features and rearranging them in the height dimension, different linear transformations are performed on the rearranged sliced local text features to obtain a local text key vector and a local text value vector respectively. After that, a matrix dot product is performed on the local text query vector and the local text key vector, followed by a Softmax operation to obtain a weight distribution matrix of the local attention mechanism. Then, a matrix dot product is performed on the weight distribution matrix of the local attention mechanism and the local text value vector to obtain a local text feature aggregation vector output by the current local self-attention module. The calculation formula is as follows:
[0018]
[0019] Ki = W K Ti r
[0020] Vi = W V Ti r
[0021]
[0022] Among them, W Q W K , W Q W K W V represent the first, second, and third local linear transformation parameter matrices respectively, represents a dimension of C×C, represents the middle column vector of the i-th sliced local text feature, H′ is the height of the sliced local text feature, Ti r represents the sliced local text feature after cutting the column vectors of the i-th sliced local text feature and rearranging them in the height dimension, Qi, Ki, and Vi are the local text query vector, local text key vector, and local text value vector corresponding to the i-th sliced local text feature respectively, Ti out is the local text feature aggregation vector corresponding to the i-th sliced local text feature, C is the feature dimension of the local text feature aggregation vector, T represents the transpose operation; Softmax(·) represents the Softmax operation, specifically the normalization process.
[0023] Step 4) is specifically:
[0024] The global self-attention module concatenates multiple local text feature aggregation vectors in the width dimension according to the serial numbers of the corresponding sliced local text features to obtain a text feature aggregation vector. Then, average pooling downsampling is performed on the text feature aggregation vector in the height dimension to obtain a global text feature sequence X. After performing different linear transformations on the global text feature sequence, the query vector, key vector, and value vector of the global text are obtained respectively. After performing matrix dot product on the query vector and key vector of the global text and then performing the Softmax operation, a weight distribution matrix of the global attention mechanism is obtained. Then, after performing matrix dot product on the weight distribution matrix of the global attention mechanism and the value vector of the global text, a global text feature vector is obtained. The calculation formula is as follows:
[0025]
[0026] Q g =W g q X
[0027] K g =w g k X
[0028] V g =W g v X
[0029]
[0030] Among them, [·] represents the concatenation operation in the width dimension, and AveragePooling(·) represents the average pooling downsampling operation to 1 in the height dimension. represents the i-th local text feature aggregation vector, Q g , K g , V g are the query vector, key vector, and value vector of the global text, W g q , W g k , W g v represent the first, second, and third global linear transformation parameter matrices, X out is the global text feature vector, C is the feature dimension of the global text feature vector, and T represents the transpose operation; Softmax(·) represents the Softmax operation, specifically the normalization process.
[0031] The specific content of step 5) is as follows:
[0032] In the character decoding module, first, the global text feature vector is linearly transformed so that the feature dimension of the linearly transformed global text feature vector is equal to the number of character categories. Then, after performing a Softmax operation on the linearly transformed global text feature vector, a character probability distribution is obtained. Decoding is achieved by performing a character dictionary mapping on the linearly transformed global text feature vector according to the character probability distribution, and the recognized text recognition sequence is output.
[0033] The beneficial effects of the present invention are mainly manifested in:
[0034] Compared with the current mainstream irregular text recognition methods, the present invention uses a local self-attention module to replace the conventional attention module with a relatively high complexity, and decouples the context sequence modeling module from this part, and only performs operations on the global self-attention mechanism on the sequence downsampled to one dimension, greatly reducing the computational complexity. At the same time, it takes into account the extraction of text features in the local area and the extraction of context semantics of the global sequence, effectively improving the recognition performance of the model. At the same time, a lightweight backbone network is used, enabling the algorithm to achieve real-time effects on various computing platforms.
[0035] The present invention can not only accurately recognize irregular texts in complex scenarios, but also well support the recognition of conventional horizontally printed texts, and can handle more and wider business scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flowchart of the method of the present invention.
[0037] Figure 2 is a structural logic diagram of the overall network in the present invention.
[0038] Figure 3 is a schematic structural diagram of the backbone network in the present invention.
[0039] Figure 4 is a schematic structural diagram of the local self-attention module (LSA) in the present invention.
[0040] Figure 5 is a schematic structural diagram of the global self-attention module (SA) in the present invention.
[0041] Figure 6 is an example diagram of two preprocessed text block images in an embodiment of the present invention.
[0042] Figure 7 is Figure 6 a schematic diagram of the weight distribution matrix of the local attention mechanism in two preprocessed text block images in each local self-attention module.
[0043] Figure 8 is toFigure 7 The schematic diagram of the weight distribution matrix of the two local attention mechanisms, after being transformed and scaled to the size of the original image, corresponds to the visualized weight distribution on the original image. The whiter the area in the weight map, the richer the text features in that area. Conversely, the area is probably a background area. Detailed implementation manners
[0044] To more clearly illustrate the purpose and technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0045] As Figure 1 and Figure 2 shown, the present invention includes the following steps:
[0046] 1) After preprocessing the natural complex scene image containing text, a preprocessed text block image is obtained;
[0047] Specifically, step 1) is as follows:
[0048] First, after using a text detection algorithm to intercept the text in the natural complex scene image containing text, an original text block image is obtained. Then, the height of the original text block image is interpolated to a preset height H, and the width of the original text block image after adjusting the height is scaled according to the aspect ratio of the width and height of the original text block image, obtaining a standard text block image x ∈ R H ×W×3 , and finally, a normalization operation with a mean of 0 and a variance of 1 is performed on the standard text block image to obtain a preprocessed text block image.
[0049] 2) After inputting the preprocessed text block image into the backbone convolutional neural network for feature extraction and feature slicing, multiple sliced local text features are obtained and the serial numbers of each sliced local text feature are recorded;
[0050] As Figure 3 shown, the backbone convolutional neural network in step 2) is mainly composed of a convolutional downsampling module connected to a feature slicing module after passing through a 6-layer depth convolutional module, a first separable convolutional downsampling module, a 12-layer depth convolutional module, and a second separable convolutional downsampling module in sequence. The preprocessed text block image is input into the convolutional downsampling module, and the feature slicing module outputs multiple sliced local text features;
[0051] Among them, the input of the feature slicing module is the visual feature map output by the second separable convolution downsampling module. The feature slicing module determines the size of the sliding window according to the height of the visual feature map. The dimension of the visual feature map is (H′, L, C), where H′ and L are the height and width of the visual feature map respectively. In this embodiment, the height H′ of the visual feature map is equal to four pixels, C is the feature dimension of the visual feature map, the height and width of the sliding window are H′ and H′ + 1 respectively, and the sliding window slides on the visual feature map in pixels. Specifically, the sliding window slides in one direction in the width dimension of the visual feature map by one pixel at a time. After sliding, the visual feature map covered by the sliding window is copied and sliced as a sliced local text feature and output. Continuous sliding slicing is performed to output multiple sliced local text features T. It is indicated that the dimension of the sliced local text feature is H′×W′×C, and W′ = H′ + 1. At the initial sliding, zero vectors with a width and height of H′ are filled at both ends of the width of the visual feature map, and the serial numbers of each sliced local text feature are recorded.
[0052] 3) Multiple sliced local text features are respectively input into multiple Local Self-Attention (LSA) modules according to the serial numbers corresponding to the sliced local text features. After each local self-attention module performs local text feature aggregation in parallel, local text feature aggregation vectors are respectively output. The number of local self-attention modules is equal to the number of sliced local text features output by the backbone convolutional neural network, and is also equal to the number of column vectors in the visual feature map; the local self-attention mechanism is used in parallel to calculate the spatial attention distribution of the text in the image, realizing parameter sharing and parallel calculation of the self-attention mechanism, effectively improving the weight distribution of the traditional spatial attention, enhancing the recognition accuracy, solving the problem of attention drift, and at the same time greatly reducing the number of parameters and the amount of computation, and achieving the best results on multiple public test sets for text recognition.
[0053] In step 3), the structures of multiple local self-attention modules are the same, specifically as follows:
[0054] As Figure 4 shown, the sliced local text feature is composed of multiple column vectors spliced in the width dimension. In this embodiment, the width of the column vector is equal to one pixel. Each local self-attention module copies the middle column vector of the input sliced local text feature and linearly transforms it to obtain a local text query vector Q. Then, after the input sliced local text feature is cut into column vectors and rearranged in the height dimension, different linear transformations are performed on the rearranged sliced local text feature to respectively obtain a local text key vector K. and a local text value vector V. After performing matrix dot product on the local text query vector and the local text key vector and then performing the Softmax operation, a weight distribution matrix of the local attention mechanism is obtained. After performing matrix dot product on the weight distribution matrix of the local attention mechanism and the local text value vector, a local text feature aggregation vector output by the current local self-attention module is obtained. The calculation formula is as follows:
[0055]
[0056]
[0057] Ki = W K Ti r
[0058] Vi = W V Ti r
[0059]
[0060] where Ti represents the local text feature of the i-th slice. i is also the serial number of the column vector in the visual feature map, and i ∈ [1, 2, .., L]; represents the visual feature map covered by the sliding window when the i-th column vector is the intermediate vector of the local text feature of the slice. W Q W K W Q W K W V respectively represent the first, second, and third local linear transformation parameter matrices. represents a matrix with dimensions C×C. represents the middle column vector of the local text feature of the i-th slice. H′ is the height of the local text feature of the slice, that is, the height of the visual feature map. Ti r represents the local text feature of the slice after column vector cutting and height dimension rearrangement of the local text feature of the i-th slice. Qi, Ki, and Vi are the local text query vector, local text key vector, and local text value vector corresponding to the local text feature of the i-th slice respectively. Ti out is the local text feature aggregation vector corresponding to the i-th slice local text feature, C is the feature dimension of the local text feature aggregation vector, T represents the transpose operation; Softmax(·) represents the Softmax operation, specifically normalizing the result of the matrix dot product to between 0 and 1, which can enable the local self-attention module to assign higher weights to the text regions in the image and lower weights to the background regions, thereby excluding the interference of complex backgrounds. Different from previous methods, this invention does not generate an attention weight map in an iterative autoregressive manner, but can perform parallel calculations on the slice local text features within each sliding window to obtain a local text vector group, greatly improving the efficiency.
[0061] 4) Input multiple local text feature aggregation vectors into the global self-attention module for concatenation, downsampling, and global text feature extraction, and then output the global text feature vector;
[0062] Step 4) is specifically as follows:
[0063] As Figure 5 shown, the global self-attention module concatenates multiple local text feature aggregation vectors in the width dimension according to the serial numbers of the corresponding slice local text features to obtain a text feature aggregation vector, and then performs mean pooling downsampling on the text feature aggregation vector in the height dimension to obtain the global text feature sequence X. After performing different linear transformations on the global text feature sequence, the query vector, key vector, and value vector of the global text are obtained respectively. After performing matrix dot product on the query vector and key vector of the global text and then performing the Softmax operation, the weight distribution matrix of the global attention mechanism is obtained. Then, after performing matrix dot product on the weight distribution matrix of the global attention mechanism and the value vector of the global text, the context semantic association of the global text feature sequence is modeled to obtain the global text feature vector. The calculation formula is as follows:
[0064]
[0065] Q g =W g q X
[0066] K g =W g k X
[0067] K g =W g v X
[0068]
[0069] Among them, [·] represents the concatenation operation in the width dimension, and AveragePooling(·) represents the operation of downsampling the average pooling to 1 in the height dimension. represents the i-th local text feature aggregation vector, where i ∈ [1, 2,.., L], and Q g , K g , V g are the query vector, key vector, and value vector of the global text, and W g q , W g k , W g v represent the first, second, and third global linear transformation parameter matrices, satisfying W g q , W g k , X out is the global text feature vector, satisfying C is the feature dimension of the global text feature vector, that is, the feature dimension of the visual feature map and the feature dimension of the local text feature aggregation vector. T represents the transpose operation; Softmax(·) represents the Softmax operation, specifically normalizing the result of the matrix dot product to between 0 and 1.
[0070] 5) After inputting the global text feature vector into the character decoding module for decoding, the recognized text string is output.
[0071] Step 5) is specifically as follows:
[0072] In the character decoding module, first perform a linear transformation on the global text feature vector so that the feature dimension of the linearly transformed global text feature vector is equal to the number of text categories. Then, perform a Softmax operation on the linearly transformed global text feature vector to obtain the character probability distribution. After performing a character dictionary mapping on the linearly transformed global text feature vector according to the character probability distribution, decoding is achieved, and the recognized text recognition sequence is output. The number of text categories is specifically 5001, including 5000 common Chinese and English characters and a termination character EOS as a separate text category to mark the termination of the current string decoding. When decoding from left to right, until the last character is decoded or the first termination character EOS is decoded, slicing is performed and decoding is terminated to obtain the final text recognition sequence. The present invention does not require the use of complex autoregressive decoding. Compared with ordinary CTC decoding, since there is no blank character, the decoding efficiency is further improved.
[0073] Examples implemented according to the method of the present invention are as follows:
[0074] The training and test data sets used in this method are respectively sourced from the publicly available and general synthetic data sets SynthText (English), MJ-Synth (English), and SynthText-Chinese (Chinese) in the field of text recognition, with a total data volume of about 14 million. The text dictionary in this paper is set to 5001, including 5000 common Chinese and English characters and an end-of-string character EOS. The network model proposed in this paper is trained on a 2080Ti GPU, using the Adam optimizer and setting a batch size of 512 and 250,000 iteration rounds.
[0075] After using the text detection algorithm to intercept the text in two natural and complex scene images containing text, and then through scaling and zero-padding, images with a height of 32 and a width of 256 are obtained. After normalization, they are respectively as shown in Figure 6 (a) and (b), and then fed into the backbone convolutional neural network ( Figure 3 ) for feature extraction to obtain the visual feature map;
[0076] After passing the above visual feature map through the feature slicing module, multiple sliced local text features can be obtained, and they are respectively input into multiple local self-attention modules to obtain corresponding multiple local text feature aggregation vectors.
[0077] After splicing these local text feature aggregation vectors in the width dimension, the visualization effects of their feature distributions are respectively Figure 7 (a) and (b). Transforming and scaling the feature map to the size of the original image and performing normalization and visualization can respectively obtain the weight distribution maps of Figure 8 (a) and (b). Then, the above-spliced features are upsampled to 1 in height to obtain the global text feature.
[0078] Performing global self-attention operation on the global text feature to model its context relationship. The network structure diagram of this part is as shown in Figure 5 , and the global text feature vector can be obtained in this part.
[0079] The output after passing the above global text feature vector through the character decoding module is mapped through the text dictionary to obtain the finally recognized string.
Claims
1. A real-time irregular text recognition method in complex scenarios, characterized in that, Including the following steps: 1) After preprocessing the natural complex scene image containing text, a preprocessed text block image is obtained; 2) After inputting the preprocessed text block image into the backbone convolutional neural network for feature extraction and feature slicing, multiple sliced local text features are obtained and the serial numbers of each sliced local text feature are recorded; 3) The multiple sliced local text features are respectively input into multiple local self-attention modules according to the serial numbers corresponding to the sliced local text features. After each local self-attention module performs local text feature aggregation in parallel, local text feature aggregation vectors are respectively output; 4) The multiple local text feature aggregation vectors are simultaneously input into the global self-attention module for splicing, downsampling, and global text feature extraction, and then a global text feature vector is output; 5) The global text feature vector is input into the character decoding module for decoding, and then the recognized text string is output.
2. The real-time irregular text recognition method in complex scenarios according to claim 1, characterized in that, The specific content of step 1) is as follows: First, the natural complex scene image containing text is intercepted by using a text detection algorithm to obtain the original text block image. Then, the height of the original text block image is interpolated to the preset height H, and the width of the original text block image after adjusting the height is scaled according to the aspect ratio of the width and height of the original text block image to obtain the standard text block image. Finally, the standard text block image is normalized with a mean of 0 and a variance of 1 to obtain the preprocessed text block image.
3. The real-time irregular text recognition method in complex scenarios according to claim 1, characterized in that, The backbone convolutional neural network in step 2) is mainly composed of a convolutional downsampling module connected to a feature slicing module after passing through a 6-layer depth convolutional module, a first separable convolutional downsampling module, a 12-layer depth convolutional module, and a second separable convolutional downsampling module in sequence. The preprocessed text block image is input into the convolutional downsampling module, and the feature slicing module outputs multiple sliced local text features and records the serial numbers of each sliced local text feature.
4. The real-time irregular text recognition method in complex scenarios according to claim 3, characterized in that, The input of the feature slicing module is the visual feature map output by the second separable convolutional downsampling module. The feature slicing module determines the size of the sliding window according to the height of the visual feature map, uses the sliding window to slide pixels in the visual feature map, and after sliding, copies and slices the visual feature map covered by the sliding window as a sliced local text feature and outputs it. Continuous sliding slicing is performed to output multiple sliced local text features and record the serial numbers of each sliced local text feature.
5. The real-time irregular text recognition method in complex scenarios according to claim 1, characterized in that, The structures of the multiple local self-attention modules in step 3) are the same. Specifically: The local text features of the slice are formed by splicing multiple column vectors in the width dimension. Each local self-attention module linearly transforms the middle column vector of the input local text features of the slice after copying to obtain a local text query vector. Then, after cutting the column vectors of the input local text features of the slice and rearranging them in the height dimension, different linear transformations are performed on the rearranged local text features of the slice to obtain a local text key vector and a local text value vector respectively. After that, a matrix dot product is performed on the local text query vector and the local text key vector and then a Softmax operation is carried out to obtain a weight distribution matrix of the local attention mechanism. Then, a matrix dot product is performed on the weight distribution matrix of the local attention mechanism and the local text value vector to obtain a local text feature aggregation vector output by the current local self-attention module. The calculation formula is as follows: Ki = W K Ti r Vi = W V Ti r Among them, W Q , W K , W Q , W K , W V represent the first, second, and third local linear transformation parameter matrices respectively, represents a matrix of dimension C×C, represents the middle column vector of the i-th slice local text feature, H ′ is the height of the slice local text feature, Ti r represents the slice local text feature after column vector cutting and height dimension rearrangement of the i-th slice local text feature. Qi, Ki, and Vi are the local text query vector, local text key vector, and local text value vector corresponding to the i-th slice local text feature respectively. Ti out is the local text feature aggregation vector corresponding to the i-th slice local text feature. C is the feature dimension of the local text feature aggregation vector. T represents the transpose operation; Softmax(·) represents the Softmax operation, specifically the normalization process.
6. The real-time irregular text recognition method in complex scenarios according to claim 1, characterized in that, The specific content of step 4) is as follows: The global self-attention module splices multiple local text feature aggregation vectors in the width dimension according to the serial numbers of the corresponding local text features of the slice to obtain a text feature aggregation vector. Then, average pooling downsampling is performed on the text feature aggregation vector in the height dimension to obtain a global text feature sequence X. Different linear transformations are performed on the global text feature sequence to obtain a query vector, a key vector, and a value vector of the global text respectively. A matrix dot product is performed on the query vector and the key vector of the global text and then a Softmax operation is carried out to obtain a weight distribution matrix of the global attention mechanism. Then, a matrix dot product is performed on the weight distribution matrix of the global attention mechanism and the value vector of the global text to obtain a global text feature vector. The calculation formula is as follows: Among them, [·] represents the concatenation operation in the width dimension, and AveragePooling(·) represents the operation of mean pooling downsampling to 1 in the height dimension. represents the i-th local text feature aggregation vector, Q g , K g , V g are the query vector, key vector, and value vector of the global text. represents the first, second, and third global linear transformation parameter matrices, X out is the global text feature vector, C is the feature dimension of the global text feature vector, T represents the transpose operation; Softmax(·) represents the Softmax operation, specifically the normalization process.
7. The real-time irregular text recognition method in complex scenarios according to claim 1, characterized in that, The specific content of step 5) is as follows: In the character decoding module, first, a linear transformation is performed on the global text feature vector so that the feature dimension of the linearly transformed global text feature vector is equal to the number of character categories. Then, a Softmax operation is performed on the linearly transformed global text feature vector to obtain a character probability distribution. Decoding is realized by performing a character dictionary mapping on the linearly transformed global text feature vector according to the character probability distribution, and the recognized character recognition sequence is output.
Citation Information
Patent Citations
Natural scene character recognition method based on Transformer model
CN112149619A
Method and apparatus for recognizing text sequence, and storage medium
US20210232847A1