Text recognition method and text recognition system
Through the feature rotation and fusion of multi-layer convolutional neural network, the recognition difficulties caused by text image deformation in natural scenes are solved, and efficient text recognition is achieved.
Patent Information
- Application Number
- CN202011061835.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-30
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-09-30
AI Technical Summary
The existing text recognition algorithm has poor recognition effect when image deformation caused by camera shooting angle or non-rigidity of text carrier in natural scenes.
The first convolutional neural network is used for feature extraction, the feature map is rotated and feature extraction is performed through the second convolutional neural network. Combined with the weight fusion of the third convolutional neural network, feature fusion and decoding are realized to identify text.
Improve the recognition accuracy of deformed text images, and accurately recognize text in the image without preprocessing.
Smart Images

Figure CN114359679B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of image processing, and in particular, to a text recognition method and a text recognition system. Background Art
[0002] Currently, in the field of OCR (Optical Character Recognition), various text recognition technologies have emerged continuously, and most of the text recognition technologies with good performance are based on deep learning algorithms. Most text recognition algorithms have good recognition effects on text images without deformation. However, in natural scenes, due to reasons such as the camera shooting angle or the non-rigid text carrier, the text in the collected images often undergoes deformation such as tilting, perspective, and bending, which easily causes the text recognition algorithm to fail. Summary of the Invention
[0003] Embodiments of the present invention provide a text recognition method and a text recognition system for solving the problem of poor text recognition effect when the text image is deformed.
[0004] To solve the above technical problems, the present invention is implemented as follows:
[0005] In a first aspect, embodiments of the present invention provide a text recognition method, including:
[0006] Performing feature extraction on the image to be recognized by using a first convolutional neural network to obtain a first feature map of the image to be recognized;
[0007] Rotating the first feature map by N angles respectively to obtain N feature maps in different directions; performing feature extraction on the N feature maps in different directions by using a second convolutional neural network to obtain N feature vectors in different directions; processing the first feature map by using a third convolutional neural network to obtain weights of the N feature vectors in different directions respectively; where N is a positive integer greater than or equal to 2;
[0008] Performing feature fusion on the N feature vectors in different directions according to the weights of the N feature vectors in different directions respectively to obtain a one-dimensional feature vector after feature fusion, and decoding the one-dimensional feature vector after feature fusion to obtain a text recognition result.
[0009] Optionally, the first convolutional neural network is a convolutional neural network based on an attention mechanism, including a plurality of convolutional modules and a plurality of attention mechanism modules.
[0010] Optionally, the convolutional module is a first convolutional module or a second convolutional module. The first convolutional neural network includes at least two first convolutional modules and at least one second convolutional module, where the first and last convolutional modules of the first convolutional neural network are both the first convolutional modules.
[0011] Optionally, the second convolutional module includes: a first convolutional layer, a depthwise separable convolutional layer based on dilated convolution, and a second convolutional layer.
[0012] Optionally, the number of convolutional modules is less than 5.
[0013] Optionally, an attention mechanism module is provided after each convolutional module.
[0014] Optionally, before using the first convolutional neural network to extract features from the image to be recognized, it further includes:
[0015] Scaling the image to be recognized into a square image of a predetermined size.
[0016] Optionally, N is equal to 4, and the N angles are 0 degree, 90 degrees, 180 degrees, and 270 degrees respectively.
[0017] Optionally, the second convolutional neural network includes multiple convolutional layers and multiple pooling layers.
[0018] Optionally, after using the second convolutional neural network to extract features from the feature maps in the N directions respectively to obtain feature vectors in the N directions, it further includes:
[0019] Using a long short-term memory network to process the feature vectors in the N directions to obtain processed feature vectors in the N directions.
[0020] Optionally, the third convolutional neural network includes: M convolutional layers, M pooling layers, and a fully connected layer, where M is a positive integer greater than or equal to 1.
[0021] Optionally, decoding the feature vectors after feature fusion includes:
[0022] Using a long short-term memory network and an attention module to process the one-dimensional feature vectors after feature fusion to obtain processed one-dimensional feature vectors;
[0023] Using a Softmax layer to calculate the processed feature vectors to obtain a text recognition result.
[0024] In a second aspect, an embodiment of the present invention provides a text recognition system, including:
[0025] The first processing unit is used to extract features from the image to be recognized by using a first convolutional neural network, and obtain a first feature map of the image to be recognized;
[0026] The second processing unit is used to rotate the first feature map at N angles respectively to obtain feature maps in N directions; use a second convolutional neural network to extract features from the feature maps in the N directions respectively to obtain N direction feature vectors; use a third convolutional neural network to process the first feature map to obtain the weights of the N direction feature vectors respectively; where N is a positive integer greater than or equal to 2;
[0027] The third processing unit is used to perform feature fusion on the N direction feature vectors according to the weights of the N direction feature vectors respectively to obtain a feature vector after feature fusion, and decode the feature vector after feature fusion to obtain a text recognition result.
[0028] Optionally, the first convolutional neural network is a convolutional neural network based on an attention mechanism, and includes a plurality of convolutional modules and a plurality of attention mechanism modules.
[0029] Optionally, the convolutional module is a first convolutional module or a second convolutional module. The first convolutional module includes: a convolutional layer; the second convolutional module includes: a first convolutional layer, a depthwise separable convolutional layer based on dilated convolution, and a second convolutional layer.
[0030] Optionally, the first convolutional neural network includes at least two first convolutional modules and at least one second convolutional module, where the first and last convolutional modules of the first convolutional neural network are both the first convolutional modules.
[0031] Optionally, the number of convolutional modules is less than 5.
[0032] Optionally, one attention mechanism module is set after each convolutional module.
[0033] Optionally, the text recognition system further includes:
[0034] A scaling unit is used to scale the image to be recognized into a square image of a predetermined size.
[0035] Optionally, N is equal to 4, and the N angles are 0 degree, 90 degrees, 180 degrees, and 270 degrees respectively.
[0036] Optionally, the second convolutional neural network includes a plurality of convolutional layers and a plurality of pooling layers.
[0037] Optionally,
[0038] The second processing unit is configured to, after extracting feature vectors in N directions from the feature maps in the N directions by using a second convolutional neural network, process the feature vectors in the N directions by using a long short-term memory network to obtain processed feature vectors in the N directions.
[0039] Optionally, the third convolutional neural network includes: M convolutional layers, M pooling layers, and a fully connected layer, where M is a positive integer greater than or equal to 1.
[0040] Optionally, the third processing unit is configured to process the feature vectors after feature fusion by using a long short-term memory network and an attention module to obtain processed feature vectors; and use a Softmax layer to calculate the processed feature vectors to obtain a text recognition result.
[0041] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the text recognition method in the first aspect are implemented.
[0042] In a fourth aspect, an embodiment of the present invention provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the text recognition method in the first aspect are implemented.
[0043] In the embodiment of the present invention, a directional convolutional neural network can extract text features and position features of feature maps in N directions. Through feature fusion and decoding, text recognition is completed, and the recognition accuracy for various deformation problems of text in an image is improved. Whether the input image to be recognized is deformed or not, the text in the image to be recognized can be accurately recognized, and it is not necessary to perform stretching, rotation, etc. on the detected text before text recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0045] Figure 1 is a schematic flowchart of the text recognition method according to the embodiment of the present invention;
[0046] Figure 2 is a schematic structural diagram of the text recognition system according to an embodiment of the present invention;
[0047] Figure 3Schematic diagram of the structure of the first convolutional module according to an embodiment of the present invention;
[0048] Figure 4 Schematic diagram of the structure of the second convolutional module according to an embodiment of the present invention;
[0049] Figure 5 Schematic diagram of the structure of the attention mechanism module according to an embodiment of the present invention;
[0050] Figure 6 Schematic diagram of the structure of the text recognition system according to another embodiment of the present invention;
[0051] Figure 7 Schematic diagram of the structure of the electronic device according to an embodiment of the present invention. Detailed implementation manners
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0053] Please refer to Figure 1 , an embodiment of the present invention provides a text recognition method, including:
[0054] Step 11: Use a first convolutional neural network to extract features from the image to be recognized, and obtain a first feature map of the image to be recognized;
[0055] Step 12: Rotate the first feature map at N angles respectively to obtain feature maps in N directions; use a second convolutional neural network to extract features from the feature maps in the N directions respectively to obtain feature vectors in the N directions; use a third convolutional neural network to process the first feature map to obtain the weights of the feature vectors in the N directions respectively; where N is a positive integer greater than or equal to 2;
[0056] The second convolutional neural network can also be called a direction convolutional neural network, which is used to extract feature vectors from the input feature map.
[0057] Step 13: According to the weights of the feature vectors in the N directions respectively, perform feature fusion on the feature vectors in the N directions to obtain a feature vector after feature fusion, and decode the feature vector after feature fusion to obtain a text recognition result.
[0058] In the embodiment of the present invention, the feature vector after feature fusion can be a one-dimensional feature vector.
[0059] In an embodiment of the present invention, the second convolutional neural network outputs feature vectors in N directions, and the third convolutional neural network outputs N weights. The feature vectors in the N directions correspond one-to-one to the N weights. Before feature fusion, it is necessary to match the feature vectors in the N directions with the N weights one-to-one.
[0060] In an embodiment of the present invention, a directional convolutional neural network can be used to extract the text features and position features of feature maps in N directions. Through feature fusion and decoding, the recognition of text is completed, improving the recognition accuracy for various deformation problems of text in images. Whether the input image to be recognized is deformed or not, the text in the image to be recognized can be accurately recognized, and it is not necessary to perform stretching, rotation, etc. on the detected text before text recognition.
[0061] In an embodiment of the present invention, optionally, before performing the above step 11, it further includes:
[0062] Step 10: Scale the image to be recognized into an image of a unified predetermined size.
[0063] The above three steps will be described in detail below.
[0064] Step 11:
[0065] In an embodiment of the present invention, optionally, the first convolutional neural network is a convolutional neural network based on an attention mechanism. A convolutional neural network using an attention mechanism can improve the problem that a traditional convolutional neural network only focuses on local image information in the image due to the limitation of the convolutional kernel and ignores the influence of global image information on the current position. It can enhance the role of global image information in the feature extraction process, improve the ability to judge similar characters, thereby improving the recognition accuracy, reducing the calculation amount, and accelerating the operation efficiency.
[0066] In an embodiment of the present invention, the convolutional neural network based on an attention mechanism includes: a plurality of convolutional modules and a plurality of attention mechanism modules (Convolutional Block Attention Module, CBAM). Optionally, one attention mechanism module is provided after each convolutional module.
[0067] Optionally, the number of convolutional modules is less than 5, so that the features of the image to be recognized extracted by the convolutional neural network based on the attention mechanism are low-level visual features. The low-level visual features include, for example, at least one of the following: color, texture, position, size, etc.
[0068] Optionally, the convolution module is the first convolution module (Convolution) or the second convolution module (Convblock). The first convolution module includes a convolutional layer, and the second convolution module includes: a first convolutional layer, a depthwise separable convolutional layer based on dilated convolution, and a second convolutional layer. The second convolution module can perform more complex processing. Using the second convolution module for convolution processing can be applicable to the text recognition of images collected in natural scenes.
[0069] Optionally, the first convolutional neural network includes at least two first convolution modules and at least one second convolution module, wherein the first and / or the last convolution module of the first convolutional neural network are both the first convolution module.
[0070] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of the text recognition system according to an embodiment of the present invention. In this embodiment, the convolutional application network based on the attention mechanism includes, in cascade: convolutional layer_1, attention mechanism module, convolutional block_2, attention mechanism module, max pooling layer (Maxpooling), convolutional block_3, attention mechanism module, max pooling layer, convolutional layer_4, attention mechanism module. That is, it includes 4 convolution modules (the first convolution module + the second convolution module), so as to extract the low-level visual features of the image to be recognized. An attention mechanism module is arranged after each convolution module (the first convolution module + the second convolution module) to extract the global image information of the image to be recognized.
[0071] Among them, referring to Figure 3 , in the embodiment of the present invention, the first convolution module includes, in cascade: a convolutional layer and BatchNormal (batch normalization) + Relu (rectified linear unit).
[0072] Among them, referring to Figure 4 , in the embodiment of the present invention, the second convolution module includes, in cascade: a convolutional layer, BatchNormal + Relu, SeparableConv (depthwise separable convolution) + DilationConv (dilated convolution), Batch Normal + Relu, a convolutional layer, Batch Normal + Relu.
[0073] Among them, referring to Figure 5, in the embodiment of the present invention, the attention mechanism module includes: a channel attention module and a spatial attention module. Among them, the processing process of the attention mechanism module for the input features is as follows: the input features are first processed by the channel attention module to obtain the first processed feature, then the first feature and the input features are subjected to a cross product process to obtain the second processed feature, the second feature is input into the spatial attention module for processing to obtain the third processed feature, and the third feature and the second feature are subjected to a cross product process to obtain the output feature.
[0074] Step 12:
[0075] Please refer to Figure 2 , in the embodiment of the present invention, optionally, N is equal to 4, and the N angles are 0 degrees, 90 degrees, 180 degrees, and 270 degrees respectively. That is to say, the input first feature map is rotated 0 degrees, 90 degrees, 180 degrees, and 270 degrees respectively to obtain feature maps in 4 directions.
[0076] In the embodiment of the present invention, optionally, before using the first convolutional neural network to extract features from the image to be recognized, it further includes: scaling the image to be recognized into a square image of a predetermined size. Using a square image as the input enables the first feature map not to require pixel padding when rotating by 0 degrees, 90 degrees, 180 degrees, and 270 degrees. In the example of the present invention, the method of scaling the image to be recognized into a square image of a predetermined size can be stretching the image to be recognized, or filling the edge pixels, etc.
[0077] In the embodiment of the present invention, optionally, the second convolutional neural network includes multiple convolutional layers and multiple pooling layers. Optionally, the convolutional layers and the pooling layers are arranged alternately. Please refer to Figure 2 , Figure 2 In the shown embodiment, the directional convolutional neural network (the second convolutional neural network) includes four convolutional layers and four max pooling layers.
[0078] In the embodiment of the present invention, optionally, after using the second convolutional neural network to extract features from the feature maps in the N directions respectively to obtain feature vectors in the N directions, it further includes: using N long short-term memory networks (Long Short-Term Memory, LSTM) to process the feature vectors in the N directions respectively to obtain the processed feature vectors in the N directions. Further optionally, the long short-term memory network can be a bi-directional long short-term memory network (Bi-directional Long Short-Term Memory, BiLSTM).
[0079] In the embodiments of the present invention, while using a second convolutional neural network to extract features from feature maps in N directions respectively, a third convolutional neural network is also used to process the input first feature map to obtain the weights of the feature vectors in N directions respectively, so as to screen the directional features according to these weights and realize the judgment of the text direction.
[0080] Optionally, the third convolutional neural network includes: M convolutional layers, M pooling layers and a fully connected layer, optionally, the M convolutional layers and the M pooling layers are alternately arranged, and the fully connected layer is located at the end, where M is a positive integer greater than or equal to 1.
[0081] Please refer to Figure 2 , Figure 2 In the embodiment shown, the third convolutional neural network includes: two convolutional layers, two max pooling layers and two fully connected layers, the two convolutional layers and the two max pooling layers are alternately arranged, and the two fully connected layers are located at the end.
[0082] Step 13:
[0083] In the embodiments of the present invention, optionally, the method of feature fusion may be: multiplying the N feature vectors by their respective weights to obtain N products, and then adding the N products to obtain the feature vector after feature fusion.
[0084] Of course, other feature fusion methods may also be used. For example, multiplying the N feature vectors by their respective weights to obtain N products, and then inputting the N products into a convolutional module for convolutional processing to obtain the feature vector after feature fusion.
[0085] In the embodiments of the present invention, optionally, decoding the feature vector after feature fusion includes:
[0086] Step 131: Using a long short-term memory network and an attention module to process the feature vector after feature fusion to obtain a processed feature vector;
[0087] Optionally, the long short-term memory network may be a bidirectional long short-term memory network.
[0088] Optionally, the attention module may be the above-mentioned attention mechanism module.
[0089] This step further extracts features from the feature vector after feature fusion, so as to optimize the text recognition result.
[0090] Step 132: Using a Softmax layer to calculate the processed feature vector to obtain the text recognition result.
[0091] Please refer to Figure 6, an embodiment of the present invention further provides a text recognition system 60, including:
[0092] A first processing unit 61, configured to extract features from the image to be recognized by using a first convolutional neural network, and obtain a first feature map of the image to be recognized;
[0093] A second processing unit 62, configured to rotate the first feature map by N angles respectively to obtain feature maps in N directions; extract features from the feature maps in the N directions by using a second convolutional neural network respectively to obtain feature vectors in the N directions; process the first feature map by using a third convolutional neural network to obtain weights of the feature vectors in the N directions respectively; where N is a positive integer greater than or equal to 2;
[0094] A third processing unit 63, configured to perform feature fusion on the feature vectors in the N directions according to the weights of the feature vectors in the N directions respectively to obtain a feature vector after feature fusion, and decode the feature vector after feature fusion to obtain a text recognition result.
[0095] In the embodiment of the present invention, a directional convolutional neural network can be used to extract text features and position features of feature maps in N directions. Through feature fusion and decoding, text recognition is completed, and the recognition accuracy of various deformation problems of text in the image is improved. Whether the input image to be recognized is deformed or not, the text in the image to be recognized can be accurately recognized, and it is not necessary to perform stretching, rotation and other processing on the detected text before text recognition.
[0096] Optionally, the first convolutional neural network is a convolutional neural network based on an attention mechanism, including a plurality of convolutional modules and a plurality of attention mechanism modules.
[0097] Optionally, the convolutional module is a first convolutional module or a second convolutional module. The first convolutional module includes: a convolutional layer; the second convolutional module includes: a first convolutional layer, a depthwise separable convolutional layer based on dilated convolution, and a second convolutional layer.
[0098] Optionally, the first convolutional neural network includes at least two first convolutional modules and at least one second convolutional module, where the first and last convolutional modules of the first convolutional neural network are both the first convolutional modules.
[0099] Optionally, the number of convolutional modules is less than 5.
[0100] Optionally, the text recognition system further includes:
[0101] A scaling unit, configured to scale the image to be recognized into a square image of a predetermined size.
[0102] Optionally, N equals 4, and the N angles are 0 degrees, 90 degrees, 180 degrees, and 270 degrees respectively.
[0103] Optionally, the second convolutional neural network includes a plurality of convolutional layers and a plurality of pooling layers.
[0104] Optionally, the convolutional layers and the pooling layers are arranged alternately.
[0105] Optionally, after the second processing unit uses the second convolutional neural network to extract features from the feature maps in the N directions to obtain feature vectors in the N directions, it uses a long short-term memory network to process the feature vectors in the N directions to obtain processed feature vectors in the N directions.
[0106] Optionally, the third convolutional neural network includes: M convolutional layers, M pooling layers, and a fully connected layer, where M is a positive integer greater than or equal to 1.
[0107] Optionally, the M convolutional layers and the M pooling layers are arranged alternately, and the fully connected layer is located at the end.
[0108] Optionally, the third processing unit is used to process the feature vectors after feature fusion using a long short-term memory network and an attention module to obtain processed feature vectors; and use a Softmax layer to calculate the processed feature vectors to obtain a text recognition result. Please refer to Figure 7 , an embodiment of the present invention further provides an electronic device 70, including a processor 71, a memory 72, and a computer program stored on the memory 72 and executable on the processor 71. When the computer program is executed by the processor 71, it implements each process of the above text recognition method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described here again.
[0109] An embodiment of the present invention further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements each process of the above text recognition method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described here again. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0110] It should be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element.
[0111] From the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0112] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims, and all of them fall within the protection scope of the present invention.
Claims
1. A text recognition method, characterized in that, it includes: using a first convolutional neural network to extract features from the image to be recognized, obtaining a first feature map of the image to be recognized; rotating the first feature map at N angles respectively to obtain feature maps in N directions; using a second convolutional neural network to extract features from the feature maps in the N directions respectively, obtaining N direction feature vectors; using a third convolutional neural network to process the first feature map to obtain the weights of the feature vectors in the N directions respectively; where N is a positive integer greater than or equal to 2; according to the weights of the feature vectors in the N directions respectively, performing feature fusion on the feature vectors in the N directions to obtain a feature vector after feature fusion, and decoding the feature vector after feature fusion to obtain a text recognition result.
2. The text recognition method according to claim 1, characterized in that, the first convolutional neural network is a convolutional neural network based on an attention mechanism, including a plurality of convolutional modules and a plurality of attention mechanism modules.
3. The text recognition method according to claim 2, characterized in that, the convolutional module is a first convolutional module or a second convolutional module, the first convolutional neural network includes at least two first convolutional modules and at least one second convolutional module, where the first and last convolutional modules of the first convolutional neural network are both the first convolutional modules.
4. The text recognition method according to claim 3, the second convolutional module includes: a first convolutional layer, a depthwise separable convolutional layer based on dilated convolution, and a second convolutional layer.
5. The text recognition method according to claim 2, characterized in that, the number of convolutional modules is less than 5.
6. The text recognition method according to any one of claims 2-5, characterized in that, one attention mechanism module is provided after each convolutional module.
7. The text recognition method according to claim 1, characterized in that, before using the first convolutional neural network to extract features from the image to be recognized, it further includes: scaling the image to be recognized into a square image of a predetermined size.
8. The text recognition method according to claim 1, characterized in that, N is equal to 4, and the N angles are 0 degree, 90 degrees, 180 degrees, and 270 degrees respectively.
9. The text recognition method according to claim 1, characterized in that, the second convolutional neural network includes a plurality of convolutional layers and a plurality of pooling layers.
10. The text recognition method according to claim 1, characterized in that, after using the second convolutional neural network to extract features from the feature maps in the N directions respectively to obtain N direction feature vectors, it further includes: using a long short-term memory network to process the N direction feature vectors to obtain processed N direction feature vectors.
11. The text recognition method according to claim 1, characterized in that, the third convolutional neural network includes: M convolutional layers, M pooling layers, and a fully connected layer, where M is a positive integer greater than or equal to 1.
12. The text recognition method according to claim 1, characterized in that, Decoding the one-dimensional feature vector after the feature fusion includes: Processing the feature vector after the feature fusion by using a long short-term memory network and an attention module to obtain a processed feature vector; Calculating the processed feature vector by using a Softmax layer to obtain a text recognition result.
13. A text recognition system Characterized in that It includes: A first processing unit, configured to perform feature extraction on an image to be recognized by using a first convolutional neural network to obtain a first feature map of the image to be recognized; A second processing unit, configured to rotate the first feature map by N angles respectively to obtain feature maps in N directions; performing feature extraction on the feature maps in the N directions respectively by using a second convolutional neural network to obtain feature vectors in the N directions; processing the first feature map by using a third convolutional neural network to obtain weights of the feature vectors in the N directions respectively; where N is a positive integer greater than or equal to 2; A third processing unit, configured to perform feature fusion on the feature vectors in the N directions according to the weights of the feature vectors in the N directions respectively to obtain a feature vector after the feature fusion, and decoding the feature vector after the feature fusion to obtain a text recognition result.
14. An electronic device Characterized in that It includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor, and when the program or instruction is executed by the processor, the steps of the text recognition method according to any one of claims 1 to 12 are implemented.
15. A readable storage medium Characterized in that A program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the text recognition method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Image super-resolution reconstruction method based on fused attention mechanism residual network
CN111192200A
Text detection and recognition method and system for natural scene
CN111340034A