OCR-Based Text Recognition Method, Device, Storage Medium, and Electronic Device
By jointly training the text recognition network and the super-resolution network, and using the shared subnet to adjust the parameters, the problem of low-quality text image recognition is solved, and the effect of improving text recognition accuracy without increasing the inference time is achieved.
Patent Information
- Application Number
- CN202210864937.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-21
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-07-21
AI Technical Summary
Existing text recognition methods have low recognition accuracy when processing low-quality text images, making it difficult to effectively identify text in low-resolution images.
A text recognition method based on OCR is proposed. By obtaining text image sample sets and super-resolution image samples, the pre-constructed text recognition network and super-resolution network are trained, and the shared subnet is used to adjust network parameters during the training process to improve the accuracy of text recognition.
When the super-resolution network branch is removed during the inference stage and the inference time remains unchanged, the recognition accuracy and recognition effect of the text recognition network for low-quality text images is effectively improved.
Smart Images

Figure CN115188000B_ABST
Abstract
Description
Technical Field
[0002] The present invention relates to the technical field of image processing, and in particular, to an OCR-based text recognition method, device, storage medium, and electronic device.
Background Art
[0004] Computer vision is a science that studies how to enable machines to "see". Further, it refers to machine vision that uses cameras and computers to replace human eyes to identify, track, and measure targets, etc., and further performs graphic processing to make the computer process into images that are more suitable for human eyes to observe or transmitted to instruments for detection.
[0005] OCR (Optical Character Recognition) is a classic topic in the field of computer vision and is widely used in fields such as driverless, road sign recognition, license plate recognition, and taking pictures to search for questions in the education scenario. OCR refers to the process in which an electronic device (such as a scanner or digital camera) checks the characters printed on paper, determines their shapes by detecting dark and bright patterns, and then translates the shapes into computer text using character recognition methods; that is, for printed characters, an optical method is used to convert the text in a paper document into a black and white dot matrix image file, and the text in the image is converted into a text format by an identification software for further editing and processing by a word processing software. Different from text recognition in a computer, the text images to be recognized in the OCR scenario often contain a large number of low-quality images (mainly referring to low-resolution images), and existing text recognition methods are difficult to effectively recognize low-quality text images, and the recognition accuracy is relatively low.
Summary of the Invention
[0007] The present invention proposes an OCR-based text recognition method, device, storage medium, and electronic device, which can improve the accuracy of text recognition and has a good recognition effect.
[0008] On the one hand, an embodiment of the present invention provides an OCR-based text recognition method, including:
[0009] Obtaining a text image sample set, as well as the text label and super-resolution image sample corresponding to each text image sample in the text image sample set;
[0010] Using the text image sample set, the text label, and the super-resolution image sample to train a pre-constructed text recognition network and super-resolution network, where the text recognition network includes a first loss function, the super-resolution network includes a second loss function, and the text recognition network and the super-resolution network include at least one shared sub-network;
[0011] During the training process, the network parameters of the text recognition network and the super-resolution network are adjusted according to the first loss function and the second loss function;
[0012] When the training is completed, the trained text recognition network is used to recognize the text in the text image to be recognized.
[0013] On the other hand, an embodiment of the present invention further provides an OCR-based text recognition device, including:
[0014] An acquisition unit, configured to acquire a text image sample set, as well as the text label and the super-resolution image sample corresponding to each text image sample in the text image sample set;
[0015] A training unit, configured to use the text image sample set, the text label, and the super-resolution image sample to train a pre-constructed text recognition network and a super-resolution network. The text recognition network includes a first loss function, the super-resolution network includes a second loss function, and the text recognition network and the super-resolution network include at least one shared sub-network; during the training process, the network parameters of the text recognition network and the super-resolution network are adjusted according to the first loss function and the second loss function;
[0016] An identification unit, configured to, when the training is completed, use the trained text recognition network to recognize the text in the text image to be recognized.
[0017] In some embodiments, the text recognition network includes a connected feature extraction sub-network and a feature recognition sub-network, and the super-resolution network includes a connected feature extraction sub-network and a super-resolution sub-network. The training unit is specifically configured to:
[0018] Determine the feature map corresponding to each text image sample through the feature extraction sub-network;
[0019] Generate a predicted image result corresponding to the feature map through the super-resolution sub-network;
[0020] Generate a predicted text result corresponding to the feature map through the feature recognition sub-network;
[0021] Adjust the parameters of the text recognition network and the super-resolution network according to the predicted image result, the predicted text result, the text label, the super-resolution image sample, the first loss function, and the second loss function.
[0022] In some embodiments, the training unit is further configured to:
[0023] Calculate a first error value according to the first loss function, the predicted text result, and the text label;
[0024] Calculate a second error value according to the second loss function, the predicted image result, and the super-resolution image sample;
[0025] Use the formula \(L = L\) rec +\(\lambda L\) sr to calculate the total error value, where \(L\) is the total error value, \(L\) rec is the first error value, \(L\) sr is the second error value, and \(\lambda\) is a hyperparameter;
[0026] Backward adjust the network parameters of the text recognition network and the super-resolution network according to the total error value.
[0027] In some embodiments, the feature extraction sub-network includes a first feature extraction block, a plurality of cascaded residual blocks, and a feature enhancement block, and the training unit is further configured to:
[0028] Determine a first shallow feature map corresponding to each text image sample through the first feature extraction block;
[0029] Process the first shallow feature map through the plurality of residual blocks;
[0030] Through the feature enhancement block, obtain the residual feature maps output after the processing of each residual block, and respectively downsample the first shallow feature map and the residual feature maps to obtain corresponding downsampled feature maps, and then perform channel fusion on all the downsampled feature maps to obtain the feature map corresponding to the text image sample.
[0031] In some embodiments, both the text recognition network and the super-resolution network further include a text correction sub-network connected to the feature extraction sub-network, and the training unit is further configured to:
[0032] Determine multiple key point information on each text image sample through the text correction sub-network, and correct the text image sample according to a preset interpolation algorithm and the key point information to obtain a corresponding corrected image;
[0033] The step of determining a first shallow feature map corresponding to each text image sample through the first feature extraction block specifically includes: performing shallow feature extraction on each corrected image through the feature extraction sub-network to obtain a first shallow feature map.
[0034] In some embodiments, the super-resolution sub-network includes a second feature extraction block, a plurality of cascaded sequence residual blocks, and a pixel recombination block. The training unit is further configured to:
[0035] Generate a binary map corresponding to the text image sample;
[0036] Perform channel fusion on the feature map and the binary map to generate a fused feature map;
[0037] Determine a second shallow feature map corresponding to the fused feature map through the second feature extraction block;
[0038] Process the second shallow feature map through the sequence residual block to obtain a deep feature map;
[0039] Perform pixel recombination on the deep feature map and the second shallow feature map through the pixel recombination block to obtain a corresponding predicted image result.
[0040] In some embodiments, the super-resolution sub-network further includes a center alignment block. Before determining the second shallow feature map corresponding to the fused feature map through the second feature extraction block, the training unit is further configured to:
[0041] Generate an aligned feature map corresponding to the fused feature map through the center alignment block;
[0042] The training unit is specifically configured to: perform shallow feature extraction from the aligned feature map through the second feature extraction block to obtain a second shallow feature map.
[0043] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded by a processor to execute the OCR-based text recognition method described in any one of the above.
[0044] On the other hand, an embodiment of the present invention further provides an electronic device, including a coupled memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program in the memory to execute the steps in the OCR-based text recognition method described in any one of the above.
[0045] The text recognition method, device, storage medium and electronic device based on OCR provided by the embodiments of the present invention obtain a text image sample set, as well as the text label and super-resolution image sample corresponding to each text image sample in the text image sample set. Then, using the text image sample set, text labels and super-resolution image samples, the pre-constructed text recognition network and super-resolution network are trained. Among them, the text recognition network includes a first loss function, the super-resolution network includes a second loss function, and the text recognition network and the super-resolution network include at least one shared sub-network. At the same time, during the training process, according to the first loss function and the second loss function, the network parameters of the text recognition network and the super-resolution network are adjusted. When the training is completed, the trained text recognition network is used to perform text recognition on the text image to be recognized. Therefore, only a super-resolution branch is added for feature learning in the training stage, and this branch is removed in the inference stage, so that the inference time remains unchanged, and the recognition accuracy and recognition effect of the text recognition network for low-quality text images are effectively improved.
Description of the Drawings
[0047] In order to more clearly illustrate the embodiments of the present invention or related technologies, the following drawings will be described when briefly introducing the embodiments. Obviously, the drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without any creative work.
[0048] Figure 1 It is a schematic flowchart of the text recognition method based on OCR provided by the embodiments of the present invention.
[0049] Figure 2 It is a schematic flowchart of the working processes of the text recognition network and the super-resolution network in the training stage and the inference stage provided by the embodiments of the present invention.
[0050] Figure 3 It is a schematic flowchart of the specific process of step S102 provided by the embodiments of the present invention.
[0051] Figure 4 It is a schematic flowchart of the working processes of the text recognition network and the super-resolution network in the training stage provided by the embodiments of the present application.
[0052] Figure 5 It is a schematic structural diagram of a single sequence residual block provided by the embodiments of the present invention.
[0053] Figure 6 It is a schematic structural diagram of the text recognition device based on OCR provided by the embodiments of the present application.
[0054] Figure 7 It is a schematic structural diagram of the electronic device provided by the embodiments of the present application.
[0055] Figure 8 It is another structural schematic diagram of the electronic device provided by the embodiment of the present application.
Specific Embodiments
[0057] Next, with reference to the accompanying drawings and embodiments, the present invention will be further described in detail. It should be specifically noted that the following embodiments are only used to illustrate the present invention, but do not limit the scope of the present invention. Similarly, the following embodiments are only partial embodiments of the present invention rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0058] In the description herein, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. Unless otherwise clearly specified in the context, the singular forms "a" and "an" used herein are also intended to include the plural. The meaning of "a plurality" is two or more. It should also be understood that the terms "comprising" and / or "including" used herein specify the presence of the stated features, integers, steps, operations, units, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, units, components, and / or combinations thereof.
[0059] The embodiment of the present invention provides an OCR-based text recognition method, device, storage medium, and electronic device.
[0060] Please refer to Figure 1 , Figure 1 which is a flowchart of the OCR-based text recognition method provided by the embodiment of the present invention. This text recognition method is applied to an electronic device, which may include a terminal device or a server with text recognition capabilities. Specifically, the text recognition method includes the following steps S101-S104, where:
[0061] S101. Obtain a text image sample set, as well as the text label and super-resolution image sample corresponding to each text image sample in the text image sample set.
[0062] Among them, the text image sample is a low-resolution RGB image, and the super-resolution image sample is a high-resolution RGB image. The text label refers to the label value of each pixel point on the text image sample. For example, when the actual text content included in the text image sample is "same", the label value of the pixel point where "same" is located can be 1, and the label values of the remaining pixel points can be 0.
[0063] A high-resolution camera and a low-resolution camera, or the same camera with different focal lengths, can be used to capture the same text file to obtain a text image sample and a super-resolution image sample respectively. Alternatively, after obtaining a super-resolution image sample by using a high-resolution camera, a text image sample can be obtained by blurring the super-resolution image sample, such as mean filtering, Gaussian filtering, etc.
[0064] S102. Use the text image sample set, the text label, and the super-resolution image sample to train a pre-constructed text recognition network and a super-resolution network. The text recognition network includes a first loss function, the super-resolution network includes a second loss function, and the text recognition network and the super-resolution network include at least one shared sub-network.
[0065] Among them, the text recognition network is used to identify the text content included in the image, and the super-resolution network is used to improve the resolution of the image to obtain a super-resolution image. Both the text recognition network and the super-resolution network are composed of multiple sub-networks. By making at least one sub-network shared by both, the training of the text recognition network and the training of the super-resolution network are associated. That is, when training one of the networks, such as the super-resolution network, the change of the network parameters in the shared part will inevitably affect the other network, such as the training result of the text recognition network.
[0066] S103. During the training process, adjust the network parameters of the text recognition network and the super-resolution network according to the first loss function and the second loss function.
[0067] Among them, the process of network training is also the process of continuously adjusting network parameters in an iterative manner through the loss function.
[0068] In some embodiments, please refer to Figure 2 , Figure 2 shows a schematic diagram of the working processes of the text recognition network and the super-resolution network in the training stage and the inference stage. Specifically, the text recognition network may include a connected feature extraction sub-network and a feature recognition sub-network, and the super-resolution network may include a connected feature extraction sub-network and a super-resolution sub-network. At this time, please refer to Figure 3 , the above step S102 may specifically include the following steps S1021 - S1023, where:
[0069] S1021. Determine the feature map corresponding to each text image sample through the feature extraction sub-network.
[0070] Among them, the feature extraction sub-network is mainly used to extract visual features. It is a sub-network shared by the text recognition network and the super-resolution network, while the feature recognition sub-network and the super-resolution sub-network are unique parts of these two networks respectively. The feature extraction sub-network is mainly used to extract visual features. In training stage ①, when the feature extraction sub-network extracts the feature map of the low-resolution text image sample, this feature map will enter two branches for processing. One branch is based on this feature map, and the super-resolution prediction image result is generated through the super-resolution sub-network. The other branch is based on this feature map, and the predicted text result is generated through the feature recognition sub-network.
[0071] For example, in the above Figure 2 , for the low-resolution text image sample P1 with the text content "serve", in training stage ①, when the feature map is obtained by processing this text image sample through the shared feature extraction sub-network, this feature Figure 1 will be processed by the feature recognition sub-network to obtain the text "serve", and on the other hand, it will be processed by the super-resolution network to obtain the super-resolution image P2.
[0072] In some embodiments, the feature extraction sub-network may use the structure of a residual network (ResNet) as the backbone network, which may include a first feature extraction block, a plurality of cascaded residual blocks, and a feature enhancement block. At this time, step S1021 above may specifically include:
[0073] Determine the first shallow feature map corresponding to each text image sample through this first feature extraction block;
[0074] Process this first shallow feature map through this plurality of residual blocks;
[0075] Obtain the residual feature map output after the processing of each residual block through this feature enhancement block, and respectively downsample this first shallow feature map and this residual feature map to obtain the corresponding downsampled feature maps, and then perform channel fusion on all these downsampled feature maps to obtain the feature map corresponding to this text image sample.
[0076] Among them, the first feature extraction block, the residual block (RB), and the feature enhancement block can all include convolutional groups. Each convolutional group contains at least one basic convolutional calculation process. The number of convolutional calculations in different convolutional groups depends on the number of convolutional kernels. The number of convolutional kernels in different convolutional groups can be the same or different. A convolutional kernel is a matrix used for convolutional calculation, and convolutional calculation is a series of operation operations performed on each pixel in an image using a convolutional kernel. The first feature extraction block, the residual block, and the feature enhancement block can be cascaded in sequence. The cascading order mainly indicates the connection order between these modules. Other connection modules can also be inserted between two adjacent cascaded modules, which is not restricted here.
[0077] The feature enhancement block can include a downsampling block and a channel fusion block. A downsampling block can be connected to the output of the first feature extraction block and each residual block. The downsampling block can perform downsampling of a preset multiple, such as 4 times downsampling, on the output feature maps (the first shallow feature map and the residual feature map) at the output. Then, all the downsampling blocks are connected to a channel fusion block to perform channel fusion on the sampled feature maps after downsampling of different output layers through the channel fusion block.
[0078] Since the number of convolutional kernels in different convolutional groups is the same as their output channel numbers (dimensions), such as 64, 128, 256, etc., and the number of output channel numbers is the same as the number of output feature maps, and the larger the number of output channels of a convolutional group, the smaller the image size of its output feature map. For example, if the number of output channels of a convolutional group is 64 and the number of output channels of the next convolutional group is 128, then the size of the output feature map processed by the convolutional group with 128 output channels is half of the size of the output feature map processed by the convolutional group with 64 output channels. Therefore, the channel fusion here is mainly to unify the sampled feature maps of different sizes and dimensions to obtain a feature map with the expected dimension and size. For example, finally, a feature map of 8×25×992 is obtained, where 992 represents the dimension and 8×25 represents the image size (height × width).
[0079] It should be noted that since the structure of the residual network uses a direct connection (shortcut connection), for example, a direct connection is made between two residual blocks at intervals, so the input information can bypass through the direct connection to the output, which can better protect the original features of the input and the integrity of the input information (that is, the residual network includes a direct mapping part). At the same time, since the output of the previous residual block will be used as the input of the next residual block between two adjacent cascaded residual blocks, the deep features of the input information can be further extracted (that is, the residual network includes a residual part), and the entire residual network only needs to learn the part of the difference between the input and the output, which simplifies the learning objective and difficulty. At the same time, by adding a feature enhancement block to process and obtain the final output feature map, compared with the scheme of directly using the output feature map of the last convolutional group as the final output feature map, the addition of the feature enhancement block can make the final output feature map obtain diversified semantic information, which is beneficial to improving the accuracy of subsequent text recognition.
[0080] In some embodiments, when capturing a text image in a natural scene, it will be affected by many factors, resulting in the text arrangement on the captured text image not necessarily being horizontally arranged, and there are various deformations, such as text bending and perspective. To improve the accuracy of subsequent text recognition, these deformations can be corrected in advance.
[0081] For example, please refer to Figure 4 , Figure 4 is a schematic diagram of the working process of the text recognition network and the super-resolution network in the training stage provided by the embodiments of the present application. Among them, both the text recognition network and the super-resolution network can further include a text correction sub-network connected to the feature extraction sub-network. At this time, before the above step S1021, the OCR-based text recognition method can further include:
[0082] Through the text correction sub-network, determine the key point information of multiple positions on each text image sample, and according to a preset interpolation algorithm and the key point information, correct the text image sample to obtain a corresponding corrected image;
[0083] Determining the feature map corresponding to each text image sample through the feature extraction sub-network specifically includes: generating the feature map corresponding to each corrected image through the feature extraction sub-network.
[0084] Among them, the text correction sub-network is also a sub-network shared by the text recognition network and the super-resolution network, and its main function is to correct the text into a horizontal arrangement. The TPS (thin plate spline transformation) interpolation method can be used for correction, and this correction method has a good correction effect on these two types of deformed texts, namely perspective and bending.
[0085] Specifically, N key points A on the text image sample can be predicted first through a trained convolutional neural network n , such as the positions of 20 key points. The upper and lower edges of the text are constrained by the positions of these key points. Then, through the TPS interpolation method, these N key points A n are deformed into the corresponding N points B n . In this process, an interpolation method that minimizes the thin plate bending energy is used to calculate new pixel values and perform pixel filling, and finally a new text image (corrected image) is generated to perform a flexible transformation on the original text image sample. For example, in the above Figure 4 , when the text "serve" on the text image sample is a curved and deformed text, after being processed by the text correction sub-network, the text on the corrected image becomes the horizontal text "serve". In other embodiments, the text correction sub-network can also adopt other horizontal arrangement correction methods, such as Aff affine transformation (affine transformation).
[0086] S1022: Through the super-resolution sub-network, generate the predicted image result corresponding to the feature map; through the feature recognition sub-network, generate the predicted text result corresponding to the feature map.
[0087] In some embodiments, please continue to refer to Figure 4 , the super-resolution sub-network may include a second feature extraction block, a plurality of cascaded sequence residual blocks, and a pixel recombination block. Generating the predicted image result corresponding to the feature map through the super-resolution sub-network includes:[[]]
[0088] Generate a binary map corresponding to the text image sample;
[0089] Perform channel fusion on the feature map and the binary map to generate a fused feature map;
[0090] Determine the second shallow feature map corresponding to the fused feature map through the second feature extraction block;
[0091] Process the second shallow feature map through the sequence residual block to obtain a deep feature map;
[0092] Perform pixel recombination on the deep feature map and the second shallow feature map through the pixel recombination block to obtain the corresponding predicted image result.
[0093] Among them, the second feature extraction block, the sequential residual block, and the pixel recombination block are cascaded in sequence. They can all include at least one convolutional group (conv). The convolutional group in the second feature extraction block can include only one convolutional kernel. The binary image is to set the grayscale value of the pixel points on the text image sample to 0 (black) or 255 (white). It can better analyze the shape and contour of the text. First, the average value K of the pixel points on the text image sample can be calculated, and then each pixel value is scanned. If the pixel value is greater than K, the pixel value is set to 255. If the pixel value is less than or equal to K, the pixel value is set to 0.
[0094] It should be noted that since the feature map reflects the spatial information and color information of the text in the RGB text image sample, while the binary image focuses on reflecting the shape and contour of the text, by performing channel fusion on the two, that is, connecting the channels of the two in series, the obtained fused feature map is more conducive to the generation of the super-resolution image of the subsequent text part.
[0095] The number of sequential residual blocks (Sequential Residual Block, SRB) can be determined based on requirements, such as being set to 5. The sequential residual block can also be set according to the structure of the residual network. For example, two adjacent sequential residual blocks are cascaded, and two non-adjacent sequential residual blocks are directly connected to extract deeper and sequence-related functions.
[0096] Please refer to Figure 5 , Figure 5 shows the structural schematic diagram of a single sequential residual block. Among them, each sequential residual block is modified from the residual block RB. Based on the residual block, a bidirectional LSTM (Long Short-Term Memory) structure (BLSTM) is introduced in the horizontal and vertical directions at the end. BLSTM can propagate the error differential, convert the fused feature map into a feature sequence, and feedback it back to the convolutional layer. Through BLSTM, semantic information feature extraction is performed on the fused feature map. Using the horizontal convolutional feature and the vertical convolutional feature as sequence inputs, its internal state is repeatedly updated in the hidden layer, so that the super-resolution network is also robust to oblique text.
[0097] The pixel recombination block can include an upsampling block and 1 convolutional kernel (such as a 1*1 convolution). It is used to convert the deep feature maps of multiple channels and the second shallow feature map into RGB image blocks of the expected size. The image composed of all image blocks is the prediction image result predicted by the super-resolution network.
[0098] In some embodiments, when paired text image samples and super-resolution image samples are obtained using different cameras, considering the shooting errors between the cameras, it will inevitably lead to pixel misalignment between the paired images. At this time, a central alignment block can be introduced. That is, the super-resolution sub-network can also include a central alignment block. Before the above step of "determining the second shallow feature map corresponding to the fusion feature map through the second feature extraction block", the text recognition method may further include:
[0099] Generating an aligned feature map corresponding to the fusion feature map through the central alignment block;
[0100] The above "determining the second shallow feature map corresponding to the fusion feature map through the second feature extraction block" includes: performing shallow feature extraction from the aligned feature map through the second feature extraction block to obtain the second shallow feature map.
[0101] Among them, the central alignment block can be a trained Spatial Transformer Networks (STN). By performing alignment correction on the fusion feature map before feature extraction, the problem of pixel misalignment between the text image sample and the super-resolution image sample can be solved.
[0102] In some embodiments, please continue to refer to Figure 4 , the feature recognition sub-network may include a feature compression block, an encoder, and an attention mechanism-based decoder connected in series. Generating a predicted text result corresponding to the feature map through the feature recognition sub-network includes:
[0103] Generating a one-dimensional feature vector corresponding to the feature map through the feature compression block;
[0104] Generating a feature sequence corresponding to the one-dimensional feature vector through the encoder;
[0105] Generating a predicted text result corresponding to the feature sequence through the decoder.
[0106] Among them, the feature compression block can obtain a one-dimensional vector from the feature map through 1×1 dimensionality reduction and recombination. For example, if the size of the feature map is 8×25×992, the size of the one-dimensional vector can be 25×1024. The encoder can be an encoder based on bidirectional LSTM. By introducing an attention mechanism-based decoder and a bidirectional LSTM-based encoder, the text recognition accuracy can be effectively improved.
[0107] S1023. Adjust the parameters of the text recognition network and the super-resolution network according to the predicted image result, the predicted text result, the text label, the super-resolution image sample, the first loss function, and the second loss function.
[0108] In some embodiments, step S1023 may specifically include:
[0109] Calculating a first error value according to a first loss function, a predicted text result, and a text label;
[0110] Calculating a second error value according to a second loss function, a predicted image result, and a super-resolution image sample;
[0111] Using the formula \(L = L\) rec +\(\lambda L\) sr to calculate a total error value, where \(L\) is the total error value, \(L\) rec is the first error value, \(L\) sr is the second error value, and \(\lambda\) is a hyperparameter;
[0112] Adjusting the network parameters of the text recognition network and the super-resolution network in reverse according to the total error value.
[0113] Among them, since in the recognition scenario of OCR, only the boundary between characters and the background in the text image needs to be noted, in the super-resolution network, a gradient loss can be introduced to sharpen the text edge, making the characters in the generated text image clearer. The first loss function may include a cross-entropy loss, and the second loss function may include a mean squared error loss and a gradient loss. The range of the hyperparameter can be 0 to 1, and the hyperparameter is used to adjust the weights of the two error values.
[0114] Specifically, the calculation formulas of \(L\) rec and \(L\) sr can be as follows:
[0115]
[0116] where \(MN\) is the size of the text image sample, \(M\) is the image length, and \(N\) is the image width. \(y\) i,j represents the label value (actual value) of the pixel point \((i, j)\) in the text image sample, and \(S\) i,j represents the model prediction value of the pixel point \((i, j)\) in the predicted text result. \(\nabla I_{hr}(x)\) represents the gradient of the pixel in the super-resolution image sample (HR image), \(\nabla I_{sr}(x)\) represents the gradient of the pixel in the predicted image result (SR image), the gradient refers to the spatial gradient of the RGB values of the image pixels, and "|| 1 " is the L1 norm function, that is, the sum of the absolute values of the vector elements, and \(E_x\) represents the image pixel expectation (which can be calculated by the mean squared error loss function).
[0117] S104. When the training is completed, use the trained text recognition network to perform text recognition on the text image to be recognized.
[0118] Among them, when the text image to be recognized is input into the trained text recognition network, it will be processed successively by the text correction sub-network, the feature extraction sub-network, and the feature recognition sub-network, and finally the text contained in the text image will be recognized, such as "serve".
[0119] It should be noted that, please continue to refer to Figure 2 , in this embodiment, when training the text recognition network (i.e., training stage ①), a super-resolution network branch is introduced for joint training, so that the expression of features in the feature space can be effectively improved. For example, after introducing the super-resolution network for joint training, the network parameters of the feature extraction sub-network and the text correction sub-network are improved. Then, the features in the feature maps obtained by these improved sub-networks have significantly improved feature resolution, which is conducive to improving the text recognition ability of the text recognition network for low-resolution images in inference stage ②. At the same time, after training is completed (i.e., inference stage ②), by removing the super-resolution network branch, it is ensured that when recognizing low-resolution text images, no additional computational amount will be added and the inference time will not change, ensuring high efficiency of text recognition.
[0120] As can be seen from the above, the OCR-based text recognition method provided by the embodiments of the present invention obtains a text image sample set, as well as the text label and super-resolution image sample corresponding to each text image sample in the text image sample set. Then, the pre-constructed text recognition network and super-resolution network are trained using the text image sample set, text labels, and super-resolution image samples. Among them, the text recognition network includes a first loss function, the super-resolution network includes a second loss function, and the text recognition network and the super-resolution network include at least one shared sub-network. At the same time, during the training process, according to the first loss function and the second loss function, the network parameters of the text recognition network and the super-resolution network are adjusted. When the training is completed, the trained text recognition network is used to recognize the text image to be recognized, so that only the super-resolution branch is added for feature learning in the training stage, and this branch is removed in the inference stage, effectively improving the recognition accuracy and recognition effect of the text recognition network for low-quality text images without changing the inference time.
[0121] According to the method described in the above embodiments, this embodiment will be further described from the perspective of an OCR-based text recognition device. The text recognition device can be specifically implemented as an independent entity and can be applied to electronic devices such as servers or mobile terminals.
[0122] Please refer to Figure 6 , Figure 6Specifically describes a text recognition device based on OCR provided by an embodiment of the present application. The text recognition device based on OCR may include: an acquisition unit 10, a training unit 20, and a recognition unit 30, where:
[0123] The acquisition unit 10 is configured to acquire a text image sample set, as well as the text label and the super-resolution image sample corresponding to each text image sample in the text image sample set;
[0124] The training unit 20 is configured to use the text image sample set, the text label, and the super-resolution image sample to train a pre-constructed text recognition network and a super-resolution network. The text recognition network includes a first loss function, the super-resolution network includes a second loss function, and the text recognition network and the super-resolution network include at least one shared sub-network; during the training process, adjust the network parameters of the text recognition network and the super-resolution network according to the first loss function and the second loss function;
[0125] The recognition unit 30 is configured to, when the training is completed, use the trained text recognition network to perform text recognition on the text image to be recognized.
[0126] In some embodiments, the text recognition network includes a connected feature extraction sub-network and a feature recognition sub-network, and the super-resolution network includes the connected feature extraction sub-network and a super-resolution sub-network. At this time, the training unit 20 is specifically configured to:
[0127] Determine the feature map corresponding to each text image sample through the feature extraction sub-network;
[0128] Generate a predicted image result corresponding to the feature map through the super-resolution sub-network;
[0129] Generate a predicted text result corresponding to the feature map through the feature recognition sub-network;
[0130] Adjust the parameters of the text recognition network and the super-resolution network according to the predicted image result, the predicted text result, the text label, the super-resolution image sample, the first loss function, and the second loss function.
[0131] In some embodiments, the training unit 20 is further configured to:
[0132] Calculate a first error value according to the first loss function, the predicted text result, and the text label;
[0133] Calculate a second error value according to the second loss function, the predicted image result, and the super-resolution image sample;
[0134] Use the formula L = L rec+λL sr Calculate the total error value, where L is the total error value, L rec is the first error value, L sr is the second error value, and λ is a hyperparameter;
[0135] According to the total error value, reversely adjust the network parameters of the text recognition network and the super-resolution network.
[0136] Specifically, the range of the hyperparameter can be 0 to 1, L rec and L sr The calculation formulas can be as follows:
[0137]
[0138] where MN is the size of the text image sample, M is the image length, and N is the image width. y i,j represents the label value (actual value) of the pixel point (i, j) in the text image sample, S i,j represents the model prediction value of the pixel point (i, j) in the predicted text result. ∇Ihr(x) represents the gradient of the pixel in the super-resolution image sample (HR image), ∇Isr(x) represents the gradient of the pixel in the predicted image result (SR image), and the gradient refers to the spatial gradient of the RGB values of the image pixels. "|||| 1 " is the first norm function, that is, the sum of the absolute values of the vector elements, and Ex represents the image pixel expectation (which can be calculated by the mean square error loss function).
[0139] In some embodiments, the feature extraction sub-network includes a first feature extraction block, a plurality of cascaded residual blocks, and a feature enhancement block. At this time, the training unit 20 is further configured to:
[0140] Determine the first shallow feature map corresponding to each text image sample through the first feature extraction block;
[0141] Process the first shallow feature map through the plurality of residual blocks;
[0142] Obtain the residual feature map output after the processing of each residual block through the feature enhancement block, respectively downsample the first shallow feature map and the residual feature map to obtain the corresponding downsampled feature maps, and then perform channel fusion on all the downsampled feature maps to obtain the feature map corresponding to the text image sample.
[0143] In some embodiments, both the text recognition network and the super-resolution network further include a text correction sub-network connected to the feature extraction sub-network. At this time, before determining the first shallow feature map corresponding to each text image sample through the first feature extraction block, the training unit 20 is further configured to:
[0144] Through this text correction sub-network, multiple key point information on each text image sample is determined, and according to a preset interpolation algorithm and this key point information, the text image sample is corrected to obtain a corresponding corrected image;
[0145] Correspondingly, the above step of "determining the first shallow feature map corresponding to each text image sample through this first feature extraction block" specifically includes: performing shallow feature extraction on each corrected image through this feature extraction sub-network to obtain a first shallow feature map.
[0146] In some embodiments, this super-resolution sub-network includes a second feature extraction block, a plurality of cascaded sequence residual blocks, and a pixel recombination block. At this time, this training unit 20 is further configured to:
[0147] Generate a binary map corresponding to the text image sample;
[0148] Perform channel fusion on this feature map and this binary map to generate a fused feature map;
[0149] Determine the second shallow feature map corresponding to this fused feature map through this second feature extraction block;
[0150] Process this second shallow feature map through this sequence residual block to obtain a deep feature map;
[0151] Perform pixel recombination on this deep feature map and this second shallow feature map through this pixel recombination block to obtain a corresponding predicted image result.
[0152] In some embodiments, this super-resolution sub-network further includes a center alignment block. At this time, before determining the second shallow feature map corresponding to this fused feature map through this second feature extraction block, this training unit 20 is further configured to:
[0153] Generate an aligned feature map corresponding to this fused feature map through this center alignment block;
[0154] Correspondingly, the above step of "determining the second shallow feature map corresponding to this fused feature map through this second feature extraction block" includes: performing shallow feature extraction from this aligned feature map through this second feature extraction block to obtain a second shallow feature map.
[0155] In some embodiments, this feature recognition sub-network may include a feature compression block, an encoder, and an attention mechanism-based decoder cascaded in sequence. At this time, this training unit 20 is further configured to:
[0156] Generate a one-dimensional feature vector corresponding to this feature map through this feature compression block;
[0157] Through this encoder, a feature sequence corresponding to the one-dimensional feature vector is generated;
[0158] Through this decoder, a predicted text result corresponding to the feature sequence is generated.
[0159] In specific implementation, each of the above modules / units can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of each of the above modules / units, reference can be made to the method embodiments described above, which will not be elaborated here.
[0160] In addition, an embodiment of the present application further provides an electronic device, which can be a device such as a smart phone, a tablet computer, a server, etc. As Figure 7 shown, the electronic device 200 includes a processor 201 and a memory 202. Among them, the processor 201 is electrically connected to the memory 202.
[0161] The processor 201 is the control center of the electronic device 200, connecting various parts of the entire electronic device through various interfaces and lines. By running or loading application programs stored in the memory 202, and calling data stored in the memory 202, it executes various functions of the electronic device and processes data, thereby monitoring the entire electronic device.
[0162] In this embodiment, the processor 201 in the electronic device 200 will load instructions corresponding to the processes of one or more application programs into the memory 202 according to the following steps, and the processor 201 will run the application programs stored in the memory 202 to implement various functions:
[0163] Obtain a text image sample set, as well as the text label and super-resolution image sample corresponding to each text image sample in the text image sample set; use the text image sample set, the text label, and the super-resolution image sample to train a pre-constructed text recognition network and super-resolution network. The text recognition network includes a first loss function, the super-resolution network includes a second loss function, and the text recognition network and the super-resolution network include at least one shared sub-network; during the training process, adjust the network parameters of the text recognition network and the super-resolution network according to the first loss function and the second loss function; when the training is completed, use the trained text recognition network to perform text recognition on the text image to be recognized.
[0164] Figure 8 The specific structural block diagram of the electronic device provided by the embodiment of the present invention is shown. This electronic device can be used to implement the OCR-based text recognition and generation method provided in the above embodiment. The electronic device 300 can include a smart phone or a server.
[0165] The electronic device may include components such as a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, a radio frequency (RF) circuit 303, a power supply 304, an input unit 305, and a display unit 306. Those skilled in the art can understand that the structure of the electronic device shown in the figure does not limit the electronic device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:
[0166] The processor 301 is the control center of the electronic device. Among them, the processor uses various interfaces and circuits to connect all parts of the entire electronic device, and by running or executing software programs and / or modules stored in the memory 302, and calling data stored in the memory 302, it executes various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor may include one or more processing cores; preferably, the processor may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor.
[0167] The memory 302 can be used to store software programs (computer programs) and modules. The processor 301 executes various functional applications and data processing by running the software programs and modules stored in the memory 302. The memory 302 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store data created according to the use of the electronic device. In addition, the memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 302 may also include a memory controller to provide the processor 301 with access to the memory 302.
[0168] The RF circuit 303 can be used for receiving and transmitting signals during the process of receiving and sending information. In particular, after receiving the downlink information of the base station, it is handed over to one or more processors 301 for processing. Additionally, the data related to the uplink is sent to the base station. Generally, the RF circuit 303 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a subscriber identity module (SIM) card, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 303 can also communicate with the network and other devices through wireless communication. This wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobilecommunication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0169] The electronic device further includes a power supply 304 (such as a battery) for powering each component. Preferably, the power supply 304 can be logically connected to the processor 301 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 304 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator, etc.
[0170] The electronic device may further include an input unit 305, which may be configured to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls. Specifically, in a specific embodiment, the input unit 305 may include a touch-sensitive surface and other input devices. The touch-sensitive surface, also known as a touch display screen or a touchpad, may collect touch operations of a user thereon or nearby (such as operations of the user using a finger, a stylus or any suitable object or accessory on or near the touch-sensitive surface), and drive corresponding connection devices according to a preset program. Optionally, the touch-sensitive surface may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects signals brought by the touch operation, and transmits the signals to the touch controller; the touch controller receives touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 301, and can receive commands sent by the processor 301 and execute them. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave may be used to implement the touch-sensitive surface. In addition to the touch-sensitive surface, the input unit 305 may further include other input devices. Specifically, the other input devices may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.
[0171] The electronic device may further include a display unit 306, which may be configured to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces may be composed of graphics, text, icons, videos, and any combination thereof. The display unit 306 includes a plurality of hardware display processing units, a video frame processing module, a display screen, etc. Among them, the plurality of hardware display processing units and the video frame processing module may be integrated in a processing chip. Among them, the display screen may include a display panel. Optionally, the display panel may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch-sensitive surface may cover the display panel. After the touch-sensitive surface detects a touch operation thereon or nearby, it is transmitted to the processor 301 to determine the type of touch event. Subsequently, the processor 301 provides corresponding visual output on the display panel according to the type of touch event. Although in the figure, the touch-sensitive surface and the display panel are implemented as two independent components to achieve input and input functions, in some embodiments, the touch-sensitive surface and the display panel may be integrated to achieve input and output functions.
[0172] Although not shown, the electronic device may further include a camera, a Bluetooth module, etc., which will not be elaborated herein. The electronic device further includes a first splicing module, which includes a signal processing module, a plurality of image processing modules connected to the signal processing module, and an image splicing module connected to the plurality of image processing modules. Each of the image processing modules is connected to a corresponding first display pixel interface. Specifically, in this embodiment, the processor 301 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 302 according to the following instructions, and the processor 301 will run the application programs stored in the memory 302 to implement various functions as follows:
[0173] Obtain a text image sample set, as well as the text labels and super-resolution image samples corresponding to each text image sample in the text image sample set; use the text image sample set, the text labels, and the super-resolution image samples to train a pre-constructed text recognition network and a super-resolution network. The text recognition network includes a first loss function, the super-resolution network includes a second loss function, and the text recognition network and the super-resolution network include at least one shared sub-network; during the training process, adjust the network parameters of the text recognition network and the super-resolution network according to the first loss function and the second loss function; when the training is completed, use the trained text recognition network to perform text recognition on the text image to be recognized.
[0174] The electronic device can implement the steps in any of the embodiments of the OCR-based text recognition method provided in this application embodiment. Therefore, it can achieve the beneficial effects that any text recognition method provided in this application embodiment can achieve. For details, please refer to the previous embodiments and will not be elaborated herein.
[0175] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions (computer programs), or the relevant hardware can be controlled by instructions (computer programs). The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. For this purpose, an embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored, and the computer program can be loaded by a processor to execute the steps in any of the embodiments of the OCR-based text recognition method provided in the embodiments of the present invention.
[0176] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), a magnetic disk, an optical disc, etc.
[0177] Since the computer program stored in the storage medium can execute the steps in any of the embodiments of the OCR-based text recognition method provided by the embodiments of the present invention, the beneficial effects achievable by any of the OCR-based text recognition methods provided by the embodiments of the present invention can be realized. For details, refer to the previous embodiments and will not be elaborated herein.
[0178] The above has introduced in detail an OCR-based text recognition method, device, electronic device, and storage medium provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An OCR-based text recognition method, characterized in that, it includes: obtaining a text image sample set, as well as the text label and super-resolution image sample corresponding to each text image sample in the text image sample set; using the text image sample set, the text label, and the super-resolution image sample to train a pre-constructed text recognition network and super-resolution network. The text recognition network includes a first loss function, a connected feature extraction sub-network and a feature recognition sub-network. The super-resolution network includes a second loss function, and the connected feature extraction sub-network and super-resolution sub-network, and the text recognition network and the super-resolution network include at least one shared sub-network; during the training process, through the feature extraction sub-network, determine the feature map corresponding to each text image sample; wherein, the feature extraction sub-network includes a first feature extraction block, a plurality of cascaded residual blocks, and a feature enhancement block. The process of determining the feature map corresponding to each text image sample through the feature extraction sub-network includes: determining the first shallow feature map corresponding to each text image sample through the first feature extraction block; processing the first shallow feature map through the plurality of residual blocks; through the feature enhancement block, obtaining the residual feature map output after each residual block processes, and respectively downsampling the first shallow feature map and the residual feature map to obtain the corresponding downsampled feature map, and then performing channel fusion on all the downsampled feature maps to obtain the feature map corresponding to the text image sample; generating a predicted image result corresponding to the feature map through the super-resolution sub-network; generating a predicted text result corresponding to the feature map through the feature recognition sub-network; adjusting the network parameters of the text recognition network and the super-resolution network according to the predicted image result, the predicted text result, the text label, the super-resolution image sample, the first loss function, and the second loss function; when the training is completed, use the trained text recognition network to perform text recognition on the text image to be recognized.
2. The text recognition method according to claim 1, characterized in that, the adjusting the network parameters of the text recognition network and the super-resolution network according to the predicted image result, the predicted text result, the text label, the super-resolution image sample, the first loss function, and the second loss function includes: calculating a first error value according to the first loss function, the predicted text result, and the text label; calculating a second error value according to the second loss function, the predicted image result, and the super-resolution image sample; Use the formula \(L = L\) rec +\(\lambda L\) sr to calculate the total error value, where \(L\) is the total error value, \(L\) rec is the first error value, \(L\) sr is the second error value, and \(\lambda\) is a hyperparameter; inversely adjusting the network parameters of the text recognition network and the super-resolution network according to the total error value.
3. The text recognition method according to claim 1, characterized in that, Both the text recognition network and the super-resolution network further include a text correction sub-network connected to the feature extraction sub-network. Before determining the first shallow feature map corresponding to each text image sample through the first feature extraction block, it further includes: Determining, through the text correction sub-network, multiple key point information on each text image sample, and correcting the text image sample according to a preset interpolation algorithm and the key point information to obtain a corresponding corrected image; The step of determining, through the first feature extraction block, the first shallow feature map corresponding to each text image sample specifically includes: performing shallow feature extraction on each corrected image through the feature extraction sub-network to obtain a first shallow feature map.
4. The text recognition method according to claim 1, wherein, The super-resolution sub-network includes a second feature extraction block, a plurality of cascaded sequence residual blocks, and a pixel recombination block. The step of generating a predicted image result corresponding to the feature map through the super-resolution sub-network includes: Generating a binary map corresponding to the text image sample; Performing channel fusion on the feature map and the binary map to generate a fused feature map; Determining, through the second feature extraction block, a second shallow feature map corresponding to the fused feature map; Processing the second shallow feature map through the sequence residual blocks to obtain a deep feature map; Performing pixel recombination on the deep feature map and the second shallow feature map through the pixel recombination block to obtain a corresponding predicted image result.
5. The text recognition method according to claim 4, wherein, The super-resolution sub-network further includes a center alignment block. Before determining, through the second feature extraction block, the second shallow feature map corresponding to the fused feature map, it further includes: Generating, through the center alignment block, an aligned feature map corresponding to the fused feature map; The step of determining, through the second feature extraction block, the second shallow feature map corresponding to the fused feature map includes: performing shallow feature extraction from the aligned feature map through the second feature extraction block to obtain a second shallow feature map.
6. An OCR-based text recognition device, wherein, it includes: An acquisition unit for acquiring a text image sample set, as well as the text label and the super-resolution image sample corresponding to each text image sample in the text image sample set; A training unit, configured to train a pre-constructed text recognition network and a super-resolution network by using the text image sample set, the text labels, and the super-resolution image samples. The text recognition network includes a first loss function, a feature extraction sub-network and a feature recognition sub-network connected in series. The super-resolution network includes a second loss function, the feature extraction sub-network and a super-resolution sub-network connected in series, and the text recognition network and the super-resolution network include at least one shared sub-network. During the training process, through the feature extraction sub-network, a feature map corresponding to each text image sample is determined. Wherein, the feature extraction sub-network includes a first feature extraction block, a plurality of cascaded residual blocks, and a feature enhancement block. Determining, through the feature extraction sub-network, a feature map corresponding to each text image sample includes: determining, through the first feature extraction block, a first shallow feature map corresponding to each text image sample; processing the first shallow feature map through the plurality of residual blocks; through the feature enhancement block, obtaining the residual feature maps output after the processing of each residual block, and respectively downsampling the first shallow feature map and the residual feature maps to obtain corresponding downsampled feature maps, and then performing channel fusion on all the downsampled feature maps to obtain the feature map corresponding to the text image sample; generating, through the super-resolution sub-network, a predicted image result corresponding to the feature map; generating, through the feature recognition sub-network, a predicted text result corresponding to the feature map; adjusting the parameters of the text recognition network and the super-resolution network according to the predicted image result, the predicted text result, the text labels, the super-resolution image samples, the first loss function, and the second loss function; An identification unit, configured to, when the training is completed, perform text recognition on a text image to be recognized by using the trained text recognition network.
7. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a plurality of instructions, and the instructions are adapted to be loaded by a processor to execute the OCR-based text recognition method according to any one of claims 1 to 5.
8. An electronic device, characterized in that, comprising a coupled memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program in the memory to execute the steps in the OCR-based text recognition method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Super-resolution-based vehicle detection method, device and equipment, and storage medium
CN112016507A
Super-resolution-based low-quality image recognition method and device, equipment and medium
CN113962862A