A text image super-resolution reconstruction method based on text assistance
Through the text-assisted text image super-resolution reconstruction method, the problem of low-resolution text image recognition accuracy is solved, and higher text recognition accuracy and image reconstruction quality are achieved.
Patent Information
- Application Number
- CN202310244778.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-03-15
AI Technical Summary
When the prior art recognizes low-resolution text images, the recognition accuracy rate has dropped sharply, making it difficult to effectively improve the accuracy of text recognition.
Using text-assisted text image super-resolution reconstruction method, the pre-trained text image super-resolution reconstruction model is used to super-resolution reconstruction of low-resolution text images to improve image resolution, thereby improving the accuracy of text recognition.
By introducing attention mechanism and gated text detection module, the text image super-resolution reconstruction model can better focus on the text part, improving the reconstruction quality and accuracy of text recognition.
Smart Images

Figure CN116258632B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a text image super-resolution reconstruction method based on text assistance, belonging to the technical field of image processing. Background Art
[0002] Images are important carriers of information, and image processing technology is an important part of information processing technology. Among them, image resolution is an important factor in image processing. Image resolution affects the difficulty and visual effect of visual tasks. High-resolution images can provide clearer details and recognition, which facilitates people's analysis, decision-making, and viewing. However, obtaining high-resolution images often requires high-cost image acquisition equipment. A practical and effective way to solve this problem is image super-resolution reconstruction technology.
[0003] Image super-resolution is to reconstruct a low-resolution image into a high-resolution image through a certain algorithm. Image super-resolution reconstructs a high-resolution image only through computer processing, reducing the additional cost of high-precision image acquisition equipment. At the same time, super-resolution technology can make up for the image degradation caused by poor image information acquisition or compression during transmission as much as possible, so the research on image super-resolution algorithm is of great significance.
[0004] Traditional image super-resolution algorithms can be divided into two categories: interpolation-based methods and reconstruction-based methods. With the vigorous development of deep learning in recent years, image super-resolution algorithms have also begun to be closely integrated with deep learning. The main idea of the image super-resolution algorithm based on deep learning is to use high-resolution and low-resolution images to train the neural network model to characterize the mapping relationship between the two. Among them, the deep residual network extracts features from the image through multiple residual modules, and uses skip-layer links to connect the network input to the network output to ensure the stability of the entire network, making it easier for the entire model to converge during training.
[0005] Text recognition is the most basic and important task in computer vision, and it provides a foundation for subsequent text-related applications. Existing text recognizers have achieved satisfactory results on clear scene text images. However, when recognizing low-resolution text images, the recognition accuracy drops sharply. Summary of the invention
[0006] Objective: To overcome the deficiencies in the prior art, the present invention provides a text image super-resolution reconstruction method based on text assistance, which can use the super-resolution of text images as a preprocessing task for text recognition, thereby improving the accuracy of text recognition.
[0007] Technical solution: To solve the above technical problems, the technical solution adopted by the present invention is:
[0008] In a first aspect, the present invention provides a method for super-resolution reconstruction of text images based on text assistance, including:
[0009] Obtain a low-resolution text image to be reconstructed;
[0010] Input the low-resolution text image into a pre-trained text image super-resolution reconstruction model, and determine the super-resolution reconstruction result of the text image according to the output of the model;
[0011] Wherein the construction and training method of the text image super-resolution reconstruction model includes:
[0012] Obtain a text image data set;
[0013] Use the text image data set to train a pre-constructed text image super-resolution reconstruction model to obtain a trained text image super-resolution reconstruction model.
[0014] In some embodiments, inputting the low-resolution text image into a pre-trained text image super-resolution reconstruction model and determining the super-resolution reconstruction result of the text image according to the output of the model includes:
[0015] Take the RGB image corresponding to the low-resolution text image and its grayscale image as a four-channel image input, first extract shallow features through a convolutional layer with a convolution kernel of 3*3 and a Relu activation layer to obtain a first feature map;
[0016] Input the first feature map into a convolutional block attention module to obtain the channel and spatial attention weights of the image, and obtain a feature map with attention weights;
[0017] Extract text sequence features from the feature map with attention weights through multiple gated text detection modules;
[0018] Perform a skip connection between the text sequence features output by the gated text detection module and the first feature map, and add them to obtain a new feature map;
[0019] Input the new feature map into a sub-pixel convolutional upsampling layer and a Tanh activation layer to obtain the output four-channel text image of super-resolution reconstruction.
[0020] In some embodiments, the construction method of the text image super-resolution reconstruction model includes:
[0021] Replace the residual module in the deep residual network with a gated text detection module, and add a convolutional block attention module in front of the gated text detection module.
[0022] In some embodiments, the processing process of the convolutional block attention module is:
[0023] The channel and spatial attention weights of the image are sequentially inferred from the first input feature map along the channel and spatial dimensions, and then multiplied by the first input feature map to achieve adaptive adjustment of the features, obtaining a feature map with attention weights.
[0024] In some embodiments, the gated text detection module extracts image features by passing the feature map through two convolutional layers with a convolutional kernel of 3×3 and a BN layer in sequence, and then uses the LSTM module to extract horizontal text features and vertical text features from the horizontal and vertical directions respectively. The image features, horizontal text features, and vertical text features are fused through a gated feature fusion method and input into the next gated text detection module.
[0025] In some embodiments, the formula for gated feature fusion is:
[0026]
[0027] where n represents the number of types of features, W i represents the trainable weight, F i represents the input feature, and F g represents the output weighted feature.
[0028] In some embodiments, the loss function adopted during the training process of the text image super-resolution reconstruction model is:
[0029] L = L 2 + αL TA
[0030] where L 2 is the mean square error loss function, L TA is the text auxiliary loss function, and α is the hyperparameter of the ratio of the two loss functions;
[0031] The formula for the mean square error loss function is:
[0032]
[0033] where MSE represents the mean square error, x and y represent two images of size M×N, and x ij , y ij represent the values of pixel points;
[0034] The formula for the text auxiliary loss function is:
[0035]
[0036] where I HR represents the original high-resolution text image, and I SR represents the super-resolution text image generated by the network. Denote the one-dimensional vector generated by the pre-trained text recognition network encoder for the image, ||·|| 2 Denote taking the 2-norm.
[0037] In some embodiments, during the training process, the super-resolution image generated by the text image super-resolution reconstruction model and the corresponding original high-resolution image are respectively input into the encoder of the pre-trained text recognition network model to obtain the corresponding one-dimensional recognition sequences, and the similarity between the two text image recognition sequences is used to measure the similarity of the text content of the two images.
[0038] In a second aspect, the present invention provides a text image super-resolution reconstruction device based on text assistance, including a processor and a storage medium;
[0039] The storage medium is used to store instructions;
[0040] The processor is used to operate according to the instructions to execute the steps of the method according to the first aspect.
[0041] In a third aspect, the present invention provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method according to the first aspect are implemented.
[0042] In a fourth aspect, the present invention provides a device, including,
[0043] One or more processors, one or more memories, and one or more programs, where one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the methods described in the first aspect.
[0044] Advantageous effects: The text image super-resolution reconstruction method based on text assistance provided by the present invention has the following advantages:
[0045] 1. The text image super-resolution model disclosed in the present invention introduces an attention mechanism to assign different weights to the feature maps. Compared with other ordinary text image super-resolution models, it enables super-resolution and subsequent text detection to better focus on the text part, thereby improving the reconstruction quality of the model.
[0046] 2. The text image super-resolution model disclosed in the present invention fuses text sequence features and image texture features. Compared with other ordinary super-resolution models, it fully excavates and utilizes the text information in the low-resolution image, which helps to improve the super-resolution reconstruction quality of the text image.
[0047] 3. In a further technical solution of the present invention, a text-assisted loss is proposed. Compared with the original loss function, the text-assisted loss reflects both the text content and the resolution of the text image, which helps to generate high-resolution text images with stronger readability. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 Schematic flowchart of a text image super-resolution reconstruction method based on text assistance according to an embodiment of the present invention;
[0049] Figure 2 Schematic diagram of the network framework of a text image super-resolution reconstruction model according to an embodiment of the present invention;
[0050] Figure 3 Schematic diagram of a gated text detection module according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The present invention will be further described below with reference to the drawings and embodiments. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and cannot be used to limit the protection scope of the present invention.
[0052] In the description of the present invention, the meaning of several is more than one, the meaning of multiple is more than two, greater than, less than, exceeding, etc. are understood as not including the present number, and above, below, within, etc. are understood as including the present number. If there is a description of first and second, it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.
[0053] In the description of the present invention, the description with reference to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0054] At present, image super-resolution based on deep learning has achieved relatively remarkable performance. However, the purpose of most of these methods is to restore the detailed texture of natural images. The task of text image super-resolution reconstruction pays more attention to improving the readability of text while enhancing the image resolution. Treating text images as natural images to perform super-resolution will ignore the classification information brought by the text itself in the image. Therefore, compared with general image super-resolution reconstruction algorithms, text image super-resolution reconstruction requires targeted algorithms to obtain satisfactory super-resolution reconstruction results.
[0055] The present invention discloses a text-assisted text image super-resolution reconstruction method. In this method, the residual module in the deep residual network (ResNet) is replaced with the gated text detection module proposed by the present invention. Image feature extraction is performed using convolutional layers, and then sequence feature extraction is carried out through a bidirectional gated recurrent unit (GRU). The original skip connection in the residual module is changed to a gated feature fusion mechanism, and multiple gated text detection modules are connected to extract deep features. On the other hand, a convolutional block attention mechanism module (CBAM) is added to the network. By inputting features with channel and spatial attention weights into the gated text detection module, the extraction of text sequence features is further enhanced. In addition, the method proposes a new loss function called text-assisted loss, which takes into account two objectives: image resolution and text readability, effectively improving the performance of the text image super-resolution model.
[0056] Embodiment 1
[0057] A text-assisted text image super-resolution reconstruction method includes:
[0058] Obtain a low-resolution text image to be reconstructed;
[0059] Input the low-resolution text image into a pre-trained text image super-resolution reconstruction model, and determine the text image super-resolution reconstruction result according to the output of the model;
[0060] Wherein the construction and training method of the text image super-resolution reconstruction model includes:
[0061] Obtain a text image dataset;
[0062] Use the text image dataset to train a pre-constructed text image super-resolution reconstruction model to obtain a trained text image super-resolution reconstruction model.
[0063] In some embodiments, a text-assisted text image super-resolution reconstruction method, as Figure 1 shown, includes the following steps:
[0064] S1: Obtain a text image dataset. In this embodiment, the text image dataset is the Textzoom dataset.
[0065] S2: Divide the dataset into a test set and a validation set, and input them into the text image super-resolution reconstruction model. The schematic diagram of the model network framework is as Figure 2 shown.
[0066] Take the RGB image corresponding to the low-resolution image and its grayscale image as a four-channel image input. First, extract shallow features through a convolutional layer with a convolution kernel of 3*3 and a Relu activation layer; input the feature map into the convolutional block attention module to obtain the channel and spatial attention weights of the image, and further extract the text sequence features through multiple gated text detection modules; perform a skip connection (Concatenation) on the features output by the gated text detection module and the features output by the first-layer convolution, and add them to obtain a new feature map; input the new feature map into the sub-pixel convolutional upsampling layer (DeConv) and the Tanh activation layer, and finally output the four-channel text image of the super-resolution reconstruction.
[0067] The convolutional block attention module infers the attention weights along the channel and spatial dimensions of the input feature map in turn, and then multiplies them with the input feature map to achieve adaptive adjustment of the features.
[0068] As Figure 3 shown, the gated text detection module proposed in this embodiment further extracts the image features by passing the feature map through two convolutional layers with a convolution kernel of 3*3 and a BN layer in turn. Then, the LSTM module is used to extract the text features horizontally and vertically respectively, and the three features are fused through a gated feature fusion method and input into the next gated text detection module.
[0069] The formula for gated feature fusion is:
[0070]
[0071] where n represents the number of types of features, W i represents the trainable weight, F i represents the input feature, and F g represents the output weighted feature.
[0072] S3: Construct the loss function of the text image super-resolution network. The overall loss function of the network is:
[0073] = L 2 + αL TA
[0074] where, L 2 is the mean square error loss function, and L TAis the text - assisted loss function, and α is the hyper - parameter of the ratio of the two loss functions.
[0075] The mean - square error function is a commonly used loss function in the field of image processing. The specific formula is:
[0076]
[0077] where x and y represent two images of size M×N, and x ij , y ij represents the value of the pixel point. The similarity at the pixel level of the two images is measured by the mean - square error.
[0078] The specific formula of the text - assisted loss function proposed in the present invention is:
[0079]
[0080] where, HR represents the original high - resolution text image, I SR represents the super - resolution text image generated by the network, represents the one - dimensional vector generated by the image through the encoder of the pre - trained text recognition network.
[0081] In this embodiment, the text recognition network selects the pre - trained RCNN text recognition network. The super - resolution image generated by the network and the original high - resolution image are respectively input into the encoder of the RCNN model to obtain the corresponding one - dimensional recognition sequences. The similarity of the text content of the two images is measured by the similarity of the recognition sequences of the two text images.
[0082] S4: Train the super - resolution network for text images, specifically including:
[0083] S4.1: Convert the low - resolution image into an RGB image and its grayscale image as a four - channel image input;
[0084] S4.2: Input the four - channel image obtained in S4.1 into the convolutional layer to extract shallow - layer features and obtain the feature map;
[0085] S4.3: Input the feature map obtained in S4.2 into the convolutional block attention module to obtain the channel and spatial attention weights of the image;
[0086] S4.4: Further extract the text sequence features from the feature map with attention weights obtained in S4.3 through multiple gated text detection modules;
[0087] S4.5: Perform skip - connection on the feature output in S4.4 and the feature output in S4.2, add them to obtain a new feature map;
[0088] S4.6: Upsample the feature map obtained in S4.5 through sub-pixel convolution and output the super-resolution result through one convolutional layer;
[0089] S4.7: Iterate the process from S4.1 to S4.6, and use the loss function to supervise the network training.
[0090] For the input low-resolution text image, first use forward propagation to calculate the total error, then use backpropagation to calculate the partial derivatives of each weight parameter, and finally update the weight parameters according to the gradient descent method. Iterate this process, save the model weight parameters when the loss function is minimized, and obtain the trained super-resolution network model.
[0091] S5: Input the low-resolution text image to be processed into the text image super-resolution reconstruction model to obtain a high-resolution text image.
[0092] In this embodiment, the model training is completed using an NVIDIA RTX 3090 GPU based on the software environment of Python 3.6.9, torch 1.10.1, and torchvision 0.11.2 on a 64-bit Ubuntu 18.04.5 operating system. The Adam optimizer is used during the training process and the learning rate is set to 10 -4 , the number of training iterations is 500 times, and the total training duration is about 40 hours.
[0093] Embodiment 2
[0094] In a second aspect, this embodiment provides a text image super-resolution reconstruction device based on text assistance, including a processor and a storage medium;
[0095] The storage medium is used to store instructions;
[0096] The processor is used to operate according to the instructions to execute the steps of the method according to Embodiment 1.
[0097] Embodiment 3
[0098] In a third aspect, this embodiment provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method according to Embodiment 1 are implemented.
[0099] Embodiment 4
[0100] In a fourth aspect, the present invention provides a device, including,
[0101] One or more processors, one or more memories, and one or more programs, where the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any of the methods implemented in the method described in Embodiment 1.
[0102] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.
[0103] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0104] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0106] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A text image super-resolution reconstruction method based on text assistance, characterized in that, it includes: Obtain a low-resolution text image to be reconstructed; Input the low-resolution text image into a pre-trained text image super-resolution reconstruction model, and determine the text image super-resolution reconstruction result according to the output of the model. Specifically, it includes: using the RGB image corresponding to the low-resolution text image and its grayscale image as a four-channel image input, first extracting shallow features through a convolutional layer with a convolution kernel of 3*3 and a Relu activation layer to obtain a first feature map; inputting the first feature map into a convolutional block attention module to obtain the channel and spatial attention weights of the image, and obtaining a feature map with attention weights; extracting text sequence features from the feature map with attention weights through multiple gated text detection modules; performing a skip connection between the text sequence features output by the gated text detection module and the first feature map, adding them to obtain a new feature map; inputting the new feature map into a sub-pixel convolutional upsampling layer and a Tanh activation layer to obtain the output four-channel text image of the super-resolution reconstruction; wherein the construction and training method of the text image super-resolution reconstruction model includes: obtaining a text image dataset; using the text image dataset to train a pre-constructed text image super-resolution reconstruction model to obtain a trained text image super-resolution reconstruction model.
2. The text image super-resolution reconstruction method based on text assistance according to claim 1, characterized in that, the construction method of the text image super-resolution reconstruction model includes: Replacing the residual module in the deep residual network with a gated text detection module, and adding a convolutional block attention module before the gated text detection module; The processing process of the convolutional block attention module is: Inferring the channel and spatial attention weights of the image along the channel and spatial dimensions of the input first feature map in sequence, and then multiplying them with the input first feature map to achieve adaptive adjustment of the features, and obtaining a feature map with attention weights.
3. The text image super-resolution reconstruction method based on text assistance according to claim 1, characterized in that, The gated text detection module extracts image features by passing the feature map through two convolutional layers with a convolution kernel of 3*3 and a BN layer in sequence, and then uses an LSTM module to extract horizontal text features and vertical text features from the horizontal and vertical directions respectively, and fuses the image features, horizontal text features and vertical text features through a gated feature fusion method, and inputs them into the next gated text detection module.
4. The text image super-resolution reconstruction method based on text assistance according to claim 3, characterized in that, The formula for gated feature fusion is: where n represents the number of types of features, and W i represents the trainable weights, and F i represents the input features, and F g represents the weighted output features.
5. The text image super-resolution reconstruction method based on text assistance according to claim 1, characterized in that, The loss function used in the training process of the text image super-resolution reconstruction model is: L = L 2 + αL TA Among them, L 2 is the mean squared error loss function, and L TA is the text auxiliary loss function, and α is the hyperparameter of the ratio of the two loss functions; The formula for the mean square error loss function is: where MSE represents the mean square error, and x and y represent two images of size M×N, and x ij , y ij represents the value of the pixel point; The formula for the text assistance loss function is: Among them, I HR represents the original high-resolution text image, I SR represents the super-resolution text image generated by the network, represents the one-dimensional vector generated by the image through the pre-trained text recognition network encoder, ||·|| 2 represents taking the 2-norm.
6. The text image super-resolution reconstruction method based on text assistance according to claim 5, characterized in that, During the training process, the super-resolution images generated by the text image super-resolution reconstruction model and the corresponding original high-resolution images are respectively input into the encoder of the pre-trained text recognition network model to obtain the corresponding one-dimensional recognition sequences, and the similarity of the text contents of the two images is measured by the similarity of the recognition sequences of the two text images.
7. A text image super-resolution reconstruction device based on text assistance, characterized in that, it includes a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 6.
8. A storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer device, characterized in that: it includes, one or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any one of the methods according to claims 1 to 6.
Citation Information
Patent Citations
Super-resolution reconstruction method for Chinese text image
CN114626984A
Text recognition method and device based on OCR (Optical Character Recognition), storage medium and electronic equipment
CN115188000A
Text perception loss-based attention text super-resolution method
CN115713464A