Text image super-resolution method and system based on lightweight convolutional neural network
Through the combination of lightweight convolutional neural network and LSTM network, the super resolution problem of low-quality text images is solved, efficient and clear text images are generated, and the OCR recognition rate is improved, which is suitable for mobile devices.
Patent Information
- Application Number
- CN202510823738.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The existing text super-resolution network model has large parameters and consumes a lot of computing resources, making it difficult to achieve efficient recognition on low-quality text images, and traditional natural scene super-resolution methods cannot effectively utilize the characteristics of text images.
Lightweight convolutional neural network is adopted, and lightweight convolutional neural network is trained, and memory information of multi-level feature maps is extracted using the LSTM network, and the error between feature maps is calculated as training losses, so as to realize super-resolution reconstruction of low-quality text images. Combined with image enhancement and multi-level feature fusion, a clear and recognizable text image is generated.
With low parameters and low computing resources, high-quality text images are generated, which improves the OCR recognition rate and is suitable for mobile devices, realizing efficient super-resolution reconstruction of text images.
Smart Images

Figure CN120339077A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image super-resolution, and particularly relates to a method and system for text image super-resolution based on a lightweight convolutional neural network. Background Art
[0002] OCR (Optical Character Recognition) refers to the process of using devices (such as scanners or digital cameras) to record text and text information in different scenarios, and then determining the position and shape of the text by detecting dark and bright patterns or edges, and then translating the shape into computer text using character recognition methods. It is a technology that converts the text in a scene into a black-and-white dot-matrix image file in an optical manner, and converts the text in the image into a text format through recognition software for further editing and processing by word processing software. How to improve the character recognition accuracy is the most important topic in OCR.
[0003] Nowadays, OCR text recognition technology has achieved excellent results on high-quality text images. However, when recognizing low-quality blurred text images, the recognition rate performance of OCR drops sharply. The reasons for low quality include low resolution, noise, blur, jpeg compression, artifacts, ringing phenomena, etc. The main difficulty in recognizing blurred text images lies in the lack of detailed information and edge texture information about characters in the image. Super-resolution is a reasonable method to solve this problem. However, traditional natural scene super-resolution methods reconstruct the global image without distinguishing between foreground and background, while text super-resolution needs to focus on text information, and at the same time cannot utilize the characteristic that the context of the text is relevant to help better information recovery and reconstruction. Therefore, natural scene super-resolution is not suitable for directly being used for blurred text tasks. Compared with natural images, scene text has arbitrary poses, illuminations, and blurs. Therefore, super-resolving super-blurred text images is more challenging. Therefore, a text content-aware text super-resolution network is needed to generate clear and recognizable text images for recognition.
[0004] The current text super-resolution network models have a large number of parameters and a large number of floating-point operations, require a large amount of computing resources and complex operators, and it is difficult to balance performance and resources. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention discloses a method and system for text image super-resolution based on a lightweight convolutional neural network, and the technical solutions adopted are as follows:
[0006] In a first aspect, the present invention proposes a method for text image super-resolution based on a lightweight convolutional neural network, including:
[0007] Train a lightweight convolutional neural network using low-quality text image-ground truth image pairs to achieve super-resolution reconstruction of low-quality text images;
[0008] The lightweight convolutional neural network includes a first network for extracting shallow features of low-quality text images, a second network for extracting first deep features and second deep features from the semantic information of shallow features, a first fusion network for fusing semantic information and its two deep features, a second fusion network for fusing shallow features and the output result of the first fusion network, and a third network for reconstructing high-resolution text images using the output result of the second fusion network;
[0009] The construction process of the loss function for training the lightweight convolutional neural network includes:
[0010] Extract multi-level feature maps of the reconstructed high-resolution text images and their ground truth images, use the LSTM network to extract the memory information of the feature map with the smallest spatial resolution in the multi-level feature maps, and generate the embedding vector of the memory information; calculate the errors at each level between the multi-level feature maps of the reconstructed high-resolution text images and their ground truth images, as well as the memory error between the embedding vectors of the memory information, and use the weighted result between the errors at each level and the memory error as the training loss.
[0011] Furthermore, when training the lightweight convolutional neural network, it also includes the step of image enhancement for low-quality text images.
[0012] Furthermore, the first network first expands the number of channels of the low-quality text image from 3 to 64, and then sequentially uses a convolutional layer and several stacked first residual blocks to extract the shallow features of the image.
[0013] Furthermore, the second network includes a convolutional layer, several stacked second residual blocks, and two parallel and identically structured deep extraction branches. First, use the convolutional layer and several stacked second residual blocks to extract the semantic information of the shallow features, and then use the two deep extraction branches to separately extract the first deep feature and the second deep feature from the semantic information.
[0014] Furthermore, the deep extraction branch is a multi-level structure. In each layer, the output feature of the previous layer is used as the input feature to perform a convolutional operation. After non-linearly activating the result of the convolutional operation, it is concatenated with the input feature in the channel dimension to obtain the output feature of the current layer.
[0015] Furthermore, the first fusion network first concatenates the semantic information, the first deep feature, and the second deep feature in the channel dimension, and then uses several stacked third residual blocks to extract the first fusion feature from the concatenated result.
[0016] Further, the second fusion network first concatenates the shallow features and the output result of the first fusion network in the channel dimension, performs a convolution operation on the concatenated result, then performs a non-linear activation on the result of the convolution operation and then executes a 1×1 convolution to obtain second fusion features containing global information.
[0017] Further, the third network first performs a sub-pixel convolution operation on the output result of the second fusion network to improve the resolution, and then performs a convolution operation to generate a reconstructed high-resolution text image.
[0018] Further, when training the lightweight convolutional neural network, a multi-level text recognition network is used to extract multi-level feature maps of the image. In each layer, the output features of the previous layer are used as input features to perform a convolution operation, and the results of the convolution operation are sequentially subjected to non-linear activation and pooling to obtain the output features of the current layer. The output features of different layers constitute a multi-level feature map, and the last-level feature map is the feature map with the smallest spatial resolution.
[0019] In a second aspect, the present invention proposes a text image super-resolution system based on a lightweight convolutional neural network for implementing the above-mentioned text image super-resolution method based on a lightweight convolutional neural network.
[0020] In a third aspect, the present invention proposes a computer electronic device, which is characterized by comprising a memory and a processor;
[0021] The memory is used to store a computer program;
[0022] The processor is used to implement the above-mentioned text image super-resolution method based on a lightweight convolutional neural network when executing the computer program.
[0023] The beneficial effects of the present invention are as follows:
[0024] The present invention proposes a text image super-resolution method based on a lightweight convolutional neural network. The lightweight convolutional neural network adopted has a parameter quantity of less than 1M and a floating-point operation of less than 4G. By inputting a low-quality text image into the lightweight convolutional neural network for super-resolution, the obtained high-resolution image and the ground-truth image are input into the text recognition network to calculate the feature difference in different feature dimensions and the embedding vector difference of the memory information of the feature map with the smallest spatial resolution as the text loss function, and the text loss function is used as a part of the loss to participate in the backpropagation of the gradient. The present invention uses a lightweight image super-resolution network in series with a text recognition network as a text prior to correct the quality of text image super-resolution, so that the super-resolved text image has rich detail information and excellent visual effects. The semantic information of the present invention makes the trained lightweight convolutional neural network very convenient to be integrated on mobile devices as an upstream task of the OCR task to improve the recognition rate of the OCR. Description of the Drawings
[0025] Figure 1 It is a flowchart of the image super-resolution method based on a lightweight convolutional neural network in the embodiments of the present invention;
[0026] Figure 2 It is a structural block diagram of the composition of each module in the lightweight convolutional neural network proposed by the present invention;
[0027] Figure 3 It is a schematic diagram of an electronic device proposed by the present invention. Detailed Embodiments
[0028] The present invention will be further described and explained below in conjunction with the detailed embodiments. The embodiments are only examples of the present disclosure content and do not delimit the scope of limitation. The technical features of each embodiment in the present invention can be combined correspondingly on the premise of no mutual conflict.
[0029] The drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0030] The flowcharts shown in the drawings are only illustrative descriptions and do not necessarily include all steps. For example, some steps can be decomposed, while some steps can be combined or partially combined. Therefore, the actual execution order may be changed according to the actual situation.
[0031] The present invention realizes text image super-resolution based on a lightweight convolutional neural network. Using a low-quality text image dataset as the training set and the corresponding ground-truth image of the low-quality text image as the label image, a lightweight convolutional neural network is trained, and the trained lightweight convolutional neural network is used to generate the super-resolution result of the low-quality text image. The lightweight convolutional neural network takes a low-quality text image as the input, and sequentially extracts shallow features, semantic information, deep features, fusion features from the text image by using a series-connected first network, second network, first fusion network, second fusion network, and third network, and reconstructs it into a high-resolution text image. The present invention uses a lightweight image super-resolution network in series with a text recognition network as a text prior to correct the quality of text image super-resolution, so that the super-resolved text image has rich detail information and excellent visual effects.
[0032] As Figure 1 shown, the main steps of the text image super-resolution method based on a lightweight convolutional neural network include:
[0033] S01, extract the shallow features of the low-quality text image after data enhancement using the first network.
[0034] S02, extract semantic information from the shallow features of the low-quality text image, and use a multi-layer dual-branch network to extract the first deep feature and the second deep feature from the semantic information respectively.
[0035] S03, fuse the semantic information, the first deep feature and the second deep feature using the first fusion network as the first fusion feature.
[0036] S04, fuse the shallow feature and the first fusion feature using the second fusion network as the second fusion feature.
[0037] S05, reconstruct the second fusion feature into a high-resolution text image using the third network.
[0038] S06, extract the multi-level feature maps of the reconstructed high-resolution text image and the ground truth image, use the LSTM network to extract the memory feature map of the feature map with the smallest spatial resolution in the multi-level feature maps, calculate the hierarchical error between the multi-level feature maps of the high-resolution text image and the ground truth image respectively, and the memory error between the memory feature maps, and use the weighted result between the memory error and the multiple hierarchical errors as the training loss to update the model parameters of the lightweight convolutional neural network.
[0039] S07, take the low-quality text image as the input, and use the trained lightweight convolutional neural network model to generate a high-resolution text image as the super-resolution result.
[0040] Exemplarily, the image enhancement in step S01 includes mirror symmetry, horizontal 90° flipping, vertical 90° flipping, etc. Image enhancement refers to the process of improving the visual effect of an image or highlighting specific information through a series of technical means, aiming to enhance the image quality, enhance key features or meet the requirements of subsequent processing. Its core goal is to solve the defects of the original image caused by factors such as shooting environment and equipment limitations (such as blurring, low contrast, noise interference, etc.), or to specifically enhance the detailed information in a specific area.
[0041] It should be noted that the original text image is a three-channel image, and the number of channels remains unchanged after image enhancement processing. Before extracting the shallow features of the low-quality text image after data augmentation, the number of channels of the low-quality text image after data augmentation is first increased from 3 channels to 64 channels, and then shallow feature extraction is performed. In a specific implementation of the present invention, the first network first increases the number of channels of the low-quality text image, and then sequentially uses a convolutional layer and several stacked first residual blocks to extract the shallow features of the image. Since these features are obtained through simple convolution, the information of the original image can be relatively completely retained. These shallow features serve as guiding information in the subsequent calculation process of reconstructing features, which can help the network converge quickly and achieve good performance at the same time.
[0042] The second network includes a convolutional layer, several stacked second residual blocks, and two parallel and identically structured deep extraction branches. First, the convolutional layer and several stacked second residual blocks are sequentially used to extract the semantic information of the shallow features, and then the two deep extraction branches are used to respectively extract the first deep feature and the second deep feature from the semantic information. Here, the deep extraction branch is a multi-level structure. In each layer, the output features of the previous layer are used as input features to perform a convolutional operation. After non-linearly activating the result of the convolutional operation, it is concatenated with the input features in the channel dimension to obtain the output features of the current layer. Since it is required that the number of network parameters and the number of floating-point operations are small while ensuring performance, the present invention adopts the structure of a residual concatenation block to extract and calculate features, and concatenates the calculated output features with the input features in the channel dimension as the input of the new structure, and repeats the above process to obtain the semantic information and deep features of the low-resolution text image.
[0043] The first fusion network first concatenates the semantic information, the first deep feature, and the second deep feature in the channel dimension, and then uses several stacked third residual blocks to extract the first fusion feature from the concatenated result. The second fusion network first concatenates the shallow features and the output result of the first fusion network in the channel dimension, performs a convolutional operation on the concatenated result, and then non-linearly activates the result of the convolutional operation and performs a 1×1 convolution to obtain the second fusion feature containing global information. The present invention sequentially executes two fusion networks, first performs multi-scale feature fusion from the semantic information scale and the deep feature scale after several residual concatenation blocks, and then performs feature channel concatenation and 1×1 convolution information fusion on the fusion features and the shallow features to obtain the global features containing the information of the entire image.
[0044] The third network first performs a sub-pixel convolution operation on the output result of the second fusion network to improve the resolution, and then performs a convolutional operation to generate a reconstructed high-resolution text image. The global features are transformed into a high-resolution image through sub-pixel convolution and ordinary convolution operations.
[0045] The present invention uses real images as labels and does not adopt the traditional VGG network for computing perception because the VGG network is not designed for text recognition tasks and cannot understand the specific information in text images. To calculate the text loss, the present invention uses a multi-level text recognition network to extract multi-level feature maps of the image. In each layer, the output feature of the previous layer is used as the input feature to perform a convolution operation. The result of the convolution operation is successively subjected to non-linear activation and pooling to obtain the output feature of the current layer. The output features of different layers constitute the multi-level feature map, and the feature map of the last level is the feature map with the smallest spatial resolution.
[0046] Exemplarily, the multi-level text recognition network is a pyramid-like structure. As the network deepens, the size of the obtained feature map continuously shrinks, the receptive field continuously expands, and the feature map contains more global information. After passing through a 4-layer network structure, the high-resolution image and the ground truth respectively obtain 4 feature maps of different sizes. The finally obtained feature maps are respectively input into the LSTM unit to obtain the final feature map, and the weighted difference of a total of 5 corresponding feature maps of the high-resolution image and the ground truth is calculated as the text loss function for gradient backpropagation.
[0047] Based on the same inventive concept, the present invention proposes a text image super-resolution system based on a lightweight convolutional neural network. As Figure 2 shown, the system mainly includes six modules: a text image data preprocessing module, a shallow feature calculation module, a semantic information and deep feature calculation module, a fusion feature calculation module, a super-resolution reconstruction module, and a training module based on text prior. Among them, the shallow feature calculation module, the semantic information and deep feature calculation module, the fusion feature calculation module, and the super-resolution reconstruction module constitute the lightweight convolutional neural network module. Each module is respectively used to execute the steps in the above-mentioned text image super-resolution method based on the lightweight convolutional neural network.
[0048] In a specific implementation of the present invention, the text image data preprocessing module is used to preprocess the input original low-resolution text image. Specifically, it executes the method in the following step S11.
[0049] S11, obtain the original low-resolution text image and perform image enhancement, and denote the low-resolution text image after image enhancement as .
[0050] Here, the image enhancement process can adopt the method in the above S01.
[0051] The shallow feature calculation module is used to extract the shallow features of the low-resolution text image. Specifically, it executes the method in the following step S12.
[0052] S12, the low-resolution text image obtained in step S11 The channel dimension of the text image is changed from 3 to 64 using convolution, and then shallow feature extraction is performed.
[0053] Exemplarily, the shallow feature extraction process is expressed as:
[0054]
[0055] Among them, represents the shallow features of the low-resolution text image ; represents the stacked first residual block network for extracting shallow features, represents the activation function, represents the convolutional layer. The spatial resolution remains unchanged during the shallow feature calculation process.
[0056] The semantic information and deep feature calculation module is used to extract semantic information, the first deep feature, and the second deep feature from the shallow features of the low-quality text image . Specifically, it executes the method in step S13 below.
[0057] S13. Further feature extraction is performed on the shallow features of the low-resolution text image obtained in S12 to obtain the semantic information of the low-resolution text image :
[0058]
[0059] Among them, represents the semantic information of the low-resolution text image , represents the stacked second residual block network for extracting semantic information, represents the activation function, represents the convolutional layer, represents the shallow features of the low-resolution text image .
[0060] The semantic information of the low-resolution text image is input into two parallel deep extraction branches for deep feature extraction, respectively obtaining the first deep feature and the second deep feature of the low-resolution text image :
[0061]
[0062]
[0063] Among them, and respectively represent the low-resolution text image The first deep feature and the second deep feature and represent two parallel deep extraction branches with the same structure. In this embodiment, and adopt a five-layer calculation structure, and its calculation process is shown as follows:
[0064]
[0065] Among them, represents the intermediate state of is the low-resolution text image semantic information of , is the low-resolution text image deep feature of or ; represents concatenating features in the channel dimension, represents the activation function, represents the convolutional layer. The deep extraction branch is a multi-layer structure, and the overall calculation idea of each layer is to input the features of the previous layer into the convolutional layer, non-linearly activate the result through the activation function, and finally concatenate it with itself in the channel to obtain the new feature .
[0066] The fusion feature calculation module is used to perform multi-scale fusion on the semantic information, the first deep feature and the second deep feature of the low-resolution text image to obtain the fusion feature, and assist in super-resolution of the low-resolution text image. Specifically, it executes the method in step S14 below.
[0067] S14, fuse the semantic information, the first deep feature and the second deep feature of the low-resolution text image obtained in step S13, specifically:
[0068]
[0069] Among them, represents the first fusion feature of the low-resolution text image , represents the stacked third residual block network for feature extraction, represents concatenating features in the channel dimension.
[0070] The super-resolution reconstruction module is used to combine the shallow feature with the fusion feature The fusion is performed to obtain a second fusion feature, which is a global feature and can be further reconstructed into a high-resolution text image. Specifically, the method in step S15 is described below.
[0071] S15, the low-resolution text image obtained in step S12 is Shallow features And the fusion feature obtained in step S14 Further fusion to obtain low-resolution text image The global features of . Exemplarily, the fusion process is expressed as:
[0072]
[0073] in, Represents a low-resolution text image The global characteristics of represents a 1×1 convolution for fusion features, represents the activation function, represents the convolutional layer, Indicates concatenating features in the channel dimension.
[0074] Convert low-resolution text images Global features Perform convolution reconstruction to obtain a low-resolution text image The third deep feature:
[0075]
[0076] in, Represents a low-resolution text image The third deep feature of is the reconstructed high-resolution text image, whose resolution is 2 times; represents the convolutional layer, Represents sub-pixel convolution.
[0077] A text prior-based training module for computing reconstructed high-resolution text images The difference between multiple pairs of feature maps of the true value image GT, and the weighted sum of the difference between the feature maps is used as the text loss To be used for training the above modules. Specifically, it executes the method in S16 below.
[0078] S16, reconstructing the high-resolution text image obtained in step S15 The true value image GT corresponding to the low-quality text image is respectively input into the multi-level text recognition network to extract the multi-level feature map. For example, the multi-level text recognition network is a four-layer pyramid convolution network, and the calculation process of the multi-level feature map is as follows:
[0079]
[0080]
[0081] Among them, represents the multi-level feature map of the reconstructed high-resolution text image When i = 1, is ; represents the multi-level feature map of the ground truth image GT. When i = 1, is GT. represents pooling, and the downsampling rate is 2; represents the activation function, represents the convolutional layer.
[0082] Using the feature map with the smallest spatial resolution of the reconstructed high-resolution text image and the feature map with the smallest spatial resolution of the ground truth image GT input into the same LSTM network to calculate the memory feature map and extract word embeddings:
[0083]
[0084]
[0085] Among them, , respectively represent the memory information embedding vectors of the reconstructed high-resolution text image and the ground truth image GT; represents the word embedding operation, represents the long short-term memory unit.
[0086] Using the multi-level feature map of the high-resolution text image obtained above, the multi-level feature map of the ground truth image GT , } and the memory feature map , calculate the loss:
[0087]
[0088]
[0089]
[0090] Among them, represents the high-resolution text image in the space of the text recognition network The text loss between the true value image GT The weighted weights assigned when calculating different feature differences Indicates the mean squared error loss, which is used to measure the similarity between two features. The calculated As part of the total loss function, it performs backpropagation of the gradient.
[0091] When using the trained lightweight convolutional neural network to implement the super-resolution reconstruction of low-quality text images, the final super-resolution result Is the high-resolution text image output by the image super-resolution method based on the lightweight convolutional neural network of the present invention for the low-resolution text image.
[0092] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The implementation methods of the remaining modules are not elaborated here. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative work.
[0093] The embodiments of the system of the present invention can be applied to any device with data processing capabilities. The any device with data processing capabilities can be a device or apparatus such as a computer. The system embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation.
[0094] In addition, it should be noted that the above-mentioned method for super-resolution of text images based on a lightweight convolutional neural network can essentially be executed by a computer program. Therefore, similarly, based on the same inventive concept, in another preferred embodiment of the present invention, a computer electronic device corresponding to the method provided in the above embodiment is also provided, such as Figure 3 As shown, it includes a memory and a processor;
[0095] The memory is used to store a computer program;
[0096] The processor is used to implement the method for super-resolution of text images based on a lightweight convolutional neural network in the above embodiment when executing the computer program.
[0097] In addition, when the logical instructions in the above-mentioned memory can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention.
[0098] The above-described embodiments are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical fields can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.
Claims
1. A text image super-resolution method based on a lightweight convolutional neural network, characterized in that Including: Training a lightweight convolutional neural network using low-quality text image-ground truth image pairs to achieve super-resolution reconstruction of low-quality text images; The lightweight convolutional neural network includes a first network for extracting shallow features of low-quality text images, a second network for extracting first deep features and second deep features from the semantic information of the shallow features, a first fusion network for fusing semantic information and its two deep features, a second fusion network for fusing shallow features and the output result of the first fusion network, and a third network for reconstructing high-resolution text images using the output result of the second fusion network; The construction process of the loss function for training the lightweight convolutional neural network includes: Extracting multi-level feature maps of the reconstructed high-resolution text images and their ground truth images, using an LSTM network to extract the memory information of the feature map with the smallest spatial resolution in the multi-level feature maps, and generating an embedded vector of the memory information; Calculating the errors at each level between the multi-level feature maps of the reconstructed high-resolution text images and their ground truth images respectively, and the memory error between the embedded vectors of the memory information, and taking the weighted result between the errors at each level and the memory error as the training loss.
2. The method for super-resolution of text images based on a lightweight convolutional neural network according to claim 1, wherein The first network first amplifies the number of channels of the low-quality text image, and then sequentially uses a convolutional layer and several stacked first residual blocks to extract the shallow features of the image.
3. The text image super-resolution method based on a lightweight convolutional neural network according to claim 1, characterized in that, The second network includes a convolutional layer, several stacked second residual blocks, and two parallel and identically structured deep extraction branches. First, the convolutional layer and several stacked second residual blocks are sequentially used to extract the semantic information of the shallow features, and then the first deep feature and the second deep feature are respectively extracted from the semantic information using the two deep extraction branches.
4. The method for super-resolution of text images based on a lightweight convolutional neural network according to claim 3, wherein The deep extraction branch is a multi-level structure. In each layer, the output feature of the previous layer is used as the input feature to perform a convolutional operation. After non-linearly activating the result of the convolutional operation, it is concatenated with the input feature in the channel dimension to obtain the output feature of the current layer.
5. The method for super-resolution of text images based on a lightweight convolutional neural network according to claim 1, characterized in that, The first fusion network first concatenates the semantic information, the first deep feature, and the second deep feature in the channel dimension, and then uses several stacked third residual blocks to extract the first fusion feature from the concatenated result.
6. The method for super-resolution of text images based on a lightweight convolutional neural network according to claim 1, characterized in that The second fusion network first concatenates the shallow features and the output result of the first fusion network in the channel dimension, performs a convolutional operation on the concatenated result, and then non-linearly activates the result of the convolutional operation and performs a 1×1 convolution to obtain the second fusion feature containing global information.
7. The method for super-resolution of text images based on a lightweight convolutional neural network according to claim 1, characterized in that The third network first performs a sub-pixel convolution operation on the output result of the second fusion network to improve the resolution, and then performs a convolutional operation to generate the reconstructed high-resolution text image.
8. The method for text image super-resolution based on a lightweight convolutional neural network according to claim 1, characterized in that When training the lightweight convolutional neural network, a multi-level text recognition network is used to extract the multi-level feature maps of the image. In each layer, the output feature of the previous layer is used as the input feature to perform a convolutional operation. The result of the convolutional operation is sequentially non-linearly activated and pooled to obtain the output feature of the current layer. The output features of different layers form the multi-level feature maps, and the feature map of the last level has the smallest spatial resolution.
9. A text image super-resolution system based on a lightweight convolutional neural network, characterized in that, Including: A lightweight convolutional neural network module, comprising a first network for extracting shallow features of a low-quality text image, a second network for extracting first deep features and second deep features from the semantic information of the shallow features, a first fusion network for fusing the semantic information and its two deep features, a second fusion network for fusing the shallow features and the output result of the first fusion network, and a third network for reconstructing a high-resolution text image by using the output result of the second fusion network; A training module based on text prior, which is used to train a lightweight convolutional neural network by using a low-quality text image-ground truth image pair to achieve super-resolution reconstruction of the low-quality text image; wherein, the construction process of the loss function for training the lightweight convolutional neural network includes: extracting multi-level feature maps of the reconstructed high-resolution text image and its ground truth image, using an LSTM network to extract the memory information of the feature map with the smallest spatial resolution in the multi-level feature maps to generate an embedded vector of the memory information; respectively calculating the errors at each level between the multi-level feature maps of the reconstructed high-resolution text image and its ground truth image, and the memory error between the embedded vectors of the memory information, and taking the weighted result between the errors at each level and the memory error as the training loss.
10. A computer electronic device, characterized in that, Comprising a memory and a processor; The memory is used for storing a computer program; The processor is used for, when executing the computer program, implementing the text image super-resolution method based on the lightweight convolutional neural network according to any one of claims 1 to 8.
Citation Information
Patent Citations
End-to-end quality enhancement method and device based on binocular stereo image
CN110399881A
Lightweight image super-division method and system based on attention feedback mechanism
CN113409191A
Light-weight multi-scale infrared image super-resolution reconstruction method
CN114092330A
LDCT image super-resolution enhancement method and device based on residual convolutional neural network
CN114255168A
Text recognition method and device based on image enhancement processing, equipment and medium
CN115187456A