Low-Resolution Text Image Recognition Method Based on SROCRN Network

By integrating the improved super-resolution module and image recognition module in the OCR recognition network, the super-resolution image recognition network (SROCRN) is built, which solves the problem of low-resolution text images in OCR recognition, and achieves higher recognition accuracy.

CN112733716BActive Publication Date: 2025-06-17HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110030021.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-11
Publication Date
2025-06-17
Estimated Expiration
2041-01-11

AI Technical Summary

Technical Problem

The low resolution text images are problem that the recognition accuracy is low due to their low resolution during the OCR recognition process.

Method used

By fusing and improving the existing image super-resolution reconstruction network (SRGAN) with the text image OCR recognition network (CRNN), a super-resolution image recognition network (SROCRN) is proposed to improve the resolution of low-resolution text images and improve the accuracy of OCR recognition.

Benefits of technology

Through the improved super-resolution module and image recognition module, the SROCRN network can reconstruct high-resolution text images more accurately, thereby significantly improving the accuracy of OCR recognition, solving the problem of low-resolution text image recognition rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112733716B_ABST
    Figure CN112733716B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for recognizing low-resolution text images based on the SROCRN network. Aiming at the problem of low accuracy in OCR recognition of low-resolution text images, the present invention method fuses and improves the existing image super-resolution reconstruction network (SRGAN) and the text image OCR recognition network (CRNN), and further proposes a super-resolution image recognition network (SROCRN), thereby solving the problem of OCR recognition of low-resolution text images. By combining the improved super-resolution reconstruction technology and image recognition technology, the method for recognizing low-resolution text images based on the SROCRN network is used to recognize low-resolution text images, solving the problem of difficulty in recognizing and obtaining text sequences due to insufficient resolution during the recognition of some text images. This method is easy to implement and has good recognition effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of text image recognition, and in particular to a method for recognizing low-resolution text images based on the SROCRN network. Background Art

[0002] In today's society, text image recognition (OCR) plays an increasingly important role in various fields. However, there is currently no suitable solution to the problem of low recognition rate for low-resolution text images. Due to the influence of different compression coding methods and image degradation functions during the transmission process, the resolution of text images will decrease accordingly, thereby affecting the accuracy and integrity of text recognition. For some text images containing important information, it is very regrettable that they cannot be accurately recognized due to resolution and clarity limitations. Therefore, when using low-resolution text images for OCR recognition, it is extremely necessary to improve the resolution of low-resolution text images through technical means before performing OCR recognition. Summary of the Invention

[0003] The purpose of the present invention is to address the deficiencies of the prior art and provide a method for recognizing low-resolution text images based on the SROCRN network, aiming to solve the problem of low recognition accuracy caused by the low resolution of low-resolution text images during the OCR recognition process. To solve this problem, the method proposed by the present invention integrates and improves the existing image super-resolution reconstruction network (SRGAN) and the text image OCR recognition network (CRNN), and further proposes the super-resolution image recognition network (SROCRN), thereby solving the problem of OCR recognition of low-resolution text images.

[0004] The technical solution of the method of the present invention is divided into two processes: the construction and training of the SROCRN network model and the recognition of low-resolution text images. The specific content is as follows:

[0005] Step 1: Construct the dataset of the SROCRN network model

[0006] Obtain a number of original high-resolution text images with a resolution of W×H, and label them (the actual sequence content of the text images). Divide these high-resolution text images into group A and group B according to a ratio of 3:1. The images in group A and group B are respectively subjected to two image scaling transformations to obtain group A-1 and group B-1 of low-resolution text images with a size of 1 / 4*W×1*4*H. The images in group A and group A-1 form the training set, and the images in group B-1 form the test set. The training set and the test set together constitute the dataset of the SROCRN network model.

[0007] Step 2: Construct the SROCRN network:

[0008] 2-1 Constructing the super-resolution module of the SROCRN network:

[0009] The super-resolution module adopts an adversarial network and consists of a generator and a discriminator;

[0010] The generator is composed of a convolutional layer, an upsampling layer, five cascaded residual modules, and two cascaded upsampling layers in sequence. The input of the convolutional layer is A-1 groups of low-resolution text images;

[0011] The discriminator is composed of a convolutional layer, an activation layer, five cascaded residual modules, a feature transformation layer, and a fully connected layer in sequence. The input of the convolutional layer is the output of the activation layer and the original high-resolution text image.

[0012] The residual module includes a convolutional layer, a normalization layer, and an activation layer.

[0013] 2-2 Constructing the image recognition module of the SROCRN network:

[0014] The image recognition module adopts a combination of a convolutional network (CNN) and a short-term memory network (RNN), and is composed of a combination of a text detection (CTPN) module and a CRNN module; The CTPN module is composed of a VGG feature extraction layer, a convolutional layer, a BLSTM temporal information fusion layer, and a fully connected layer. The input of the VGG feature extraction layer is the output of the super-resolution module; The CRNN module is composed of a convolutional layer, a pooling layer, an RNN sequence feature extraction layer, and a fully connected layer. The input of the convolutional layer is the output of the CTPN module.

[0015] Step 3. Training the SROCRN model using the dataset

[0016] Step 4. Recognition of low-resolution text images (testing the network using the dataset)

[0017] 4-1 Encapsulate the test set in the dataset using the DATALOADER function and import it into the PYTHON environment;

[0018] 4-2 Load the corresponding trained SROCRN model, use the above test set as the input image and input it into the model to obtain the finally recognized text sequence.

[0019] Another object of the present invention is to provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the above method.

[0020] Another object of the present invention is to provide a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the above method is implemented.

[0021] Advantages of the present invention:

[0022] 1) In the present invention, the super-resolution module of the SROCRN network changes the normalization layer in the residual module from the batch normalization layer to the instance normalization layer on the basis of SRGAN (see formulas (8)-(10)), making the data distribution after normalization more focused on a single image itself from the original one Batch, thereby obtaining a more accurate feature map, improving the image reconstruction effect of the super-resolution module, and further improving the accuracy of OCR recognition.

[0023] 2) In the present invention, the super-resolution module of the SROCRN network changes the original 16-layer residual structure to a 5-layer residual structure on the basis of SRGAN. Based on the relatively single content of the text image, after repeated comparison, it is finally determined to use a 5-layer residual structure for training, reducing the training difficulty on the premise of ensuring the reconstruction effect, greatly reducing the model size of the super-resolution module, and ensuring the high efficiency of the training of the entire low-resolution text image recognition network.

[0024] 3) In the present invention, the image recognition module of the SROCRN network integrates the CTPN module on the basis of the original CRNN module for pre-text detection, detecting and bounding the text area, thereby improving the accuracy and precision of the CRNN to obtain text features, and further improving the accuracy of OCR recognition of low-resolution text images.

[0025] In summary, aiming at the problem of low accuracy in OCR recognition of low-resolution text images, the method of the present invention combines the improved super-resolution reconstruction technology and image recognition technology, and uses the low-resolution text image recognition method based on the SROCRN network to recognize low-resolution text images, solving the problem of difficulty in recognizing and obtaining text sequences caused by insufficient resolution in the recognition process of some text images. This method is easy to implement and has a good recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is the flowchart of the method of the present invention;

[0027] Figure 2 is the overall model construction diagram of the SROCRN network of the present invention;

[0028] Figure 3 is the super-resolution module construction diagram of the SROCRN network of the present invention;

[0029] Figure 4 is the residual module construction diagram of the SROCRN network of the present invention;

[0030] Figure 5 is the image recognition module construction diagram of the SROCRN network of the present invention;

[0031] Figure 6 This is an example diagram of the implementation effect of the embodiments of the present invention; among them, (a) is a low-resolution text image, (b) is the recognition effect of the CTPN network, (c) is the recognition effect of the CRNN network, and (d) is the recognition effect of the SROCRN network of the present invention. Specific implementation manners

[0032] The following further analyzes and describes the present invention in conjunction with the accompanying drawings and specific embodiments.

[0033] It is carried out through two processes: the construction and training of the low-resolution text image recognition model based on the SROCRN network and the recognition of low-resolution text images, as Figure 1 The specific content is as follows:

[0034] Step 1: Construction of the dataset for model training:

[0035] 1.1 Collect n (n>800) high-resolution text images with a high-resolution size of W×H. The images are all on a solid-color background. Label these images (the actual sequence content of the text images) and write them into a json document, and divide them into group A and group B according to 3:1;

[0036] 1.2 Perform two image scaling transformations on the images in group A and group B collected in 1.1 respectively to obtain low-resolution images with a resolution of 1 / 4*W×1 / 4H, and mark them as group A-1 and group B-1; among them, the image scaling transformation uses the bicubic interpolation method, and the bicubic interpolation basis function is used as the basis function. The bicubic interpolation calculation is carried out according to formula (1):

[0037] f(i+u,j+v) = ABC (1)

[0038] Among them, A, B, and C are all matrices, and the forms are as follows:

[0039] A = [s(v + 1)s(v)s(1 - v)s(2 - v)] (2)

[0040]

[0041] (i,j): Pixel coordinates of the original image, where i is the abscissa value and j is the ordinate value, and both i and j are non-negative integers;

[0042] f(i,j): Pixel gray value of the original image;

[0043] (i + u,j + v): Pixel coordinates of the new image after the scaling transformation;

[0044] f(i + u,j + v): Pixel gray value of the new image after the scaling transformation;

[0045] u: The distance between the abscissa i of the original image pixel coordinate and the abscissa of the new image pixel coordinate (i + u, j + v) along the abscissa direction;

[0046] v: The distance between the ordinate j of the original image pixel coordinate and the pixel coordinate of the new image (i + u, j + v) along the ordinate direction;

[0047] |x|: The distance of the image pixel from the origin along the x direction;

[0048] s(x): The approximation polynomial of sin(π·x) / x, which is the interpolation kernel;

[0049] 1.3 Mark the group A and group A-1 in 1.2 as the training set, and mark the group B-1 as the test set.

[0050] Step 2. As Figure 2 Construct the SROCRN network:

[0051] 2.1 As Figure 3 Construct the super-resolution module of the SROCRN network:

[0052] 2.1.1 Construct the first two layers of the generator, perform convolution and activation operations on the images in the training set to extract the feature map, and this operation is carried out according to formulas (6) and (7):

[0053] Y = F2(X) = MAX(0, w * F1(X) + b) (6)

[0054] In formula (6),

[0055] X: The training set image in 1.3;

[0056] F1: The RGB channel separation function of the image;

[0057] F2: The processing function of the convolution operation;

[0058] w: The convolution kernel of size f1×f1×n, where f1 is the spatial size of the convolution kernel and n is the number of convolution kernels;

[0059] b: The n-dimensional vector;

[0060]

[0061] In formula (7),

[0062] P: The feature map after activation processing;

[0063] Y i,j : The pixel value of the image at the point (i, j) after the convolution operation;

[0064] a i,j: Fixed parameter within the interval (1, +∞);

[0065] 2.1.2 Construct the five - layer residual module of the generator, as Figure 4 The residual module consists of a convolutional layer, a normalization layer, and an activation layer, further extracting the number of channels of the extended feature map. Among them, the convolutional layer and the activation layer are carried out according to formulas (6) and (7), and their inputs are respectively the output of the activation layer or the previous residual module;

[0066] The normalization layer is carried out according to formulas (8), (9), (10):

[0067]

[0068] In formulas (8), (9), (10),

[0069] F3: Normalization processing function;

[0070] P: Feature map of the image after the previous - layer convolution;

[0071] u: Mean value of the feature map of the image after the previous - layer convolution;

[0072] σ 2 : Variance of the feature map of the image after the previous - layer convolution;

[0073] ε: Variable parameter;

[0074] H: Width of the feature map;

[0075] W: Length of the feature map;

[0076] 2.1.2 Construct the last two up - sampling layers of the generator, magnify the feature map output by the last residual module by 4 times to generate a high - resolution image. The up - sampling layer mainly uses the method of sub - pixel periodic screening, and this method is carried out according to formulas (11), (12), (13), (14):

[0077] P1(m,n,p) = P(i,j,k)(11)

[0078]

[0079] p = k×r + i%r (14)

[0080] In formulas (11), (12), (13), (14),

[0081] P1: Pixel value of the m - th feature map after sub - pixel screening at the point (n,p);

[0082] P: Pixel value of the i - th feature map generated by the previous residual structure at the point (j,k);

[0083] m, i: Channel numbers of the feature maps;

[0084] r: Upsampling multiple of the feature maps;

[0085] n, p, j, k: Subscripts corresponding to the width and height of the feature maps;

[0086] 2.1.3 Construct the discriminator network; the discriminator network consists of a convolutional layer, an activation layer, five residual modules, a feature transformation layer, and a fully connected layer in sequence. Among them, the convolutional layer, activation layer, and residual modules are carried out according to formulas (6), (7), (8), (9), and (10) respectively, and the feature transformation layer and fully connected layer are carried out according to formulas (15) and (16):

[0087] E(x1,x2,x3,......x h×w )=F4(P h×w )(15)

[0088] In formula (15),

[0089] E(x1,x2,x3,......x h×w ): Pixel value vector transformed from the feature maps;

[0090] h, w: Height and width of the feature maps;

[0091] F4: Matrix period value conversion function;

[0092] P h×w : Feature map matrix obtained after the convolutional layer, activation layer, and residual modules;

[0093]

[0094] In formula (16),

[0095] E: Pixel value vector transformed from the feature maps by the feature transformation layer;

[0096] F5: Sigmoid function of the fully connected layer;

[0097] 2.2 As Figure 5 Construct the image recognition module of the SROCRN network:

[0098] 2.2.1 Construct the text detection (CTPN) module. The CTPN module consists of a VGG feature extraction layer, a convolutional layer, a BLSTM temporal information fusion layer, and a fully connected layer. Among them, the VGG feature extraction layer uses the VGG-16 network (a convolutional network with 16 specific convolutional kernels), and the convolutional layer and fully connected layer are carried out according to formulas (6) and (16), and the BLSTM temporal information fusion layer is carried out according to formula (17):

[0099] S t = Γ1(S t-1 ) + Γ2(S t-1 ) + S t-1 (17)

[0100] In formula (17),

[0101] S t : The t-th feature map sequence box;

[0102] S t-1 : The (t - 1)-th feature map sequence box;

[0103] Γ1: The forgetting gate processing function of BLSTM, extracting unimportant features in the current feature;

[0104] Γ2: The update gate processing function of BLSTM, extracting the information that needs to be updated in the current feature;

[0105] 2.2.2 Construct the OCR recognition module (CRNN). The CRNN module consists of a convolutional layer, a pooling layer, an RNN sequence feature extraction layer, and a fully connected layer. Among them, the convolutional layer and the fully connected layer are carried out according to formula (6) and formula (16), the RNN sequence feature extraction layer adopts the BLSTM structure and is carried out according to formula (17), and the pooling layer is carried out according to formula (18):

[0106] Q(P2) = w1 * P2 + b1(18)

[0107] In formula (18),

[0108] Q: The pooling layer processing function;

[0109] P2: The feature map extracted by the convolutional layer;

[0110] w1: The convolutional kernel of size f2 × f2 × m, where f2 is the spatial size of the convolutional kernel and m is the number of convolutional kernels;

[0111] b1: An n-dimensional vector;

[0112] 2.3 Construct the loss function of the SROCRN network. The total loss function consists of a super-resolution loss and an image recognition loss. The super-resolution loss consists of a generator loss and a discriminator loss. The image recognition loss consists of a text detection loss and an OCR recognition loss. The loss functions described above are carried out according to formulas (19), (20), and (21) respectively:

[0113] L SROCRN = L SR + L OCR (19)

[0114]

[0115]

[0116] In formula (19),

[0117] L SROCRN : The total loss function of the SROCRN network; L SR : Super-resolution loss;

[0118] L OCR : Image recognition loss;

[0119] In formula (20),

[0120] L GEN : Generator loss; L DEN : Discriminator loss; W1: The width of the image generated by the generator;

[0121] H1: The height of the image generated by the generator; I HR : The real high-resolution image; I LR : The low-resolution image; G θ : The processing function of the generator network; N: The total number of images generated by the generator; D θ : The processing function of the discriminator;

[0122] In formula (21),

[0123] L CTPN : Text detection loss; L OCR : Image recognition loss; N: The total number of images input to the CTPN module; Z S : Cross-entropy loss function; Z q : The offset of the preset text detection box and the text box obtained by actual convolution in the vertical direction; Z m : The offset of the preset text detection box and the text box obtained by actual convolution in the horizontal direction; Z a : The processing function of the text recognition network; s i : The label output by the network detection classification prediction; The real label of the detection classification; v j : The predicted height value of the network detection text box in the vertical direction; The real height value of the preset text box in the vertical direction; o k : The predicted height value of the network detection text box in the horizontal direction; The real height value of the preset text box in the horizontal direction; e f : The text sequence predicted by the text recognition network; The real text sequence label; λ1, λ2: Variable parameters;

[0124] Step 3: Use the dataset to train the SROCRN model

[0125] 3.1 Use the PYTORCH framework based on PYTHON 3.6.5 to build and train the model, and configure the relevant Pytorch environment;

[0126] 3.2 Build the models function and the train function according to the SROCRN network in Step 2;

[0127] 3.3 Import relevant toolkits of PYTHON and TORCH, including torch, torch.optim, torch.nn, torchvision, models, etc.;

[0128] 3.4 Define parameter variables and assign initial values to them. The main variables are as follows:

[0129] dataset = "Text image dataset (obtained in Step 1)"; Dataroot = " / .Data"; workers = 0;

[0130] batchsize = 64; imageSize = 100; upsampling = 4; nepochs = 1000; generatorLR = 0.0001; discriminatorLR = 0.0001; nGPU = 1, etc.;

[0131] 3.5 Import the dataset and encapsulate it into dataset using ImageFolder, and at the same time perform corresponding transform operations, and resize the images in the dataset according to imageSize in 3.4 (size reset operation);

[0132] 3.6 Import the models function and the train function, encapsulate the training set in dataset in 3.5 into dataloader (a dataset container that can randomly extract images according to the batchsize value), and start training according to the parameters initialized in 3.4. Update and save the model parameters once every 100 epochs. The final trained model is saved as SROCRN.pth through the torch.save function;

[0133] Step 4: Recognition of low-resolution text images

[0134] 4.1 Encapsulate the test set in dataset in 3.5 into dataloader, and load the trained SROCRN model in 3.6;

[0135] 4.2 Use the text images in the above test set as input images and input them into the model to obtain the finally recognized text sequence, and at the same time obtain the accuracy rate of the text sequence recognition of the test set text images.

[0136] The recognition result of a single low-resolution text image is as Figure 6 shown. (a) is the original image, (b) is the result of recognizing a single low-resolution text image using the CTPN network, (c) is the result of recognizing a single low-resolution text image using the CRNN network, and (d) is the result of recognizing a single low-resolution text image using the SROCRN network of the present invention;

[0137] The comparison of the recognition rates of batch low-resolution text images of different networks is shown in Table 1. The table respectively counts the recognition accuracies of batch low-resolution text images of the CTPN network, CRNN network, and SROCRN network under different iteration times.

[0138] It can be seen from Table 1 that the SROCRN network has a significant improvement in the accuracy of recognizing a single low-resolution text image sequence compared with the CTPN network and the CRNN network. From Figure 6 it can be seen that the SROCRN network of the present invention has a significant improvement in the recognition accuracy of batch low-resolution text images compared with the CTPN network and the CRNN network. It can be obtained therefrom that the SROCRN network in the present invention can solve the problem of low recognition rate of low-resolution text images and the model has strong adaptability and strong generalization ability.

[0139] Table 1 Comparison of batch recognition accuracies of different recognition networks for low-resolution text images

[0140]

Claims

1. A method for recognizing low - resolution text images based on the SROCRN network, characterized in that The method includes the following steps: Step 1. Construct a dataset for the SROCRN network model: 1.1 Obtain a number of original high-resolution text images with a resolution of W×H, label them, and then divide them into Group A and Group B; 1.2 Perform two image scaling transformations on the images in Group A and Group B collected in 1.1 respectively to obtain low-resolution images with a resolution of 1 / 4*W×1 / 4H, and label them as Group A-1 and Group B-1; 1.3 Label Group A and Group A-1 in 1.2 as the training set, and label Group B-1 as the test set; Step 2. Construct an SROCRN network to recognize the text sequence in the low-resolution image: The SROCRN network includes a super-resolution module and an image recognition module; The super-resolution module adopts an adversarial network, which consists of a generator and a discriminator; the generator consists of a convolutional layer, an upsampling layer, five cascaded residual modules, and two cascaded upsampling layers, where the input of the convolutional layer is the low-resolution text images in Group A-1; the discriminator consists of a convolutional layer, an activation layer, five cascaded residual modules, a feature transformation layer, and a fully connected layer, where the input of the convolutional layer is the output of the activation layer and the original high-resolution text images; The image recognition module adopts a combination of a convolutional network CNN and a short-term memory network RNN, and consists of a text detection CTPN module and a CRNN module; the CTPN module consists of a VGG feature extraction layer, a convolutional layer, a BLSTM temporal information fusion layer, and a fully connected layer, where the input of the VGG feature extraction layer is the output of the super-resolution module; the CRNN module consists of a convolutional layer, a pooling layer, an RNN sequence feature extraction layer, and a fully connected layer, where the input of the convolutional layer of the CRNN module is the output of the CTPN module; In the generator of the super-resolution module, the convolutional layer and the upsampling layer perform convolution and activation operations on the images in the training set to extract feature maps, and this operation is carried out according to Formula (6) and Formula (7): Y = F2(X) = MAX(0, w*F1(X)+b) (6) In Formula (6), X: The training set images in 1.3; F1: The RGB channel separation function of the image; F2: The processing function of the convolution operation; w: A convolution kernel with a size of f1×f1×n, where f1 is the spatial size of the convolution kernel and n is the number of convolution kernels; b: An n-dimensional vector; In Formula (7), P: The feature map after activation processing; Y i,j : The pixel value of the image at point (i, j) after the convolution operation; a i,j : A fixed parameter in the interval (1, +∞); Each of the five residual modules consists of a convolutional layer, a normalization layer, and an activation layer, and further extracts the number of channels of the extended feature map, where the convolutional layer and the activation layer are carried out according to Formulas (6) and (7); the normalization layer is carried out according to Formulas (8), (9), and (10): In Formulas (8), (9), and (10), F3: The normalization processing function; P: The feature map of the image after the previous layer of convolution; u: The mean value of the feature map of the image after the previous layer of convolution; σ 2 : Variance of the feature map of the image after the previous layer of convolution; ε: A variable parameter; H: The width of the feature map; W: The length of the feature map; The last two cascaded upsampling layers quadruple the magnification of the feature map output by the last residual module to generate a high-resolution image. The upsampling layer mainly uses the method of sub-pixel periodic screening and proceeds according to formulas (11), (12), (13), and (14): P1(m,n,p)=P(i,j,k)(11) p=k×r+i%r(14) In formulas (11), (12), (13), and (14), P1: The pixel value of the m-th feature map after sub-pixel screening at the point (n,p); P: The pixel value of the i-th feature map generated by the previous residual structure at the point (j,k); m, i: The channel numbers of the feature maps; r: The upsampling multiple of the feature map; n, p, j, k: The subscripts corresponding to the width and height of the feature map; Step 3. Use the dataset in Step 1 to train and test the SROCRN model in Step 2.

2. The low-resolution text image recognition method based on the SROCRN network according to claim 1, wherein In the discriminator of the super-resolution module in Step 2, the convolutional layer, activation layer, and residual module proceed according to formulas (6), (7), (8), (9), and (10) respectively, and the feature transformation layer and fully connected layer proceed according to formulas (15) and (16): E(x1,x2,x3,......x h×w )=F4(P h×w )(15) In formula (15), E(x1,x2,x3,......x h×w ): The pixel value vector converted from the feature map; h, w: The height and width of the feature map; F4: The matrix periodic value conversion function; P h×w : The feature map matrix obtained after passing through the convolutional layer, activation layer, and residual module; In formula (16), E: The pixel value vector obtained by converting the feature map by the feature transformation layer; F5: The sigmoid function of the fully connected layer.

3. The low-resolution text image recognition method based on the SROCRN network according to claim 2, wherein In the CTPN module of the image recognition module in Step 2, the VGG feature extraction layer uses the VGG-16 network, and the convolutional layer and fully connected layer proceed according to formulas (6) and (16) respectively, and the BLSTM time series information fusion layer proceeds according to formula (17): S t = Γ1(S t-1 ) + Γ2(S t-1 ) + S t-1 (17) In formula (17), S t : The t-th feature map sequence box; S t-1 : The (t - 1)th feature map sequence box; Γ1: The forgetting gate processing function of BLSTM, which extracts the unimportant features in the current feature; Γ2: The update gate processing function of BLSTM, which extracts the information that needs to be updated in the current feature.

4. The low-resolution text image recognition method based on the SROCRN network according to claim 3, wherein In the CRNN module of the image recognition module in Step 2, the convolutional layer and fully connected layer proceed according to formulas (6) and (16) respectively, the RNN sequence feature extraction layer uses the BLSTM structure and proceeds according to formula (17), and the pooling layer proceeds according to formula (18): Q(P2)=w1*P2+b1(18) In formula (18), Q: The pooling layer processing function; P2: The feature map extracted by the convolutional layer; w1: The convolutional kernel of size f2×f2×m, where f2 is the spatial size of the convolutional kernel and m is the number of convolutional kernels; b1: The n-dimensional vector.

5. The low-resolution text image recognition method based on the SROCRN network according to claim 1 or 4, wherein The total loss function of the SROCRN network in Step 2 consists of the super-resolution loss and the image recognition loss. The super-resolution loss consists of the generator loss and the discriminator loss, and the image recognition loss consists of the text detection loss and the OCR recognition loss. The loss functions described above proceed according to formulas (19), (20), and (21) respectively: L SROCRN = L SR + L OCR (19) In formula (19), L SROCRN : The total loss function of the SROCRN network; L SR : Super-resolution loss; L OCR : Image recognition loss; In formula (20), L GEN : Generator loss; L DEN : Discriminator loss; W1: Width of the image generated by the generator; H1: The height of the image generated by the generator; I HR : The real high-resolution image; I LR : The low-resolution image; G θ : Generator network processing function; N: Total number of images generated by the generator; D θ : Discriminator processing function; In formula (21), L CTPN : Text detection loss; L OCR : Image recognition loss; N: The total number of input images to the CTPN module; Z S : Cross-entropy loss function; Z q : The offset of the preset text detection box and the text box obtained by actual convolution in the vertical direction; Z m : The offset of the preset text detection box and the text box obtained by actual convolution in the horizontal direction; Z a : Text recognition network processing function; s i : The label output by the network detection classification prediction; The true label of the detection classification; v j : The predicted height value of the network detection text box in the vertical direction; The true height value of the preset text box in the vertical direction; o k : The predicted height value of the network detection text box in the horizontal direction; The true height value of the preset text box in the horizontal direction; e f : The text sequence predicted by the text recognition network; The true text sequence label; λ1, λ2: Variable parameters.

6. A computer-readable storage medium having a computer program stored thereon, which, when executed on a computer, causes the computer to execute the method according to any one of claims 1-5.

7. A computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • A text image super-resolution reconstruction method based on a conditional generative adversarial network

    CN109410239A

  • Low-resolution license plate recognition method based on generative adversarial network

    CN111461134A