Low-resolution text recognition method based on text position and content guidance

Through the low-resolution text recognition method based on text position and content guidance, the problem of low-resolution text recognition accuracy in complex backgrounds is solved through spatial transformation and multi-module fusion technology, and a more efficient text recognition effect is achieved.

CN120496095APending Publication Date: 2025-08-15DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510620827.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing low-resolution text recognition methods treat character areas and non-character areas equally in the forward process, ignoring the negative impact of complex backgrounds, resulting in low recognition accuracy, especially in complex backgrounds that easily interfere with character positioning and reconstruction.

Method used

A low-resolution text recognition method based on text position and content guidance is adopted, and geometric alignment is performed through the spatial transformation module, combining shallow feature extraction, information generation of text position and content guidance, text area enhancement, direction-aware fusion and upsampling modules, and the text position and content guidance information are used to generate the fused image structure, character position information and semantic information, and the recognition and verification is carried out through the dual-channel verification module.

Benefits of technology

It improves the accuracy of low-resolution text recognition, effectively reduces the interference of complex backgrounds on character recognition, and improves text recognition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496095A_ABST
    Figure CN120496095A_ABST
Patent Text Reader

Abstract

The invention discloses a low-resolution text recognition method based on text position and content guidance. The method comprises the following steps: S1, acquiring a low-resolution text image to be recognized; s2, inputting the low-resolution text image into a pre-trained low-resolution text recognition network model so as to realize text recognition of the low-resolution text image; the method provided by the invention solves the problems that an existing text image super-resolution technology equally treats a character region and a non-character region in a forward process and neglects the negative influence of a complex background, and the background not only has no information amount for a downstream recognition task, but also possibly interferes the positioning and reconstruction of characters, and the recognition efficiency is low. And thus, the low-resolution text recognition precision is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to a low-resolution text recognition method based on text position and content guidance. Background Art

[0002] Currently, existing low-resolution text recognition methods are usually enhanced with text super-resolution technology and can be mainly divided into the following categories:

[0003] The first category is end-to-end CNN-based methods. These methods utilize convolutional neural networks to directly learn the mapping from low-resolution to high-resolution, and are computationally efficient and easy to deploy. Representative works such as SRCNN first introduced CNNs to the super-resolution task, while the subsequent TBSRN designed boundary supervision and gradient loss based on text characteristics, further improving the clarity of character edges.

[0004] The second category is GAN-based methods, which generate more realistic high-frequency details through the adversarial training mechanism of generative adversarial networks. For example, SRGAN first introduced GAN to the super-resolution task, significantly improving visual quality. Subsequent TG and TPGSR further combined text perception modules and semantic guidance to optimize the generation of character structures.

[0005] The third category is Transformer-based methods, which use the self-attention mechanism to model long-range dependencies and are more suitable for processing global text structures and large-scale scaling tasks. Models such as TATT and TSRN significantly improve the reconstruction of irregular text (such as artistic text) by combining the local feature extraction of CNNs with the global context modeling of Transformers.

[0006] In addition, multi-task joint learning methods, frequency domain methods, reference image-based methods, and physical degradation modeling methods have also been widely used in text super-resolution tasks.

[0007] However, existing text image super-resolution techniques treat character and non-character regions equally in their forward pass, ignoring the negative impact of complex backgrounds. These backgrounds not only lack information for downstream recognition tasks but can also interfere with character positioning and reconstruction, leading to low accuracy in low-resolution text recognition. For example, when low-resolution text is placed against a complex background, the complex background can interfere with recognition. This means the image background may be mistakenly identified as characters, leading to incorrect reconstruction of the text. Alternatively, background noise may obscure character outlines, affecting reconstruction quality. Summary of the Invention

[0008] The present invention provides a low-resolution text recognition method based on text position and content guidance to overcome the above technical problems.

[0009] In order to achieve the above object, the technical solution of the present invention is:

[0010] A low-resolution text recognition method based on text position and content guidance, characterized by comprising the following steps:

[0011] S1: Obtain a low-resolution text image to be recognized;

[0012] S2: Inputting the low-resolution text image into a pre-trained low-resolution text recognition network model to achieve text recognition of the low-resolution text image;

[0013] The pre-trained low-resolution text recognition network model includes a spatial transformation module, a shallow feature extraction module, a text position and content-guided information generation module, a text region enhancement module, a direction-aware fusion module, an upsampling module, and a two-way verification module.

[0014] S21: Performing geometric alignment processing on the low-resolution text image through a spatial transformation module to obtain a geometrically aligned output image;

[0015] S22: Perform text feature extraction on the output image after geometric alignment through a shallow feature extraction module to obtain a shallow feature map;

[0016] S23: The information generation module guides the text position and content of the attention map sequence and shallow feature map of the low-resolution text image to obtain a text guidance feature map;

[0017] The attention map sequence is obtained by extracting features from a low-resolution text image according to a text probability distribution using a preset text recognizer;

[0018] S24: Fusing the text guidance feature map and the shallow feature map through the text region enhancement module to obtain a fused feature map;

[0019] S25: extracting direction-aware text features in the vertical and horizontal directions from the fused feature map through the direction-aware fusion module, and obtaining a direction-aware fused feature map based on the direction-aware text features;

[0020] S26: Obtain a super-resolution image by fusing the feature map according to the direction perception based on pixel reshuffling through the upsampling module;

[0021] S27: The direction-aware fusion feature map and the super-resolution image are recognized and verified through a two-way verification module to achieve text recognition of the low-resolution text image.

[0022] Furthermore, the method for obtaining the geometrically aligned output image in S21 includes the following steps:

[0023] S211: extracting features from the low-resolution text image to obtain an output feature map;

[0024] S212: Flatten the output feature map to obtain a two-dimensional tensor of the low-resolution text image;

[0025] S213: Perform a full connection operation on the two-dimensional tensor through the first fully connected layer, and obtain a first intermediate feature vector by combining a batch normalization function and a ReLU activation function;

[0026] S214: Performing a fully connected operation on the first intermediate feature vector through a second fully connected layer to obtain a second intermediate feature vector, i.e., the image after spatial transformation;

[0027] S215: obtaining a set of required text pixel control points regularly arranged on the image plane after spatial transformation according to the size of the low-resolution text image, and obtaining a kernel matrix according to the set of text pixel control points;

[0028] And the formula for obtaining the kernel matrix is

[0029]

[0030] Where: L represents the kernel matrix; N represents the number of required text pixel control points; R represents the distance parameter matrix, and its internal elements d n,n′ represents the Euclidean distance between the desired text pixel control points n,n′ and n∈{1,2,...,N},n′∈{1,2,...,N};

[0031] S216: Constructing a transformation matrix according to the kernel matrix and the second intermediate eigenvector;

[0032] And the transformation matrix is constructed as

[0033]

[0034] Where: W represents the transformation matrix; L -1 represents the inverse kernel matrix; P s represents the required text pixel control point set obtained according to the second intermediate feature vector;

[0035] S217: Obtain a two-dimensional coordinate set of each pixel point according to the low-resolution text image;

[0036] Get the distance parameter r between each pixel point and the required text pixel control point according to the two-dimensional coordinate set q,k , and the distance parameter r q,k The expression is

[0037]

[0038] Where: d q,k Represents the Euclidean distance between the qth pixel and the kth text pixel control point;

[0039] S218: According to the distance parameter r q,k Construct feature vector S q And S q =[1,x q ,y q ,r q,1 ,...,r q,N ], where x q ,y q represents the corresponding position of the qth pixel in the low-resolution image; r q,1 represents the Euclidean distance between the first pixel and the kth text pixel control point; r q,N Represents the Euclidean distance between the Nth pixel and the kth text pixel control point;

[0040] And according to the transformation matrix and eigenvector S q Solve the corresponding positions of the text pixels in the low-resolution text image to obtain the output image after geometric alignment Where B represents the batch size; C represents the number of channels; H and W represent the height and width of the low-resolution text image respectively; represents the channel dimension;

[0041] And the formula for solving the position of text pixel points is

[0042] [x q ,y q ]=S q W.

[0043] Furthermore, the method for obtaining the shallow feature map in S22 is specifically as follows:

[0044] The feature extraction of the output image after geometric alignment is performed through the convolution layer with a preset convolution kernel size, and the activation operation of the output image after feature extraction is performed through the ReLU activation function to obtain the shallow feature map

[0045] Furthermore, the method for obtaining the text guidance feature map in S23 specifically includes the following steps:

[0046] S231: The attention map sequence is obtained by extracting features from the low-resolution text image using a preset text recognizer, and the attention map sequence is recorded as h attn ;

[0047] And obtain the foreground character response map according to the attention map sequence, and the acquisition formula of the foreground character response map is

[0048] h pos =Softmax(Conv(h attn ))

[0049] Where: h pos represents the foreground character response map;

[0050] S232: performing instance normalization processing on the shallow feature map, and obtaining a character perception feature map based on the normalization processing and the foreground character response map;

[0051] And the formula for obtaining the character perception feature map is

[0052] F pos =IN(F v )⊙h pos

[0053] Where: F pos represents the character perception feature map; IN(·) represents the instance normalization function; ⊙ represents the element-wise multiplication operator;

[0054] S234: Taking the attention graph sequence as the semantic basis, the preset self-attention module is used to extract the semantic features of the character sequence in the attention graph sequence, and the semantic features of the character sequence are recorded as h text ;

[0055] S234: Constructing a cascaded cross-attention mechanism network,

[0056] The cascaded cross-attention mechanism network module includes a first cross-attention mechanism module and a second cross-attention mechanism module connected in sequence;

[0057] Based on the cascaded cross-attention mechanism network, the text guidance feature map is obtained according to the semantic features of the character sequence and the character perception feature map;

[0058] Specifically include:

[0059] S2341: Character sequence semantic feature h text As the input of the query vector function Query in the first cross attention mechanism module, the character perception feature map F pos They are respectively used as the input of the key vector function Key and the value vector function Value in the first cross attention mechanism module to obtain the attention output h key ;

[0060] S2342: Output attention h key As the input of the key vector function Key in the second cross attention mechanism module, the character perception feature map Fpos As the input of the query vector function Query in the second cross attention mechanism module and the character sequence semantic feature h text As the input of the value vector function Value in the second cross attention mechanism module to obtain the text guidance feature map

[0061] Furthermore, the method for fusing feature maps in S24 specifically includes the following steps:

[0062] S241: The text guide feature map F g With the shallow feature map F v Connect along the channel dimension to obtain the channel connection feature map;

[0063] The channel connection feature map is projected by setting three convolutional layer modules in parallel to obtain three different spatial projection feature maps and record them as X1, X2, and X3 respectively;

[0064] S242: Obtaining a fusion feature map based on each spatial projection feature map

[0065] And the formula for obtaining the fusion feature map is

[0066] X f =X3+X2⊙Sigmoid(MLP(GDWConv(X1)))

[0067] Where: X f Represents the fused feature map; GDWConv represents the global depth convolution symbol; MLP represents the multi-layer perceptron; ⊙ represents the element-by-element multiplication operator symbol.

[0068] Furthermore, the method for obtaining the direction-aware fusion feature map in S25 includes the following steps:

[0069] S251: Extracting the direction-aware text features in the vertical and horizontal directions from the fused feature map, and obtaining the hidden states of the GRU in the horizontal and vertical directions through the bidirectional GRU module;

[0070] The formula for obtaining the hidden state of the GRU is:

[0071]

[0072] Where: represents a state transition function; Represents the hidden state of GRU in the horizontal and vertical directions respectively; Represents the hidden state of GRU in the horizontal and vertical directions respectively; X rRepresents the feature map after a 3×3 convolution operation; X h ,X v They represent the feature maps in the horizontal and vertical directions after two consecutive 3×3 convolutions; W represents the width of the fused feature map; H represents the height of the fused feature map;

[0073] S252: The hidden state of GRU Splicing is performed on the channel dimension to obtain the splicing feature map, and the splicing feature map is linearly fused through the linear layer to obtain the direction-aware fusion feature map.

[0074] Furthermore, in S27, the direction-aware fusion feature map and the super-resolution image are recognized and verified by a two-way verification module to realize the text recognition method for the low-resolution text image. Specifically,

[0075] Acquire super-resolution images Where r represents the upsampling factor;

[0076] The super-resolution image X SR Fusion feature map F with direction perception s Input them into the preset text recognizer for text recognition and obtain text recognition results p1 and p2;

[0077] Confirm whether the text recognition results p1 and p2 are the same;

[0078] If they are the same, output the recognition result p1;

[0079] Otherwise, the text recognition results p1, p2 are output and the difference positions are marked for subsequent manual verification of the text recognition results.

[0080] Beneficial effect: The present invention provides a low-resolution text recognition method based on text position and content guidance, which realizes text recognition of low-resolution text images by inputting low-resolution text images into a pre-trained low-resolution text recognition network model. That is, the text position and content guidance information is used to generate guidance information that integrates image structure, character position information and semantic information, namely, a text guidance feature map, and the context information is extracted in combination with the text area enhancement module and the direction perception module, thereby improving the text recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0082] Figure 1 is a flow chart of the low-resolution text recognition method of the present invention;

[0083] Figure 2 This is a structural diagram of the spatial transformation network head in this embodiment;

[0084] Figure 3 This is a structural diagram of the text position and content guidance information generation module in this embodiment;

[0085] Figure 4 This is a structural diagram of the text area enhancement module in this embodiment;

[0086] Figure 5 This is a structural diagram of the direction perception fusion module in this embodiment;

[0087] Figure 6 This is a flow chart of the dual-path verification module in this embodiment;

[0088] Figure 7 This is a visualization result diagram of the low-resolution text recognition method in this embodiment. DETAILED DESCRIPTION

[0089] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0090] This embodiment provides a low-resolution text recognition method based on text position and content guidance, which is characterized by specifically comprising the following steps:

[0091] S1: Obtain a low-resolution text image to be recognized;

[0092] S2: Inputting the low-resolution text image into a pre-trained low-resolution text recognition network model to achieve text recognition of the low-resolution text image;

[0093] The pre-trained low-resolution text recognition network model includes a spatial transformation module, a shallow feature extraction module, a text position and content-guided information generation module, a text region enhancement module, a direction-aware fusion module, an upsampling module, and a two-way verification module.

[0094] This embodiment proposes a low-resolution text recognition method based on text position and content guidance, which aims to efficiently recognize low-resolution text images. The flowchart of the low-resolution text recognition method is shown in FIG. Figure 1 As shown, low-resolution images First, the aligned image is obtained through the spatial transformation module Where B represents the batch size, C represents the number of channels, H and W represent the height and width of the image respectively; then the shallow feature extraction module is used to extract shallow features Then, the shallow feature F v First, it is sent to the text position and content guidance information generation module to generate guidance information Afterwards F v With F g After full fusion through the text area enhancement module, the fused features are obtained F f The fused context information is obtained through the direction-aware fusion module F s Get super-resolution images through upsampling modules Where r represents the upsampling factor, and finally, F s and X SR The final recognition result is sent to the two-way verification module. The method specifically includes the following steps:

[0095] S21: Geometrically aligning the low-resolution text image using a spatial transformation module to obtain a geometrically aligned output image. In this embodiment, the spatial transformation module is used to address the misalignment and deformation issues in the text image and mainly includes two parts: a spatial transformation network head and a thin plate spline interpolation.

[0096] The specific steps include:

[0097] S211: extracting features from the low-resolution text image to obtain an output feature map;

[0098] Specifically, the spatial transformation network head structure is as follows: Figure 2 As shown, first, the low-resolution text image x is input into a feature extraction network consisting of 6 consecutive 3×3 convolutional layers and maximum pooling layers to obtain the output feature map The spatial height of the low-resolution text image x has been compressed to 1, and only the feature representation in the horizontal direction is retained;

[0099] S212: Flatten the output feature map to obtain a two-dimensional tensor of the low-resolution text image;

[0100] Specifically, for F STN Perform a flattening operation to convert it into a two-dimensional tensor 8W means that 256 channels are used. The parameter quantity calculated from the spatial compression ratio;

[0101] S213: Perform a full connection operation on the two-dimensional tensor through the first fully connected layer, and obtain a first intermediate feature vector by combining a batch normalization function and a ReLU activation function;

[0102] Specifically, the two-dimensional tensor F′ STN It is then fed into the first fully connected layer, whose output dimension is [B,512], and batch normalization and ReLU activation function are used to generate the first intermediate feature vector with rich semantic information, denoted as

[0103] S214: Performing a fully connected operation on the first intermediate feature vector through a second fully connected layer to obtain a second intermediate feature vector, i.e., the image after spatial transformation;

[0104] Specifically, in order to enhance the stability of training, this embodiment changes F feat Multiply by a coefficient of 0.1 and then send it to the second fully connected layer. The output dimension of this layer is [B, N × 2], where N represents the number of required text pixel control points. Each point contains two coordinate values (x, y). All required text pixel control points are used to obtain the second intermediate feature vector, which is recorded as And P s is input to the thin plate spline interpolation module;

[0105] S215: Obtain the required text pixel control point set arranged regularly on the image plane after spatial transformation according to the size of the low-resolution text image through the thin plate spline interpolation module, denoted as And obtain the kernel matrix according to the text pixel control point set;

[0106] And the formula for obtaining the kernel matrix is

[0107]

[0108] Where: represents the kernel matrix; N represents the number of required text pixel control points; Represents the distance parameter matrix, and its internal elements d n,n′ represents the Euclidean distance between the desired text pixel control points n,n′ and n∈{1,2,...,N},n′∈{1,2,...,N};

[0109] S216: Constructing a transformation matrix according to the kernel matrix and the second intermediate eigenvector;

[0110] And the transformation matrix is constructed as

[0111]

[0112] Where: represents the transformation matrix; L -1 represents the inverse kernel matrix; P s represents the required text pixel control point set obtained according to the second intermediate feature vector;

[0113] S217: Obtain the two-dimensional coordinate set of each pixel according to the size (H, W) of the low-resolution text image The qth pixel is represented by p′ q (x′ q ,y′ q )express;

[0114] Get the distance parameter r between each pixel point and the required text pixel control point according to the two-dimensional coordinate set q,k , and the distance parameter r q,k The expression is

[0115]

[0116] Where: d q,k Represents the Euclidean distance between the qth pixel and the kth text pixel control point; note p q The corresponding point of the text pixel to be solved in the low-resolution text image is p q (x q ,y q );

[0117] S218: According to the distance parameter r q,k Construct feature vector S q And S q =[1,x q ,y q ,r q,1 ,...,r q,N ], where x q ,y q represents the corresponding position of the qth pixel in the low-resolution image; r q,1 represents the Euclidean distance between the first pixel and the kth text pixel control point; r q,N Represents the Euclidean distance between the Nth pixel and the kth text pixel control point;

[0118] And according to the transformation matrix and eigenvector S q Solve the corresponding positions of the text pixels in the low-resolution text image to obtain the output image after geometric alignment

[0119] Where B represents the batch size; C represents the number of channels; H and W represent the height and width of the low-resolution text image respectively; represents the channel dimension;

[0120] And the formula for solving the position of text pixel points is

[0121] [x q ,y q ]=S q W

[0122] This embodiment uses a formula to solve the position of text pixels to calculate the coordinates of each pixel sampling point in the target image (the image after spatial transformation) in the source image (the low-resolution text image to be recognized). The coordinates are range-normalized and fed into a bilinear interpolation sampling function to extract the corresponding pixel values from the source image to obtain the geometrically aligned output image X.

[0123] S22: Perform text feature extraction on the output image after geometric alignment through a shallow feature extraction module to obtain a shallow feature map;

[0124] Specifically, the feature extraction of the output image after geometric alignment is performed through a convolution layer with a preset convolution kernel size (convolution kernel size 9×9), and the activation operation of the output image after feature extraction is performed through the ReLU activation function to obtain the shallow feature map

[0125] S23: The information generation module guides the text position and content of the attention map sequence and shallow feature map of the low-resolution text image to obtain a text guidance feature map;

[0126] The attention map sequence is obtained by extracting features from low-resolution text images according to the text probability distribution through a preset text recognizer; the overall structure of the information generation module is as follows: Figure 3 As shown, the main function is to generate guidance information that integrates text position and content;

[0127] The specific steps include:

[0128] S231: The attention map sequence is obtained by extracting features from the low-resolution text image using a preset text recognizer, and the attention map sequence is recorded as Where T represents the maximum sequence length of the attention map sequence;

[0129] And obtain the foreground character response map according to the attention map sequence, and the acquisition formula of the foreground character response map is

[0130] h pos =Softmax(Conv(h attn ))

[0131] Where: h pos represents the foreground character response map;

[0132] In this embodiment, h is processed by 3×3 convolution.attn Further feature extraction is performed, and then the foreground character response map is obtained through the Softmax operation

[0133] S232: performing instance normalization processing on the shallow feature map, and obtaining a character perception feature map based on the normalization processing and the foreground character response map;

[0134] And the formula for obtaining the character perception feature map is

[0135] F pos =IN(F v )⊙h pos

[0136] Where: F pos represents the character perception feature map; IN(·) represents the instance normalization function; ⊙ represents the element-wise multiplication operator;

[0137] In order to promote the information fusion between image and text features, this embodiment uses the shallow feature map F of the input image to v The instance normalization operation is applied to remove the interference caused by factors such as style and texture in the image, so that the model can focus more on the text content itself. Then, the foreground character response map h pos Perform element-by-element multiplication with the shallow features of the image after instance normalization to obtain the character perception feature map

[0138] S234: Taking the attention graph sequence as the semantic basis, the preset self-attention module is used to extract the semantic features of the character sequence in the attention graph sequence, and the semantic features of the character sequence are recorded as h text ;

[0139] S234: Constructing a cascaded cross-attention mechanism network,

[0140] The cascaded cross-attention mechanism network module includes a first cross-attention mechanism module and a second cross-attention mechanism module connected in sequence;

[0141] Based on the cascaded cross-attention mechanism network, the text guidance feature map is obtained according to the semantic features of the character sequence and the character perception feature map;

[0142] Specifically include:

[0143] S2341: Character sequence semantic feature h text As the input of the query vector function Query in the first cross attention mechanism module, the character perception feature map F pos They are respectively used as the input of the key vector function Key and the value vector function Value in the first cross attention mechanism module to obtain the attention output h key;

[0144] S2342: Output attention h key As the input of the key vector function Key in the second cross attention mechanism module, the character perception feature map F pos As the input of the query vector function Query in the second cross attention mechanism module and the character sequence semantic feature h text As the input of the value vector function Value in the second cross attention mechanism module to obtain the text guidance feature map

[0145] In this embodiment, for semantic information guidance, the text probability distribution output by the pre-trained text recognizer is used as the semantic basis, and the self-attention mechanism is introduced to perform semantic modeling on it, thereby extracting the contextual semantic relationship within the character sequence and obtaining the semantic feature representation. In order to further enhance the matching relationship between character regions and their corresponding semantics, two cascaded cross attention mechanisms are designed. The first cross attention mechanism is based on h text For query, F pos is the key and value, which is used to learn the explicit projection of semantic information on the character image, so that each character can actively locate its response area in the image; then, the attention output With the original semantic feature h text and character features F pos Then perform the second cross attention operation: this time with F pos For query, h key is the key, h text The character perception ability is further enhanced and the semantic features are selectively focused, thereby obtaining the guidance information F that integrates the image structure, character position information and semantic information. g That is, the text guidance feature map;

[0146] S24: Fusing the text guidance feature map and the shallow feature map through the text region enhancement module to obtain a fused feature map, specifically including the following steps:

[0147] S241: The text guide feature map F g With the shallow feature map F v Connect along the channel dimension to obtain the channel connection feature map;

[0148] The channel connection feature map is projected by setting three convolutional layer modules in parallel to obtain three different spatial projection feature maps and record them as X1, X2, and X3 respectively;

[0149] S242: Obtaining a fusion feature map based on each spatial projection feature map

[0150] And the formula for obtaining the fusion feature map is

[0151] X f =X3+X2⊙Sigmoid(MLP(GDWConv(X1)))

[0152] Where: X f Represents the fused feature map; GDWConv represents the global depth convolution symbol; MLP represents the multi-layer perceptron; ⊙ represents the element-by-element multiplication operator symbol.

[0153] In this embodiment, the structure of the text area enhancement module is as follows: Figure 4 As shown, first the shallow feature F v and guidance information F g Connect along the channel dimension, and then project the image features into three different feature spaces through three parallel 1×1 convolutions, which are represented as as well as Then perform the channel attention mechanism on X1 and multiply the resulting attention weight by X2 to generate the channel attention feature, which is added to X3 to obtain the fusion feature Throughout the entire process, in order to better utilize the characteristics of the spatial distribution of character regions in scene text images, this embodiment uses global depth convolution GDWConv;

[0154] S25: extracting direction-aware text features in the vertical and horizontal directions from the fused feature map through the direction-aware fusion module, and obtaining a direction-aware fused feature map based on the direction-aware text features;

[0155] The specific steps include:

[0156] S251: Extracting the direction-aware text features in the vertical and horizontal directions from the fused feature map, and obtaining the hidden states of the GRU in the horizontal and vertical directions through the bidirectional GRU module;

[0157] The formula for obtaining the hidden state of the GRU is:

[0158]

[0159] Where: represents a state transition function; Represents the hidden state of GRU in the horizontal and vertical directions respectively; Represents the hidden state of GRU in the horizontal and vertical directions respectively; Represents the feature map after a 3×3 convolution operation; They represent the feature maps in the horizontal and vertical directions after two consecutive 3×3 convolutions; W represents the width of the fused feature map; H represents the height of the fused feature map;

[0160] The structure of the direction perception fusion module in this embodiment is as follows: Figure 5 As shown, the feature X f After multiple 3×3 convolutions, the vertical and horizontal direction-aware text features are obtained respectively. Then, the vertical and horizontal direction-aware text features are input into the bidirectional GRU module to extract context information and obtain the hidden state of the GRU.

[0161] S252: The hidden state of GRU Splicing is performed on the channel dimension to obtain the splicing feature map, and the splicing feature map is linearly fused through the linear layer to obtain the direction-aware fusion feature map.

[0162] S26: Obtain a super-resolution image by fusing the feature map according to the direction perception based on pixel reshuffling through the upsampling module;

[0163] In a specific embodiment, F s After 3×3 convolution, we get Then the super-resolution image X is obtained by pixel shuffling. SR The pixel reshuffling method adopts the channel-space dimension conversion mechanism to convert Cr 2 The information of the channel dimension is redistributed to the spatial dimension, that is, the channel data in each r×r area is rearranged into corresponding spatial pixels to obtain a super-resolution image, thereby improving the resolution of the feature map without introducing additional learnable parameters;

[0164] S27: Use the dual-path verification module to perform recognition verification on the direction-aware fusion feature map and the super-resolution image to achieve text recognition of low-resolution text images, specifically including

[0165] Acquire super-resolution images Where r represents the upsampling factor;

[0166] The super-resolution image X SR Fusion feature map F with direction perception s Input them into the preset text recognizer for text recognition and obtain text recognition results p1 and p2;

[0167] Confirm whether the text recognition results p1 and p2 are the same;

[0168] If they are the same, output the recognition result p1;

[0169] Otherwise, output the text recognition results p1, p2 and mark the difference positions for subsequent manual verification of the text recognition results;

[0170] In order to ensure the reliability of the recognition result, this embodiment adopts a two-way verification module, the structure of which is as follows: Figure 6 As shown; X SR With F s Input into the same text recognizer to obtain two recognition results, p1 and p2. In this embodiment, ABI Net is preferably used as the text recognizer. If p1 and p2 are consistent, p1 is output; otherwise, both p1 and p2 are output and the differences are marked for subsequent manual verification. The method of marking the differences in this embodiment is a well-known technical means and is not the invention of this application, so it will not be detailed here.

[0171] This embodiment also includes the training process of the low-resolution text recognition network model:

[0172] S100: Obtain batches of low-resolution text images with identification labels and randomly divide them into a data set and a test set;

[0173] S101 performs model training on the constructed low-resolution text recognition network model according to the training set to obtain a trained low-resolution text recognition network model;

[0174] S102: confirming whether the output of the trained low-resolution text recognition network model converges based on the test set, so as to evaluate the trained low-resolution text recognition network model;

[0175] The cross entropy loss function is used to determine whether the output of the trained low-resolution text recognition network model is convergent.

[0176] If it is confirmed that the output of the trained low-resolution text recognition network model converges, then the trained low-resolution text recognition network model is confirmed to be the optimal low-resolution text recognition network model; the optimal low-resolution text recognition network model is the pre-trained low-resolution text recognition network model;

[0177] Otherwise, the parameter weights of the trained low-resolution text recognition network model are adaptively adjusted based on the back-propagation method, and step S101 is repeated.

[0178] Compared with the prior art, the method described in this embodiment performs geometric alignment processing on low-resolution text images through a spatial transformation module to obtain a geometrically aligned output image; performs text feature extraction on the geometrically aligned output image through a shallow feature extraction module to obtain a shallow feature map; performs text position and content guidance on the attention map sequence and shallow feature map of the low-resolution text image through an information generation module to obtain a text guidance feature map; fuses the text guidance feature map and the shallow feature map through a text region enhancement module to obtain a fusion feature map; extracts the vertical and horizontal direction-aware text features in the fusion feature map through a direction-aware fusion module, and obtains a direction-aware fusion feature map based on the direction-aware text features; obtains a super-resolution image based on the direction-aware fusion feature map through an upsampling module based on pixel reshuffling; performs recognition and verification on the direction-aware fusion feature map and the super-resolution image through a two-way verification module to realize text recognition of low-resolution text images; that is, generates guidance information that integrates image structure, character position information and semantic information through text position and content guidance information; and extracts context information in combination with the text region enhancement module and the direction perception module, thereby improving text recognition effect. Figure 7 The figure shows the visualization result of the low-resolution text recognition method, where the left side is a low-resolution text image, the middle part shows the result without using the method proposed in this embodiment, and the right side shows the recognition result using the method proposed in this embodiment. Figure 7 It can be seen that the method described in this embodiment has high accuracy in correct recognition of low-resolution text.

[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A low-resolution text recognition method based on text position and content guidance, characterized in that: The specific steps include: S1: Obtain a low-resolution text image to be recognized; S2: Inputting the low-resolution text image into a pre-trained low-resolution text recognition network model to achieve text recognition of the low-resolution text image; The pre-trained low-resolution text recognition network model includes a spatial transformation module, a shallow feature extraction module, a text position and content-guided information generation module, a text region enhancement module, a direction-aware fusion module, an upsampling module, and a two-way verification module. S21: Performing geometric alignment processing on the low-resolution text image through a spatial transformation module to obtain a geometrically aligned output image; S22: Perform text feature extraction on the output image after geometric alignment through a shallow feature extraction module to obtain a shallow feature map; S23: The information generation module guides the text position and content of the attention map sequence and shallow feature map of the low-resolution text image to obtain a text guidance feature map; The attention map sequence is obtained by extracting features from a low-resolution text image according to a text probability distribution using a preset text recognizer; S24: Fusing the text guidance feature map and the shallow feature map through the text region enhancement module to obtain a fused feature map; S25: extracting direction-aware text features in the vertical and horizontal directions from the fused feature map through the direction-aware fusion module, and obtaining a direction-aware fused feature map based on the direction-aware text features; S26: Obtain a super-resolution image by fusing the feature map according to the direction perception based on pixel reshuffling through the upsampling module; S27: The direction-aware fusion feature map and the super-resolution image are recognized and verified through a two-way verification module to achieve text recognition of the low-resolution text image.

2. The low-resolution text recognition method based on text position and content guidance according to claim 1, characterized in that: The method for obtaining the geometrically aligned output image in S21 specifically includes the following steps: S211: extracting features from the low-resolution text image to obtain an output feature map; S212: Flatten the output feature map to obtain a two-dimensional tensor of the low-resolution text image; S213: Perform a full connection operation on the two-dimensional tensor through the first fully connected layer, and obtain a first intermediate feature vector by combining a batch normalization function and a ReLU activation function; S214: Performing a fully connected operation on the first intermediate feature vector through a second fully connected layer to obtain a second intermediate feature vector, i.e., the image after spatial transformation; S215: obtaining a set of required text pixel control points regularly arranged on the image plane after spatial transformation according to the size of the low-resolution text image, and obtaining a kernel matrix according to the set of text pixel control points; And the formula for obtaining the kernel matrix is Where: L represents the kernel matrix; N represents the number of required text pixel control points; R represents the distance parameter matrix, and its internal elements d n,n′ represents the Euclidean distance between the desired text pixel control points n,n′ and n∈{1,2,...,N},n′∈{1,2,...,N}; S216: Constructing a transformation matrix according to the kernel matrix and the second intermediate eigenvector; And the transformation matrix is constructed as Where: W represents the transformation matrix; L -1 represents the inverse kernel matrix; P s represents the required text pixel control point set obtained according to the second intermediate feature vector; S217: Obtain a two-dimensional coordinate set of each pixel point according to the low-resolution text image; Get the distance parameter r between each pixel point and the required text pixel control point according to the two-dimensional coordinate set q,k , and the distance parameter r q,k The expression is Where: d q,k Represents the Euclidean distance between the qth pixel and the kth text pixel control point; S218: According to the distance parameter r q,k Construct feature vector S q And S q =[1,x q ,y q ,r q,1 ,...,r q,N ], where x q ,y q represents the corresponding position of the qth pixel in the low-resolution image; r q,1 represents the Euclidean distance between the first pixel and the kth text pixel control point; r q,N Represents the Euclidean distance between the Nth pixel and the kth text pixel control point; And according to the transformation matrix and eigenvector S q Solve the corresponding positions of the text pixels in the low-resolution text image to obtain the output image after geometric alignment Where B represents the batch size; C represents the number of channels; H and W represent the height and width of the low-resolution text image respectively; represents the channel dimension; And the formula for solving the position of text pixel points is [x q ,y q ]=S q W。 3. The low-resolution text recognition method based on text position and content guidance according to claim 2, characterized in that: The method for obtaining shallow feature maps in S22 is as follows: The feature extraction of the output image after geometric alignment is performed through the convolution layer with a preset convolution kernel size, and the activation operation of the output image after feature extraction is performed through the ReLU activation function to obtain the shallow feature map 4. The low-resolution text recognition method based on text position and content guidance according to claim 3, characterized in that: The method for obtaining the text guidance feature map in S23 specifically includes the following steps: S231: The attention map sequence is obtained by extracting features from the low-resolution text image using a preset text recognizer, and the attention map sequence is recorded as And obtain the foreground character response map according to the attention map sequence, and the acquisition formula of the foreground character response map is h pos =Softmax(Conv(h attn )) Where: h pos represents the foreground character response map; S232: performing instance normalization processing on the shallow feature map, and obtaining a character perception feature map based on the normalization processing and the foreground character response map; And the formula for obtaining the character perception feature map is F pos =IN(F v )⊙h pos Where: F pos represents the character perception feature map; IN(·) represents the instance normalization function; ⊙ represents the element-wise multiplication operator; S234: Taking the attention graph sequence as the semantic basis, the preset self-attention module is used to extract the semantic features of the character sequence in the attention graph sequence, and the semantic features of the character sequence are recorded as h text ; S234: Constructing a cascaded cross-attention mechanism network, The cascaded cross-attention mechanism network module includes a first cross-attention mechanism module and a second cross-attention mechanism module connected in sequence; Based on the cascaded cross-attention mechanism network, the text guidance feature map is obtained according to the semantic features of the character sequence and the character perception feature map; Specifically include: S2341: Character sequence semantic feature h text As the input of the query vector function Query in the first cross attention mechanism module, the character perception feature map F pos They are respectively used as the input of the key vector function Key and the value vector function Value in the first cross attention mechanism module to obtain the attention output h key ; S2342: Output attention h key As the input of the key vector function Key in the second cross attention mechanism module, the character perception feature map F pos As the input of the query vector function Query in the second cross attention mechanism module and the character sequence semantic feature h text As the input of the value vector function Value in the second cross attention mechanism module to obtain the text guidance feature map 5. The low-resolution text recognition method based on text position and content guidance according to claim 4, characterized in that: The method for fusing feature maps in S24 specifically includes the following steps: S241: The text guide feature map F g With the shallow feature map F v Connect along the channel dimension to obtain the channel connection feature map; The channel connection feature map is projected by setting three convolutional layer modules in parallel to obtain three different spatial projection feature maps and record them as X1, X2, and X3 respectively; S242: Obtaining a fusion feature map based on each spatial projection feature map And the formula for obtaining the fusion feature map is X f =X3+X2⊙Sigmoid(MLP(GDWConv(X1))) Where: X f Represents the fused feature map; GDWConv represents the global depth convolution symbol; MLP represents the multi-layer perceptron; ⊙ represents the element-by-element multiplication operator symbol.

6. The low-resolution text recognition method based on text position and content guidance according to claim 5, characterized in that: The method for obtaining the direction-aware fusion feature map in S25 specifically includes the following steps: S251: Extracting the direction-aware text features in the vertical and horizontal directions from the fused feature map, and obtaining the hidden states of the GRU in the horizontal and vertical directions through the bidirectional GRU module; The formula for obtaining the hidden state of the GRU is: Where: represents a state transition function; Represents the hidden state of GRU in the horizontal and vertical directions respectively; Represents the hidden state of GRU in the horizontal and vertical directions respectively; X r Represents the feature map after a 3×3 convolution operation; X h ,X v They represent the feature maps in the horizontal and vertical directions after two consecutive 3×3 convolutions; W represents the width of the fused feature map; H represents the height of the fused feature map; S252: The hidden state of GRU Splicing is performed on the channel dimension to obtain the splicing feature map, and the splicing feature map is linearly fused through the linear layer to obtain the direction-aware fusion feature map.

7. The low-resolution text recognition method based on text position and content guidance according to claim 6, characterized in that: In S27, the direction-aware fusion feature map and the super-resolution image are recognized and verified by a two-way verification module to realize the text recognition method of the low-resolution text image. Specifically, Acquire super-resolution images Where r represents the upsampling factor; The super-resolution image X SR Fusion feature map F with direction perception s Input them into the preset text recognizer for text recognition and obtain text recognition results p1 and p2; Confirm whether the text recognition results p1 and p2 are the same; If they are the same, output the recognition result p1; otherwise , output the text recognition results p1, p2 and mark the difference positions for subsequent manual verification of the text recognition results.

Citation Information

Cited By

  • Image local enhancement super-resolution method based on text prompt

    CN121169695A