Scene text image super-resolution method based on two-stage reference image guidance
Through a two-stage reference image-guided method, combined with style transfer and block matching, deformable convolution and cross-attention fusion, the problems of insufficient visual prior integration and generalization ability in scene text image super-resolution are solved, and the text recognition accuracy and texture details are improved. It is suitable for practical applications such as license plates, road signs, and tickets.
Patent Information
- Application Number
- CN202511088910.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-09-30
AI Technical Summary
Existing technologies find it difficult to effectively integrate multiple visual priors, deeply explore the feature correspondence between reference images and target images, and accurately capture text content and structural features. In addition, they lack generalization capabilities in scene text image super-resolution and are unable to meet the needs of different application scenarios and text types.
A two-stage reference image-guided method is adopted. In the first stage, the printed text reference image and the low-resolution image features are aligned, and a feature alignment module is combined with style transfer and block matching. In the second stage, deformable convolution and cross-attention fusion are used to generate high-quality super-resolution images through multi-level feature fusion and dynamic adjustment modules.
It significantly improves text clarity and texture details, improves the recognition accuracy of scene text recognition models, enhances the generalization ability of the method, and adapts to the processing capabilities of scenes of different quality and complexity, especially in the fields of license plate, road sign, and bill recognition.
Smart Images

Figure CN120725876A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and image processing, and more particularly, to a scene text image super-resolution method based on two-stage reference image guidance. Background Art
[0002] With the rapid development of deep learning technology in the field of computer vision, especially in image super-resolution reconstruction, image denoising, and image enhancement, the demand for accurate super-resolution processing of scene text images is becoming increasingly urgent. Although traditional super-resolution methods have made some progress in natural scene image reconstruction, deep learning-based super-resolution network design, and image quality assessment, they still have not systematically addressed the core issues unique to scene text images, such as text texture fidelity, semantic information preservation, and recognition accuracy improvement. Therefore, they are unable to meet the actual needs of different application scenarios, different text types, and different image quality.
[0003] Scene text image super-resolution has unique technical challenges: first, it is necessary to accurately restore the texture and content of the text to ensure that the text is clearly legible; second, the evaluation indicators focus more on the recognition accuracy of the scene text recognition model rather than simple image quality indicators; finally, it is necessary to rely on prior knowledge to guide the reconstruction process.
[0004] Existing methods primarily guide the reconstruction process by introducing text priors into the super-resolution backbone network. Specifically, they utilize a scene text recognition model to learn the probability distribution of text in low-resolution images. This probability distribution is then mapped to a feature space using a mapping network and integrated with image features. While this guidance based on text probability distribution is somewhat effective, super-resolution, as a low-level vision task, may be more effectively guided by vision-based priors, a task for which existing methods have not yet fully explored sufficient scope.
[0005] Reference image-based super-resolution technology offers a new approach to solving this problem. This technology not only relies on the information in the low-resolution image itself but also utilizes additional high-resolution reference images to improve reconstruction quality. When similarities exist between the low-resolution image and the reference image, the rich details in the reference image can effectively guide the super-resolution reconstruction process, enabling precise matching of similar regions through a feature alignment module.
[0006] Therefore, how to effectively integrate multiple visual priors, deeply explore the feature correspondence between reference images and target images, accurately capture text content and structural features, and develop a scene text image super-resolution method with strong generalization ability has become a technical problem that needs to be solved urgently. Summary of the Invention
[0007] The present invention provides a scene text image super-resolution method based on two-stage reference image guidance, which solves the technical problems in the existing technology that it is difficult to effectively integrate multiple visual priors, deeply explore the feature correspondence between reference images and target images, accurately capture text content and structural features, and have strong generalization capabilities.
[0008] The present invention provides a scene text image super-resolution method based on two-stage reference image guidance, comprising the following steps:
[0009] Phase 1: Input a low-resolution scene text image into a pre-trained scene text recognition model to obtain the probability distribution of the corresponding text. Based on the probability distribution, a printed text reference image is generated using a drawing library.
[0010] Input the printed text reference image and the low-resolution image into the shared multi-layer reference image feature extraction module to obtain the features of each layer;
[0011] A feature alignment module combining style transfer and block matching is used to align the reference image features with the low-resolution image features.
[0012] The aligned features are upsampled to the same spatial size as the low-resolution image features, fused in series, and input into the corresponding layer of the super-resolution backbone network to guide the generation of super-resolution images;
[0013] Phase 2: The super-resolution image generated in the first phase is used as a new reference image and fed into the triplet input dynamic adjustment module together with the original low-resolution image and the printed text reference image.
[0014] Extract the features of three types of images respectively;
[0015] Use the deformable convolution module to align the features of the first-stage super-resolution image with the low-resolution image;
[0016] Use the cross-attention based feature fusion module to fuse three features;
[0017] The fused features are input into the super-resolution backbone network and trained using the loss function to generate the final super-resolution image.
[0018] As a preferred solution of the present invention, the feature alignment module combining style transfer and block matching includes:
[0019] Style transfer: First, instance regularization is performed on the reference image features to eliminate their own style. Then, the convolutional blocks are used to learn the style mean and variance of the low-resolution image features, and the style distribution of the reference image features is adjusted accordingly.
[0020] Block matching part: Expand the two features into vectors respectively, calculate the cosine similarity to determine the index of the most similar feature, and realize feature alignment based on the index;
[0021] Attention-enhanced alignment mechanism: Attention weights are introduced during the feature alignment stage. The importance of different feature positions is evaluated through the attention network, making the alignment operation more targeted and avoiding noise interference caused by global uniform alignment. This mechanism works in conjunction with the cross-attention mechanism.
[0022] As a preferred solution of the present invention, the operation process of the deformable convolution module includes:
[0023] The offset and modulation mask are learned using the first stage super-resolution image features and low-resolution image features;
[0024] Adaptively deform and align features based on the learned offset and modulation mask, which provides accurately aligned feature inputs to the feature fusion module;
[0025] Since the style difference between super-resolution images and low-resolution images is small, using deformable convolution can achieve more accurate feature alignment, which is different from style transfer methods.
[0026] As a preferred embodiment of the present invention:
[0027] Two cross-attention branches, each fusing one feature into another, receive feature input from the alignment module;
[0028] The importance weights of different feature positions are calculated through the attention mechanism, combined with a multi-scale attention weight generator;
[0029] The output results of the two branches are fused in series and processed through the convolution layer to obtain the final fusion feature;
[0030] Attention-based feature fusion enhancement module: A multi-head attention mechanism is introduced to guide the feature fusion process. During the feature fusion stage, the attention mechanism is used to dynamically adjust the fusion weights of different feature channels or feature maps. The fusion is adaptively performed according to the current task requirements and feature content, thereby improving the expressiveness of the fused features and providing high-quality feature input to the backbone network.
[0031] As a preferred solution of the present invention, the shared multi-layer reference image feature extraction module adopts a deep convolutional neural network structure, which can simultaneously process different types of input images and extract feature representations at multiple scale levels, providing multi-level feature input for the feature alignment module;
[0032] The super-resolution backbone network adopts the TSRN network structure, and effectively guides the super-resolution reconstruction process by injecting the aligned and fused features into multiple layers of the network. The reconstruction process is supervised by a loss function.
[0033] As a preferred solution of the present invention, in the second stage, different feature fusion strategies are adopted for the three features of printed text reference image features, low-resolution image features, and super-resolution image features of the first stage:
[0034] For the printed reference image features and low-resolution image features, the alignment method of style transfer and block matching is continued;
[0035] For the first stage super-resolution image features and low-resolution image features, deformable convolution is used for alignment;
[0036] Effective fusion of three features is achieved through the cross-attention mechanism;
[0037] Triple input dynamic adjustment module: Based on sample attribute analysis, the weight or processing method of each element in the triple input is dynamically adjusted. By designing a classifier or attention weight generator based on sample features, the contribution of each part of the triple input is determined in real time, enabling the model to better adapt to different input situations.
[0038] As a preferred embodiment of the present invention, the loss function of the method includes:
[0039] Pixel reconstruction loss LSR, used to constrain the pixel-level difference between the super-resolution image generated by the super-resolution backbone network and the high-resolution label image;
[0040] Gradient contour loss (LGP) is used to preserve the edge and contour information of the image and to restore the text texture.
[0041] Adversarial loss Ladv, used to improve the texture details and structural information of super-resolution images and enhance text recognition accuracy;
[0042] Perceptual loss Lpercep, based on high-level feature calculation extracted from pre-trained VGG network, complements the feature fusion mechanism;
[0043] The total loss is the weighted sum of the above losses: L total =α1·L SR +α2·L GP +α3·L adv +α4·L percep , to guide the two-stage training process;
[0044] Among them, L total is the total loss function, LSR is the pixel reconstruction loss, L GP is the gradient contour loss, l adv To combat the loss, L percep is the perceptual loss, α1, α2, α3, and α4 are the weight coefficients of pixel reconstruction loss, gradient contour loss, adversarial loss, and perceptual loss, respectively.
[0045] As a preferred embodiment of the present invention, the method is particularly suitable for scene text image super-resolution tasks and can:
[0046] Correctly restore the texture and content of the text, and ensure texture consistency through style transfer;
[0047] Improve the recognition accuracy of super-resolution images in scene text recognition models thanks to feature fusion and loss function design;
[0048] The super-resolution process is guided by visual priors and the feature extraction module is used to obtain visual priors, overcoming the limitations of relying solely on text priors.
[0049] The text recognition effect has been improved in actual application scenarios such as license plate recognition, road sign recognition, and bill recognition, and it can adapt to different application scenarios through a dynamic adjustment mechanism.
[0050] As a preferred solution of the present invention, the attention-based feature alignment and fusion enhancement module further includes:
[0051] Multi-scale Attention Weight Generator: Generates attention weights at different resolution levels, and cooperates with multi-layer feature extraction to enable feature alignment to adapt to multi-scale text structures;
[0052] Adaptive fusion weight regulator: Dynamically adjusts the weight distribution of feature fusion according to the feature complexity and text content of the input sample, providing dynamic weight support for the cross-attention mechanism;
[0053] Feature Importance Evaluation Module: The learned importance scores are used to guide the feature selection and fusion process, improve the discriminability of fused features, and enhance the reconstruction capability of the backbone network.
[0054] As a preferred solution of the present invention, the triplet input dynamic adjustment module further includes:
[0055] Sample attribute analyzer: Analyzes the text complexity, noise level, image quality and other attributes of the input sample to provide an analysis basis for the adaptive fusion weight regulator;
[0056] Weight generation network: Generates dynamic weights for each element of the triplet input based on sample attributes to guide the feature alignment and fusion process;
[0057] Adaptive processing strategy selector: Automatically selects the most suitable feature processing and fusion strategy based on sample characteristics, emphasizes original data for simple samples, and focuses on using reference data for complex samples, working in conjunction with the loss function to achieve personalized super-resolution reconstruction.
[0058] The beneficial effects of the present invention are as follows: Through a two-stage reference image guidance mechanism, the present invention effectively addresses the shortcomings of existing methods in terms of text fidelity and texture detail. The first stage utilizes the structural priors of printed text reference images, combined with a feature alignment module that uses style transfer and block matching, to ensure accurate reconstruction of text structure. The second stage further optimizes texture details through deformable convolution and cross-attention fusion. The resulting super-resolution image shows significant improvements in text clarity, fidelity, and texture detail, effectively overcoming the limitations of super-resolution methods that rely solely on text priors or natural image super-resolution.
[0059] To address the problem that existing methods' evaluation metrics focus on image quality rather than recognition accuracy, this invention better preserves the content and structure of text through precise feature alignment and multi-level feature fusion. The attention-enhanced alignment mechanism avoids the noise interference caused by global uniform alignment, and the cross-attention fusion module effectively utilizes the detailed information of the reference image, significantly improving the recognition accuracy of scene text recognition models. It demonstrates greater competitiveness than traditional methods in tests such as the TextZm dataset.
[0060] By introducing multiple visual prior guidance mechanisms and deeply exploring the feature correspondence between reference and target images, this method effectively addresses the generalization issues faced by existing methods. Its two-stage architecture enables the method to adapt to input images of varying quality. Deformable convolution alignment and cross-attention fusion enhance its processing capabilities for complex scenes, providing a new and effective approach for scene text image super-resolution. This approach is widely applicable to fields such as license plate recognition, street sign recognition, and bill recognition, which require high text image quality and recognition accuracy. This advances the in-depth application and continued development of text image processing technology in real-world scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 Schematic diagram of a super-resolution backbone network and feature fusion of a scene text image super-resolution method based on two-stage reference image guidance provided in an embodiment of the present invention;
[0062] Figure 2 1 is a schematic diagram of a spatial transformation network of a scene text image super-resolution method based on two-stage reference image guidance provided in an embodiment of the present invention;
[0063] Figure 3Schematic diagram of a multi-head cross-attention mechanism for a scene image super-resolution method based on code prediction and text prior guidance provided in an embodiment of the present invention;
[0064] Figure 4 It is a schematic diagram of image processing and feature alignment processing of a scene image super-resolution method based on code prediction and text prior guidance provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.
[0066] At least one embodiment of the present invention discloses a scene text image super-resolution method based on two-stage reference image guidance, such as Figure 1 As shown, the following steps are included:
[0067] Example 1
[0068] The present invention proposes a scene text image super-resolution method based on two-stage reference image guidance, whose overall framework includes a first stage and a second stage.
[0069] Phase 1:
[0070] 1. Input Processing: Input is a low-resolution scene text image ILR and its corresponding binary mask image Im. The concatenation of the two is used to obtain ILRM. The binary mask image is synthesized using a simple clustering algorithm and can roughly display the text area. Because the pixels of the low-resolution and high-resolution image pairs in the training dataset TextZm are not aligned, which affects the learning ability of the model, a spatial transformer network (STN) layer is added as an alignment module. The ILRM passes through the STN layer to obtain the calibrated image ISTN. TextZm is a dataset focused on super-resolution of text images in real scenes.
[0071] 2. Reference image generation: Input ISTN into the pre-trained scene text recognition model to obtain the corresponding text probability distribution; based on the probability distribution, draw the corresponding printed text reference image Irender through the PIL (Pythn Image Library) library; use the common printed text font Arial.
[0072] 3. Feature extraction: Irender and ISTN are input into a shared L-layer reference image feature extraction module. The shared multi-layer reference image feature extraction module uses a deep convolutional neural network structure, specifically including:
[0073] Multi-scale convolution layer: Use convolution kernels of different sizes (1×1, 3×3, 5×5) to extract features of different scales in parallel;
[0074] Downsampling layer: Use maximum pooling or strided convolution to reduce feature dimensionality and retain key information;
[0075] Residual connection: Introducing skip connections in deep networks to prevent the gradient vanishing problem;
[0076] Batch Normalization layer: accelerates training convergence and improves model stability; this module can simultaneously process different types of input images (printed reference images and scene text images) and extract feature representations at multiple scale levels; it provides multi-level feature input to the feature alignment module, including shallow texture information and deep semantic information; the two features obtained by the i-th layer feature extraction module are Fr(i) and FSTN(i).
[0077] Feature alignment: The two features are input into the feature alignment module to obtain Fr(i)align; the feature alignment module combines style transfer and block matching technology (see Example 2 for details).
[0078] Feature fusion and super-resolution reconstruction: Fr(i)align is upsampled to the same spatial size as FSTN(i) and then fused in series. The fused features are input into the corresponding layers of the super-resolution backbone network (TSRN). TSRN uses a deep convolutional neural network structure, including:
[0079] Spatial Transformer Network (STN): performs geometric calibration on the input image to address the pixel misalignment problem of low-resolution and high-resolution image pairs in the dataset;
[0080] Shallow convolution module: extracts shallow features of the image through a 3×3 convolution kernel, including low-level visual information such as edges and textures;
[0081] Multi-layer SRB (Sequential Residual Black) module: specially designed for the characteristics of scene text images, it can effectively extract the sequential information and semantic features of text. Each SRB layer contains a combination of convolutional blocks and GRU blocks;
[0082] Upsampling module: A sub-pixel convolution layer (PixelShuffle) and a convolution layer are combined to improve the spatial resolution of the feature map. By injecting the aligned and fused reference image features into multiple layers of the network, the super-resolution reconstruction process is effectively guided. Finally, a super-resolution image (ISR) is generated.
[0083] Phase 2:
[0084] 1. Model initialization: The model in this stage is initialized using the parameters learned by the corresponding modules in the first stage to accelerate convergence. Most modules in the two stages are the same, but they do not share parameters and are not trained end-to-end. Instead, they are trained separately.
[0085] 2. Multiple reference image input: The reference image feature encoding branch of the second stage receives the printed text reference image Irender, ISTN and ISR triplets; among them, ISR is the output of the super-resolution in the first stage and can provide more visual clues.
[0086] 3. Feature extraction: The method of obtaining Fr(i), FSTN(i) and FSR(i) is the same as the first stage.
[0087] 4. Feature alignment: Fr(i) and FSTN(i) are input to the feature alignment module to obtain the alignment feature Fr(i)align of each layer. Since the background and texture of ISR and ISTN are similar, the deformable convolution module (see Example 3 for details) is used to align the two features FSTN(i) and FSR(i).
[0088] 5. Feature fusion: The two reference features are input into the feature fusion module (see Example 4 for details) for feature fusion; the obtained features are consistent with the first stage and fused into the i-th layer SRB module; ultimately, a higher quality super-resolution image is generated.
[0089] Example 2
[0090] The feature alignment module proposed in this paper combines style transfer and block matching technology to solve the problem that the style difference between the reference image and the low-resolution image is large, making it difficult to effectively align them.
[0091] 1. Style transfer part: The reference image feature Fr(i) is first subjected to instance regularization (InstanceNrm, IN) operation to eliminate its own style; then, image regularization is performed to obtain Fr(i)nrm:
[0092] Fr(i)nrm=γ(IN(Fr(i)))+β
[0093] Among them, γ and β are learned by FSTN(i) through convolution blocks, representing its mean and variance respectively; the Fr(i)nrm feature is adaptively corrected according to γ and β to provide better features for the subsequent block matching process.
[0094] 2. Block matching part: set the block size to K×K, the size of FSTN(i) feature to C×HSTN×WSTN, and the size of Fr(i)nrm feature to C×Hnrm×Wnrm;
[0095] Where K is the side length of the block, C is the feature dimension, HSTN and WSTN are the height and width of the FSTN(i) feature, and Hnrm and Wnrm are the height and width of the Fr(i)nrm feature.
[0096] First, feature expansion is performed on FSTN(i) and Fr(i)nrm with a kernel size of K×K and a step size of 1 to obtain Frefs and Flr respectively;
[0097] Among them, Frefs represents the feature block set after the reference image features are expanded, and Flr represents the feature block set after the low-resolution image features are expanded;
[0098] Calculate the cosine similarity between the kth feature of FSTN(i) and the jth feature in Fr(i)nrm, and use the cosine similarity formula to calculate the similarity of the two feature vectors;
[0099] Where k and j represent the index of the feature block respectively, cosine similarity is used to measure the similarity between two feature vectors; S is the matrix of similarity scores;
[0100] The maximum value of the mth column in the matrix S represents the maximum cosine similarity between the mth feature of FSTN(i) and Fr(i)nrm, and its index represents the index of the most similar feature;
[0101] Among them, m is the feature block index; the aligned features can be determined based on the index; the larger the similarity score, the more similar the two features are. Therefore, the obtained features are multiplied by the corresponding similarity score to obtain the final feature Fr(i)align output by the feature alignment module; the output features are upsampled and fused in series to the corresponding layer of the SRB module.
[0102] 3. Attention-enhanced alignment mechanism: Based on block matching, a multi-level attention weight calculation mechanism is introduced:
[0103] Multi l Eve a =Concat((Local a ,Global a, Context a ))
[0104] Enhanced w (p,q)=softmax(Watt m Multi l Eve a )
[0105] Among them, Multi l Eve a Represents the connection result of multi-level attention features, Local a Focus on local feature similarity, Global a Considering the global feature distribution, Context a Analyze contextual semantic information, Watt m is a learnable weight matrix, p and q are two feature position indices, Enhanced w Represents the similarity representation of different feature position indexes; combines attention weights with similarity calculation.
[0106] Achieve more precise feature alignment:
[0107] S f inal(k,j)=S(k,j)·Enhanced w (k, j) Position e (k, j)
[0108] Among them, Sfinal represents the final similarity matrix, S is the original similarity matrix, k and j are two feature block indexes, Position e Used to encode the position information of features to ensure spatial consistency during alignment.
[0109] Example 3
[0110] In the second stage, since the distributions of FSTN(i) and FSR(i) are close, the present invention uses deformable convolution to align the two features.
[0111] 1. Principle of Deformable Convolution: The kernels of traditional convolution operations are all rectangular in shape, which results in the receptive field of convolutional neural network features also being rectangular, making it difficult to perceive targets of other shapes.
[0112] Deformable convolution adaptively perceives the area of the target object by learning the offset of the convolution block; through deformable convolution, the offset of similar features in FSTN(i) can be well learned, so that the two features can be adaptively aligned.
[0113] 2. Specific implementation: First, offset (i) is learned through FSTN (i) and FSR (i), and the convolution result after the input feature concatenation is activated using the tanh activation function;
[0114] and modulation mask: M(i)=sigmid(conv((FSTN(i), FSR(i))));
[0115] Among them, M(i) represents the modulation mask of the i-th layer, sigmid is the activation function, conv represents the convolution operation, and (FSTN(i), FSR(i)) represent two different feature connections respectively.
[0116] The modulation mask M(i) represents the weight of each feature position, where similar feature positions have higher weights and dissimilar feature positions have lower weights. The two features can then be aligned based on the offset and modulation mask. The FSR(i) feature is processed using a deformable convolution (DCN) operation, combining the offset and modulation mask. DCN stands for deformable convolution.
[0117] 3. Adaptive offset learning enhancement mechanism: To improve the alignment accuracy of deformable convolution, a multi-scale offset learning strategy is introduced:
[0118] O m ulti(i)=∑ s (Weight s tanh(Conv s ((FSTN(i), FSR(i)))))
[0119] Among them, O m ulti(i) represents the multi-scale offset of the i-th layer, s represents different scale levels, and Weight s is a learnable scale weight, tanh is the hyperbolic tangent activation function, Conv s Represents the convolution operation of the sth scale.
[0120] Combine feature similarity to guide offset learning:
[0121] Similarity m =cosine s (FSTN(i), FSR(i))
[0122] Guided o =O m ulti(i)·(1+α·Similarity m )
[0123] Among them, Similarity m Represents feature similarity graph, cosines is the cosine similarity calculation function, Guided o represents the offset after guidance, and α is a learnable parameter used to adjust the influence of similarity on offset learning.
[0124] Example 4
[0125] The feature fusion module proposed in this paper is based on the cross-attention mechanism and is used to fuse the two aligned features.
[0126] 1. Basic module structure:
[0127] It contains two cross-attention branches; first, Fr(i)-dc and Fr(i)align are passed through multi-head attention, and the result is then added with Fr(i)align through a residual connection:
[0128] Ffusel=Fr(i)align+MulAtt(Q=Fr(i)align, K=Fr(i)-dc, V=Fr(i)-dc)
[0129] Among them, Ffusel represents the output of the first fusion branch, Fr(i)align represents the aligned reference image features, Fr(i)-dc represents the features after deformable convolution alignment, Q represents the query vector, K represents the key vector, V represents the value vector, and MulAtt represents the multi-head attention mechanism.
[0130] In this way, the Fr(i)-dc feature is integrated into the Fr(i)align feature. Next, the query, key, and value vectors in the formula are swapped:
[0131]
[0132] Among them, Ffuse2 represents the output of the second fusion branch; in this way, the Fr(i)align feature is fused into the Fr(i)-dc feature; finally, the results of the two branches are fused in series and processed through several stacked convolutional layers.
[0133] 2. Attention-based feature fusion enhancement module: Based on the basic cross-attention mechanism, a multi-head attention mechanism is introduced to further guide the feature fusion process. The query, key, and value vectors are generated respectively through the linear transformation layer, and then the multi-head attention calculation is performed.
[0134] 3. Dynamic adjustment of multi-head attention weights: Adaptively adjust the fusion weights according to the current task requirements and feature content:
[0135] Task w eight=FC task((Comp, Quali))
[0136] Content w eight=FC c ontent((Ffuse1, Ffuse2))
[0137] Dynamic w eight=Softmax(Task w eight+Content w eight)
[0138] Final f use=Dynamic w eight(0) * Ffusel+Dynamic w eight(1) * Ffuse2
[0139] Among them, Task w eigh represents the task-related weight, FC t Ask represents the task-related fully connected layer, Comp represents the feature complexity, Quali represents the quality score, Content w Eight represents the content-related weight, FC c Content represents the content-related fully connected layer, Dynamic w Eight represents the dynamic weight vector, Softmax is the normalization function, Final f use indicates the final fusion result.
[0140] Cooperation with the multi-scale attention weight generator: Works in conjunction with the multi-scale attention weight generator in Example 8 to achieve cross-scale feature fusion and perform weighted summation of fused features at different scales.
[0141] Adaptive feature fusion weight regulator: Dynamically adjusts the weight distribution of feature fusion according to the feature complexity and text content of the input sample:
[0142] Featur=Conv c omple((Ffuse1, Ffuse2))
[0143] Text c on=LSTM t ex((Ffuse1, Ffuse2))
[0144] Adaptiv=Sigmid(FC ada((Featur,Text c on)))
[0145] Final a da=Adaptiv * Ffuse1+(1-Adaptivt)*Ffuse2
[0146] Among them, Feature c om represents feature complexity, Conv c omple represents the complexity analysis convolutional network, Text c on represents the text content weight, LSTM t ex represents the LSTM network for text analysis, Adaptiv represents the adaptive fusion weight, Sigmid is the activation function, FC a da represents the adaptive fully connected layer, Final a da represents the final adaptive fusion result; LSTM t ex is specifically used to analyze text content features and provide semantic guidance for fusion; through adaptive weight adjustment, the fusion process can be dynamically optimized according to different text content and complexity.
[0147] Fusion effect: This mutual use of multi-head cross-attention can fully promote the fusion of the two features; through dynamic weight adjustment, the fusion process can adapt to different feature content and task requirements; so that the fused features contain both the clear texture information of the reference image and the actual features of the low-resolution image; thereby more effectively guiding the generation of super-resolution images, improving image quality, and enhancing text readability and recognition accuracy.
[0148] Example 5
[0149] In this invention, the two stages are trained separately, but the loss functions used are the same, including:
[0150] 1. Pixel reconstruction loss: The L1 norm is used to calculate the pixel-level difference between the super-resolution image ISR and the high-resolution image IHR. Here, ISR represents the super-resolution reconstructed image, IHR represents the high-resolution target image, and L1 norm represents the absolute value loss.
[0151] 2. Gradient contour loss: The L1 norm is used to calculate the difference between the gradient field of the super-resolution image and the gradient field of the high-resolution label image.
[0152] 3. Adversarial loss: The standard adversarial loss formula is used, which includes the discriminator's judgment on the authenticity of super-resolution images and the judgment on high-resolution images.
[0153] 4. Perceptual loss: Based on the lth layer features extracted by the pre-trained VGG-16 network, the L2 norm difference between the super-resolution image and the high-resolution image in the high-level feature space is calculated and normalized. Here, l represents the layer index of the VGG-16 network, and the L2 norm represents the Euclidean distance loss.
[0154] 5. Total loss: The total loss is the weighted sum of the above losses. α1 to α4 are loss balance parameters. The specific setting principles are as follows:
[0155] α1 (pixel reconstruction loss weight): usually set to 1.0, as a basic constraint to ensure basic pixel-level reconstruction quality;
[0156] α2 (gradient contour loss weight): set to 0.1-0.5 to enhance the clarity of text edges and contours;
[0157] α3 (adversarial loss weight): set to 0.01-0.1 to balance the authenticity of generated images and training stability;
[0158] α4 (perceptual loss weight): set to 0.1-0.2 to ensure the consistency of high-level features.
[0159] 6. Loss function and modules work together: the total loss function guides the entire two-stage training process:
[0160] The first stage: loss function supervises the feature alignment module and basic super-resolution reconstruction;
[0161] Second stage: the loss function further optimizes deformable convolution alignment and cross-attention fusion;
[0162] The attention enhancement mechanism learns more accurate attention weights through loss backpropagation;
[0163] The triplet dynamic adjustment module adjusts the weight distribution of different inputs according to the loss change;
[0164] The multi-scale attention weight generator optimizes the multi-scale fusion strategy through loss.
[0165] Example 6
[0166] Based on the feature alignment module, the present invention further proposes an attention-enhanced alignment mechanism to improve the accuracy and effect of feature alignment.
[0167] 1. Attention weight calculation: In the feature alignment stage, the attention network is introduced to learn the importance weights of different feature positions:
[0168] Attenti(u, v)=softmax(Watt·tanh(Wq·Frefs(p)+Wk·Flr(q)))
[0169] Among them, Attenti represents the attention weight, p and q represent the feature position index, softmax is the normalization function, Watt, Wq, Wk are learnable weight matrices, tanh is the hyperbolic tangent activation function, Frefs(p) represents the reference feature indexed by the p feature position, and Flr(q) represents the low-resolution feature indexed by the q feature position.
[0170] Weighted feature alignment: Combining attention weights with cosine similarity to achieve more accurate feature alignment:
[0171] S e nhan(k,j)=S(k,j)·Attenti(k,j)
[0172] Among them, S e nhan represents the enhanced similarity matrix, S represents the original cosine similarity matrix, and k and j represent the feature block indexes; this makes the alignment operation more targeted and avoids noise interference caused by global uniform alignment.
[0173] Multi-level attention calculation mechanism: Introducing a hierarchical attention strategy to calculate attention weights at different levels:
[0174] Local a =Conv l ocal(Concat((Frefs(p),Flr(q))))
[0175] Global a =Glob(Conv g l((Frefs(n), Flr(v))))
[0176] Seman=LSTM s em((Frefs(u), Flr(v)))
[0177] Hierarchy = Softmax(FC h ier((Loca,Global a , Seman)))
[0178] Among them, Local a Represents local attention, Convlocal represents local convolution operation, Concat represents feature connection, p and q represent feature position index, Global a represents global attention, Glob represents global average pooling, Conv gl represents the global convolution operation, Seman represents semantic attention, LSTM s em represents semantic LSTM network, Hierar represents hierarchical weight, FC h ie represents the hierarchical fully connected layer; Loca focuses on local feature similarity matching; Global a Considering the global feature distribution and consistency; Seman uses LSTM to analyze feature associations at the semantic level; and through the hierarchical attention mechanism, it achieves all-round feature alignment from local to global and from bottom to top.
[0179] Cooperation with the cross-attention mechanism: This mechanism cooperates with the subsequent cross-attention mechanism to achieve deep alignment and fusion of features through multi-level attention calculations.
[0180] Example 7
[0181] In response to the three feature processing steps in the second stage, the present invention proposes a triplet input dynamic adjustment module.
[0182] 1. Sample Attribute Analyzer: Analyzes the text complexity, noise level, and image quality of the input sample:
[0183] Complexity=Conv c om((ILR,Im))
[0184] Noiselevel=Conv n oise((ILR,ISTN))
[0185] Quality s core=Conv q uality((ISTN, ISR))
[0186] Among them, Complexity represents complexity, Conv c om represents the complexity analysis convolutional network, ILR represents the low-resolution input image, Im represents the mask image, Noiseievel represents the noise level, Conv n oise represents the noise analysis convolutional network, ISTN represents the image processed by STN, and Quality s core represents the quality score, Conv q uality represents quality analysis convolutional network, and ISR represents super-resolution image.
[0187] Weight generation network: Generates dynamic weights for each element of the triplet input based on sample attributes:
[0188] W r ender=FCr ender((Complexity,Noise l Eve, Quality s core))
[0189] W S TN=FC S TN((Complexity,Noise l Eve, Quality s core))
[0190] W S R=FC S R((Complexity,Noise l Eve, Quality s core))
[0191] Among them, W r ender represents the weight of the printed reference image, FC r ender represents the fully connected layer generated by the reference image weights, W S TN represents the weight of STN processing image, FC S TN represents the fully connected layer generated by STN image weights, W S R represents the super-resolution image weight, FC S R represents the fully connected layer for super-resolution image weight generation.
[0192] Adaptive processing strategy selector: Automatically select the most suitable feature processing and fusion strategy based on sample characteristics:
[0193] For simple samples (low complexity, low noise), the strategy focuses on the original data: W S Increased TN;
[0194] For complex samples (high complexity, high noise), the strategy focuses on using reference data: W r ender and W S R increases.
[0195] Intelligent classifier based on sample features: Introducing deep learning classifiers to intelligently analyze and classify input samples:
[0196] Sampling = CNN c la((ILR,ISTN,ISR))
[0197] Difficult=Regressi((Complexity, Noiselevel, Qualityscore))
[0198] Strat=Decisi((Sample t yPe, Difficulty l evel))
[0199] Among them, Sampl represents the sample type, CNN c la represents CNN classifier, Difficult represents difficulty level, Regressi represents regression network, Strat represents strategy selector, and Decisi represents decision tree; CNN c la is used to identify sample types (such as license plates, road signs, tickets, etc.); Regressi predicts the difficulty level of sample processing; Decisi automatically selects the optimal processing strategy based on the classification results and difficulty level.
[0200] Real-time weight dynamic adjustment mechanism: Design a real-time adjustment mechanism to dynamically determine the contribution of each part of the triplet based on the current input situation:
[0201] Real t im=Atter((W r ender, W S TN, W S R))
[0202] Adap=Real t im(0) * Fr+Real t im(1) * FSTN+Real t im(2) * FSR
[0203] Among them, Real t im represents real-time weight, Atter represents attention weight generator, Adap represents adaptive triplet feature, Fr represents reference image feature, FSTN represents feature after STN processing, and FSR represents super-resolution feature. Atter uses the attention mechanism to generate weight distribution in real time. Through real-time weight adjustment, the model can perform personalized processing based on the specific characteristics of each sample, improving the pertinence and accuracy of the processing effect.
[0204] Example 8
[0205] This paper proposes a multi-scale attention weight generator to generate attention weights at different resolution levels.
[0206] 1. Multi-scale feature extraction: extracting features at different resolution levels:
[0207] F s cale1=Con1x1(F input)
[0208] F s cale2=MaxPool(Con3x3(F i nput))
[0209] F s cale3=MaxPool(MaxPool(Con5x5(F i nput)))
[0210] Among them, F s cale1、F s cale2、F s Cale3 represents the features of the 1st, 2nd, and 3rd scales respectively, Con1x1, Con3x3, and Con5x5 represent the convolution operations of 1×1, 3×3, and 5×5 respectively, and F i nput represents input features, and MaxPool represents the maximum pooling operation.
[0211] Inter-scale attention calculation: Calculate the attention weights between different scales:
[0212] Attmulti=Concat((F s cale1,Upsample(F s cale2), Upsample(F s cale3)))
[0213] Weight m ulti=Sigmoid(Conv1x1(Att m ulti))
[0214] Among them, Att m ulti represents multi-scale attention features, Upsample represents upsampling operation, Weight m ulti represents multi-scale weights and Sigmoid is the activation function.
[0215] Adaptive fusion weight regulator: Dynamically adjusts the weight distribution of feature fusion according to the feature complexity of the input sample:
[0216] Adap=Weight m ulti·(1+δ·Complexityfactor)
[0217] Among them, Adap represents the adaptive weight, δ is the learnable parameter, and Complexityfactor represents the complexity factor, which is determined by the sample complexity.
[0218] Cross-scale feature correlation analyzer: Introducing cross-scale feature correlation analysis to improve the effectiveness of multi-scale fusion:
[0219] Cross s =Correla((F s cale1,F s cale2,F s cale3))
[0220] Enha=Weight m ulti+ε·Cross s
[0221] Final m ult=∑ r (Enhanced w eight(r)·F s cale(r))
[0222] Among them, Cross s represents cross-scale correlation, Correla represents correlation analyzer, Enha represents enhanced weight, ε is a learnable parameter used to adjust the impact of cross-scale correlation on weight, Final m ult represents the final multi-scale feature, r represents the scale index, and ∑ represents the summation operation; Correla analyzes the correlation between features of different scales; through cross-scale correlation analysis, it ensures that the consistency and complementarity of features are maintained during the multi-scale fusion process.
[0223] Example 9
[0224] Implement feature importance evaluation function to guide feature selection and fusion process.
[0225] 1. Importance score calculation: The importance score obtained through learning is used to evaluate the importance of the connected features using the Sigmid activation function.
[0226] 2. Feature selection mechanism: Feature selection is performed based on importance scores, and threshold filtering is used to select important features.
[0227] 3. Enhanced discriminability: Improve the discriminability of fused features and enhance the expression of important features through weighting.
[0228] 4. Multi-dimensional importance evaluation strategy: Evaluate feature importance from multiple dimensions to improve the accuracy of evaluation:
[0229] Spatial i =Conv s ((Fr a , FSTN))
[0230] Channel i =Global a (Conv c ((Fr a , FSTN)))
[0231] Semantic i =LSTM s ((Fr a , FSTN))
[0232] Combined i =Weighted f ((Spatial i , Channel i , Semantic i ))
[0233] Among them, Spatial i Indicates spatial importance, Conv s represents the spatial convolutional analyzer, Fr a represents the reference features after alignment, FSTN represents the features after STN processing, Channel i Indicates the importance of the channel, Global a Represents global average pooling, Conv c Represents channel convolution analyzer, Semantic i Representing semantic importance, LSTM s Represents a semantic LSTM analyzer, Combined i Indicates comprehensive importance, Weighted f Represents weighted fusion operation; Spatial i Evaluate the feature importance of spatial dimensions; Channel i Evaluate the importance of features in the channel dimension; Semantic i Evaluate the importance of features in the semantic dimension; through multi-dimensional evaluation, comprehensively analyze the importance of features and provide more accurate guidance for feature selection and fusion.
[0234] Example 10: The present invention implements a dynamic adjustment mechanism for different application scenarios.
[0235] 1. License Plate Recognition Scenario Adaptation: For license plate recognition scenarios, the dynamic adjustment mechanism includes: a sample attribute analyzer detects the lighting conditions, shooting angle, and blur level of the license plate image; for low-light scenarios, the weight of the printed reference image is increased, and standard fonts are used for guidance; for angled scenarios, the weight of the STN spatial transformation is increased to enhance geometric correction; and the style transfer module adaptively adjusts style parameters based on the license plate background color.
[0236] 2. Road sign recognition scenario adaptation: The multi-scale attention weight generator dynamically adjusts the scale weight according to the size of the road sign text. For road signs with complex backgrounds, the feature importance assessment module increases the feature weight of the text area. The cross-attention fusion module adjusts the fusion strategy according to the text density.
[0237] 3. Adaptation to bill recognition scenarios: The sample attribute analyzer identifies the bill type (invoice, receipt, etc.) and text layout characteristics; the triplet dynamic adjustment module selects the most appropriate processing strategy based on the bill quality; for bills with poor printing quality, the weight of the first-stage super-resolution image ISR is increased.
[0238] The above scene recognition and adaptive strategy adjustment: Introduce a scene intelligent recognition module to automatically identify the application scenario of the input image and adjust the processing strategy:
[0239] Scene t =Scene c ((ILR,ISTN))Scene p =Scene p er p ((Scene t , Quality s ))
[0240] Adaptive s =Strategy m ((Scene t , Scene p ))
[0241] Among them, Scene t Indicates the scene type, Scene c Represents scene classifier, ILR represents low-resolution input image, ISTN represents image processed by STN, Scene p Indicates scene parameters, Scene p er p represents the scene parameter predictor, Quality s Indicates quality score, Adaptive s Represents the adaptive strategy Strategy m Represents a strategy mapper; Scene c Use deep learning models to automatically identify scene types; Scene p er p Prediction scenario-related processing parameters; Strategy mIt means automatically mapping to the optimal processing strategy based on the scenario type and parameters; through intelligent scenario recognition, it achieves true adaptive processing and can optimize for different application scenarios without human intervention.
[0242] like Figure 3 As shown, the high-quality features of code prediction are selectively fused with the intermediate features of the backbone network (the query comes from the attention output, and the key / value comes from the text prior or high-quality features). The figure includes a "multi-head attention module, Q / K / V attention mechanism, and convolution module" (such as the "multi-head attention module-QVK convolution-attention mechanism" combination diagram in the description). This figure shows the interaction process of query, key, and value in the attention mechanism, which matches the implementation details of "multi-head cross attention fusion feature".
[0243] like Figure 4 As shown, the input image needs to be rectified by the Spatial Transformer Network (STN), and scene text recognition relies on feature extraction and alignment of text images. The accompanying figure, which includes the "spatial variation module, scene text recognition module, reference image encoder module, feature alignment module, and upsampling" (such as the combined diagram of the "spatial variation module, scene text recognition module, reference image encoder module, and feature alignment module upsampling"), illustrates the auxiliary process of image preprocessing (rectification), text feature extraction, and alignment, consistent with the requirements for input image processing and feature alignment.
[0244] Through the detailed description of the embodiments above, the technical solution of the present invention can be fully understood and implemented by those skilled in the art, meeting the requirements of the specific implementation method section of the patent application. Each example fully demonstrates the technical advantages, practicality, and scalability of the method, providing a solid technical foundation for the industrial application of the technical solution.
[0245] The present invention can be applied to the following fields:
[0246] 1. Intelligent Transportation Industry: This technology can be used to improve license plate recognition accuracy, enhancing traffic management efficiency and safety. In scenarios such as parking lot management, traffic violation monitoring, and highway toll collection, super-resolution reconstruction of low-resolution license plate images enables more accurate license plate recognition and reduces recognition errors caused by image blur or low resolution.
[0247] 2. Smart Education: This technology can improve the digitization and intelligent processing of educational resources. In online education platforms and educational publishing, it can perform super-resolution reconstruction of low-resolution textbook text images and courseware images, improving text clarity and readability, and enhancing teaching effectiveness and learning experience.
[0248] 3. Smart healthcare: This technology can be applied to medical document management, medical image processing, and other fields. By reconstructing low-resolution images such as medical reports and medical records at super-resolution, it facilitates more accurate reading and analysis by doctors, improving medical efficiency and diagnostic accuracy.
[0249] 4. Document digitization and archive management industry: During the document digitization process, super-resolution reconstruction is performed on a large number of low-resolution historical documents and archival images to improve text recognition accuracy and achieve more efficient digital storage and management.
[0250] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.
Claims
1. A scene text image super-resolution method based on two-stage reference image guidance, characterized by: The following steps are involved: Phase 1: Input a low-resolution scene text image into a pre-trained scene text recognition model to obtain the probability distribution of the corresponding text. Based on the probability distribution, a printed text reference image is generated using a drawing library. Input the printed text reference image and the low-resolution image into the shared multi-layer reference image feature extraction module to obtain the features of each layer; A feature alignment module combining style transfer and block matching is used to align the reference image features with the low-resolution image features. The aligned features are upsampled to the same spatial size as the low-resolution image features, fused in series, and input into the corresponding layer of the super-resolution backbone network to guide the generation of super-resolution images; Phase 2: The super-resolution image generated in the first phase is used as a new reference image and fed into the triplet input dynamic adjustment module together with the original low-resolution image and the printed text reference image. Extract the features of three types of images respectively; Use the deformable convolution module to align the features of the first-stage super-resolution image with the low-resolution image; Use the cross-attention based feature fusion module to fuse three features; The fused features are input into the super-resolution backbone network and trained using the loss function to generate the final super-resolution image.
2. The scene text image super-resolution method based on two-stage reference image guidance according to claim 1 is characterized in that The feature alignment module combining style migration and block matching includes: Style transfer: First, instance regularization is performed on the reference image features to eliminate their own style, and then the convolutional blocks are used to learn the style mean and variance of the low-resolution image features. Adjust the style distribution of the reference image features according to the style mean and variance; Block matching part: Expand the two features into vectors respectively, calculate the cosine similarity to determine the index of the most similar feature, and realize feature alignment based on the index; Attention-enhanced alignment mechanism: Attention weights are introduced during the feature alignment stage. The importance of different feature positions is evaluated through the attention network, making the alignment operation more targeted and avoiding noise interference caused by global uniform alignment. This mechanism works in conjunction with the cross-attention mechanism.
3. The scene text image super-resolution method based on two-stage reference image guidance according to claim 1 is characterized in that The operation process of the deformable convolution module includes: The offset and modulation mask are learned using the first stage super-resolution image features and low-resolution image features; Adaptively deform and align features based on the learned offset and modulation mask, and input the accurately aligned features into the feature fusion module; Since the style difference between super-resolution images and low-resolution images is small, deformable convolution is used to achieve more accurate feature alignment, which is different from the style transfer method.
4. The method according to claim 1, wherein The cross-attention based feature fusion module includes: Two cross-attention branches, each fusing one feature into another, receive feature input from the alignment module; The importance weights of different feature positions are calculated through the attention mechanism, combined with a multi-scale attention weight generator; The output results of the two branches are fused in series and processed through the convolution layer to obtain the final fusion feature; The attention-based feature fusion enhancement module introduces a multi-head attention mechanism to guide the feature fusion process. During the feature fusion stage, the attention mechanism is used to dynamically adjust the fusion weights of different feature channels or feature maps, and adaptively fuses them according to the current task requirements and feature content to provide high-quality feature input to the backbone network.
5. The scene text image super-resolution method based on two-stage reference image guidance according to claim 1, characterized in that: The shared multi-layer reference image feature extraction module adopts a deep convolutional neural network structure to simultaneously process different types of input images and extract feature representations at multiple scale levels, providing multi-level feature input for the feature alignment module; The super-resolution backbone network adopts the TSRN network structure, and effectively guides the super-resolution reconstruction process by injecting the aligned and fused features into multiple layers of the network. The reconstruction process is supervised by a loss function.
6. The scene text image super-resolution method based on two-stage reference image guidance according to claim 1, characterized in that: In the second stage, different feature fusion strategies are used for the three features: printed text reference image features, low-resolution image features, and super-resolution image features from the first stage: For the printed reference image features and low-resolution image features, image feature alignment is performed through style transfer and block matching; For the first stage super-resolution image features and low-resolution image features, deformable convolution is used for alignment; The aligned image features are effectively fused with the three features through the cross-attention mechanism; According to the analysis of sample attributes, the weight or processing method of each element in the triple input is dynamically adjusted. By designing a classifier or attention weight generator based on sample features, the contribution of each part of the triple input is determined in real time, making the model better adapted to different input situations.
7. The scene text image super-resolution method based on two-stage reference image guidance according to claim 1, characterized in that: The loss function of the method includes: Pixel reconstruction loss LSR, used to constrain the pixel-level difference between the super-resolution image generated by the super-resolution backbone network and the high-resolution label image; Gradient contour loss (LGP) is used to preserve the edge and contour information of the image and to restore the text texture. Adversarial loss Ladv, used to improve the texture details and structural information of super-resolution images and enhance text recognition accuracy; Perceptual loss Lpercep, based on high-level feature calculation extracted from pre-trained VGG network, complements the feature fusion mechanism; The total loss is the weighted sum of the above losses: L total =α1·L SR +α2·L GP +α3·L adv +α4·L percep , to guide the two-stage training process; Among them, L total is the total loss function, L SR is the pixel reconstruction loss, L GP is the gradient contour loss, L adv To combat the loss, L precep is the perceptual loss, α1, α2, α3, and α4 are the weight coefficients of pixel reconstruction loss, gradient contour loss, adversarial loss, and perceptual loss, respectively.
8. The scene text image super-resolution method based on two-stage reference image guidance according to claim 1, characterized in that: The method is particularly suitable for scene text image super-resolution tasks, including: Restore the texture and content of the text and obtain consistent texture and content through style migration; Through feature fusion and loss function design, super-resolution images are accurately recognized on the scene text recognition model; The super-resolution process is guided by visual priors and the feature extraction module is used to obtain visual priors, overcoming the limitations of relying solely on text priors. In the actual application scenarios of license plate recognition, road sign recognition, and bill recognition, a dynamic adjustment mechanism is used to obtain text recognition effects to adapt to different application scenarios.
9. The scene text image super-resolution method based on two-stage reference image guidance according to claim 1, characterized in that: The attention-based feature alignment and fusion enhancement module also includes: Multi-scale Attention Weight Generator: Generates attention weights at different resolution levels, and works with multi-layer feature extraction to align features to adapt to multi-scale text structures. Adaptive fusion weight regulator: Dynamically adjusts the weight distribution of feature fusion according to the feature complexity and text content of the input sample to obtain the dynamic weight of the cross-attention mechanism; Feature importance evaluation module: The importance scores obtained through learning are used to guide the feature selection and fusion process to obtain the discriminability of the fused features, thereby enhancing the reconstruction ability of the backbone network.
10. The scene text image super-resolution method based on two-stage reference image guidance according to claim 7, characterized in that: The triplet input dynamic adjustment module also includes: Sample attribute analyzer: analyzes the text complexity, noise level and image quality attributes of the input sample; Weight generation network: Generates dynamic weights for each element of the triplet input based on sample attributes to guide the feature alignment and fusion process; Adaptive processing strategy selector: Automatically selects corresponding feature processing and fusion strategies based on sample characteristics, emphasizes original data for simple samples, and focuses on using reference data for complex samples, working in conjunction with the loss function to achieve super-resolution reconstruction.