Weak supervision super-resolution method based on bimodal prior

Through the weakly supersupervised super-resolution method based on the dual-modal prior, the text semantic information and spatial dimension information of the scene text image are extracted, and global feature fusion and multi-level supervision optimization are carried out, which solves the problem of insufficient accuracy of the scene text image super-resolution method in real image processing and reliance on high-quality annotation data in the prior art, and achieves high-quality image reconstruction and text recognition performance improvement.

CN120147128AInactive Publication Date: 2025-06-13NANTONG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510217194.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing scene text image super-resolution methods have poor accuracy when processing real images, especially in spatially deformed text and complex backgrounds, and rely on high-quality labeled data, making it difficult to achieve good generalization capabilities in small samples or without labeling.

Method used

The weak-supervised super-resolution method based on the dual-modal prior is adopted, and text semantic information and spatial dimension information are extracted through dual-modal prior refinement, combined with the self-attention mechanism of the visual transformer for global feature fusion, and multi-level supervision and optimization are used for structural similarity loss function and connectionist timing classification loss function.

Benefits of technology

The reconstruction quality of super-resolution images is significantly improved, especially in the recovery effect of complex scenes and text areas. It can improve the average recognition accuracy of scene text with fewer high-resolution labels and optimize the reconstruction quality of images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147128A_ABST
    Figure CN120147128A_ABST
Patent Text Reader

Abstract

The invention discloses a weak supervision super-resolution method based on bimodal prior, and relates to the field of computer vision, and the method comprises the following steps: obtaining a scene text image, carrying out the feature extraction of the scene text image, and constructing an initial image super-resolution frame based on the extracted image features; the initial image super-resolution framework is optimized through a structural similarity loss function and a connection dominant time sequence classification loss function in combination with a back propagation mechanism, and a final image super-resolution framework is obtained; and inputting a pre-acquired scene text image into the final image super-resolution frame to obtain a reconstructed scene text image. According to the invention, through bimodal prior refinement processing, capture of a global structure and semantic information of a scene text image is enhanced, so that reconstruction quality of a super-resolution image is improved; and meanwhile, through global feature fusion processing, the overall quality of the generated high-resolution image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and more particularly, to a weakly supervised super-resolution method based on bimodal priors. Background Art

[0002] Scene text image super-resolution is a technology that focuses on reconstructing the quality of scene text images; compared with traditional single-image super-resolution methods, scene text image super-resolution pays more attention to the restoration of text regions; this technology has a wide range of applications, especially in fields such as intelligent transportation, text retrieval, and image caption generation; in recent years, with the rapid development of deep learning, text recognition technology has made significant progress by adopting methods such as visual feature extraction, language modeling, and loss function optimization.

[0003] Traditional scene text image super-resolution models are usually trained on synthetic data, but due to the fact that real images are often affected by factors such as complex backgrounds, noise, and lighting, these models show poor accuracy when reconstructing real images; in recent years, researchers have first proposed a real scene text super-resolution dataset and developed a text super-resolution network based on this dataset, effectively improving the accuracy of text recognition; however, this network still faces many challenges when dealing with spatially deformed text; for this reason, researchers have added a global attention mechanism to the deep convolutional neural network and proposed a text attention network, which achieved leading performance in the text super-resolution task of that year; however, the text attention network relies on a large number of high-quality real text images, and it is difficult to obtain these images in reality, and the cost of manually annotating data is extremely high, so this method is difficult to achieve good generalization ability in the case of small samples or unlabeled data.

[0004] In recent years, research has proposed a weakly supervised framework using coarse-grained text labels, introducing a weakly supervised learning framework in the scene text image super-resolution task; however, this framework performs poorly in high-noise or complex scenes, mainly limited by the insufficient utilization of scene prior information.

[0005] Regarding the problems in the related art, no effective solution has been proposed yet. Summary of the Invention

[0006] In view of the problems in the related art, the present invention proposes a weakly supervised super-resolution method based on bimodal priors to overcome the above-mentioned technical problems existing in the existing related art.

[0007] To this end, the specific technical solution adopted by the present invention is as follows:

[0008] A weakly supervised super-resolution method based on bimodal priors, the method comprising the following steps:

[0009] S1. Obtain a scene text image, extract features from the scene text image, and construct an initial image super-resolution framework based on the extracted image features;

[0010] S2. Optimize the initial image super-resolution framework through a structural similarity loss function and a connectionist temporal classification loss function, combined with a backpropagation mechanism, to obtain a final image super-resolution framework;

[0011] S3. Input the previously obtained scene text image into the final image super-resolution framework to obtain a reconstructed scene text image.

[0012] Further, obtaining a scene text image and extracting features from the scene text image and constructing an initial image super-resolution framework based on the extracted image features includes the following steps:

[0013] S11. Obtain a scene text image and, through dual-modal prior refinement processing, extract text semantic information and spatial dimension information;

[0014] S12. Perform global feature fusion processing on the extracted text semantic information and spatial dimension information through the self-attention mechanism of a vision transformer to obtain fusion information, and add the fusion information and the spatial dimension information element by element to obtain the comprehensive features of the image;

[0015] S13. Based on the comprehensive features of the image, use a number of sequential residual blocks in cascade, combined with sub-pixel convolution, to perform reconstruction processing on the image to obtain an initial image super-resolution framework.

[0016] Further, extracting text semantic information and spatial dimension information through dual-modal prior refinement processing includes the following steps:

[0017] Extract the text semantic information in the scene text image through the text prior branch processing of dual-modal prior refinement; the text prior branch processing includes small-sample text recognition processing and convolution processing;

[0018] Extract the spatial dimension information in the scene text image through the structural prior branch processing of dual-modal prior refinement; the structural prior branch processing includes centralized feature pyramid processing and squeeze-and-excitation network processing.

[0019] Further, the expression of the text prior branch processing is:

[0020] P text =g text (I LR );

[0021] In the formula, P text represents the text semantic information; g text represents the text prior; I LRRepresents the scene text image;

[0022] The expression processed by the structural prior branch is:

[0023] P structure = g structure (I LR );

[0024] In the formula, P structure represents the spatial dimension information; g structure represents the structural prior.

[0025] Furthermore, the extracted text semantic information and spatial dimension information are globally feature fused through the self-attention mechanism of the vision transformer to obtain the fused information, and the fused information is added element-wise to the spatial dimension information to obtain the comprehensive features of the image, including the following steps:

[0026] S121. Based on the self-attention mechanism of the vision transformer, divide the scene text image into several image patches, add position encoding to each image patch, and calculate the attention scores of each image patch;

[0027] S122. Generate a self-attention mapping function according to all the image patch attention scores, and combine the extracted text semantic information and spatial dimension information to obtain the fused information;

[0028] S123. Add the fused information and the spatial dimension information element-wise to obtain the comprehensive features of the image. Furthermore, the expression of the self-attention mapping function is:

[0029] P fused = φ(P text + P structure );

[0030] In the formula, P fused represents the fused information; φ represents the self-attention mapping function; P text represents the text semantic information; P structure represents the spatial dimension information.

[0031] Furthermore, optimize the initial image super-resolution framework through the structural similarity loss function and the connectionist temporal classification loss function, and combine the backpropagation mechanism to obtain the final image super-resolution framework, including the following steps:

[0032] S21. Generate the total loss function through the structural similarity loss function and the connectionist temporal classification loss function;

[0033] S22. Based on the initial image super-resolution framework and combine the total loss function to obtain the total loss value;

[0034] S23. Optimize the initial image super-resolution framework using the backpropagation mechanism according to the total loss value to obtain the final image super-resolution framework.

[0035] Further, the expression of the structural similarity loss function is:

[0036]

[0037] In the formula, SSIM(x, y) represents the structural similarity loss function; μ x represents the mean of image x; μ y represents the mean of image y; represents the variance of image x; represents the variance of image y; σ xy represents the covariance of images x and y; C 1 represents the calculation constant one; C 2 represents the calculation constant two.

[0038] Further, the expression of the connectionist temporal classification loss function is:

[0039]

[0040] In the formula, L CTC represents the negative log average of probabilities; l i represents the corresponding correct label of the input image; y i represents the given arrival path; i represents the i-th input image; N represents the batch size of the input images; p(l i |y i ) represents the probability of the target sequence.

[0041] Further, the expression of the total loss function is:

[0042]

[0043] In the formula, L represents the total loss function; L rec represents the reconstruction loss; λ represents the hyperparameter that weighs the structural similarity loss and the connectionist temporal classification loss; SSIM(y i , G(x i )) represents the structural similarity loss between the i-th input image and the reconstructed scene text image; p(l i |G(x i )) represents the connectionist temporal classification loss between the input image and the corresponding reconstructed scene text image; x i represents the x i -th input image; G(x i ) represents the reconstructed scene text image.

[0044] The beneficial effects of the present invention are as follows:

[0045] 1. Through the dual-modal prior refinement process, the present invention enhances the capture of the global structure and semantic information of the scene text image, thereby improving the reconstruction quality of the super-resolution image. At the same time, through the global feature fusion process, the overall quality of the generated high-resolution image is improved, especially in the restoration effect of complex scenes and text regions. In combination with the multi-level supervision strategy of the structural similarity loss function and the connectionist temporal classification loss function, the image structure and semantic information can be fully captured to help improve the text recognition performance.

[0046] 2. Through the dual-modal prior refinement process, global feature fusion process and reconstruction process, the present invention can effectively capture the multi-scale global and detailed features of the image when only using 50% of the high-resolution labels, significantly improving the average recognition accuracy of scene text and optimizing the reconstruction quality of the image at the same time.

[0047] 3. Through the text prior and structure prior branch processes, the present invention combines text semantics with image structure features, thereby being able to capture multi-scale global features and detailed information simultaneously, effectively improving the reconstruction performance of text regions. And with the multi-level supervision mechanism of the structural similarity loss function and the connectionist temporal classification loss function, the joint optimization of low-level features and high-level semantics is realized, thus significantly improving the accuracy and robustness of scene text recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 is a flowchart of a weakly supervised super-resolution method based on dual-modal prior according to an embodiment of the present invention;

[0050] Figure 2 is a framework diagram of a weakly supervised super-resolution method based on dual-modal prior according to an embodiment of the present invention;

[0051] Figure 3 is a schematic diagram of CFP in the dual-modal prior refinement process of a weakly supervised super-resolution method based on dual-modal prior according to an embodiment of the present invention;

[0052] Figure 4 is a visualization image comparison diagram of the TextZoom dataset of an embodiment of a weakly supervised super-resolution method based on dual-modal prior according to an embodiment of the present invention. Detailed implementation manners

[0053] To further illustrate each embodiment, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. They are mainly used to illustrate the embodiments and can be combined with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention.

[0054] According to an embodiment of the present invention, a weakly supervised super-resolution method based on a bimodal prior is provided.

[0055] Now, the present invention will be further described in conjunction with the accompanying drawings and specific implementation manners. As Figure 1 shown, the weakly supervised super-resolution method based on a bimodal prior according to an embodiment of the present invention includes the following steps:

[0056] S1. Obtain a scene text image, perform feature extraction on the scene text image, and construct an initial image super-resolution framework based on the extracted image features.

[0057] Specifically, obtaining a scene text image, performing feature extraction on the scene text image, and constructing an initial image super-resolution framework includes the following steps:

[0058] S11. Obtain a scene text image, and through bimodal prior refinement processing, extract text semantic information and spatial dimension information.

[0059] Specifically, extracting text semantic information and spatial dimension information through bimodal prior refinement processing includes the following steps:

[0060] Through the text prior branch processing refined by the bimodal prior, extract the text semantic information in the scene text image; the text prior branch processing includes small sample text recognition processing and convolution processing;

[0061] Through the structure prior branch processing refined by the bimodal prior, extract the spatial dimension information in the scene text image; the structure prior branch processing includes centralized feature pyramid processing and squeeze-and-excitation network processing.

[0062] Specifically, the expression of the text prior branch processing is:

[0063] P text = g text (I LR ) ;

[0064] In the formula, P text represents the text semantic information; g text represents the text prior; I LR represents the scene text image;

[0065] The expression processed by the structural prior branch is:

[0066] P structure = g structure (I LR );

[0067] In the formula, P structure represents the spatial dimension information; g structure represents the structural prior.

[0068] It should be noted that I LR represents the scene text image, that is, a low-resolution scene text image containing the target text is obtained.

[0069] It should be noted that the text prior branch processing can extract the text semantic information in the image, aiming to enhance the recognition and restoration ability of the initial image super-resolution framework for the text area and improve the restoration effect of text details; the centralized feature pyramid processing (CFP) in the structural prior branch processing can effectively enhance the global information transmission and fusion between multi-scale features and effectively enhance the global expression ability of features; while the squeeze-and-excitation network processing (SENet) can adaptively generate channel weights, improve the model's attention to key feature channels, and suppress redundant information; the combination of CFP and SENet strengthens the feature representation in the spatial dimension and channel dimension, can more accurately capture the details and structural features of the image, and thus significantly improves the overall quality of super-resolution reconstruction.

[0070] It should be noted that by constructing the text prior and structural prior branch processing, multi-scale global features and detail information can be captured simultaneously, effectively improving the reconstruction performance of the text area.

[0071] S12. Perform global feature fusion processing on the extracted text semantic information and spatial dimension information through the self-attention mechanism of the vision transformer to obtain fusion information, and add the fusion information and the spatial dimension information element by element to obtain the comprehensive features of the image.

[0072] Specifically, performing global feature fusion processing on the extracted text semantic information and spatial dimension information through the self-attention mechanism of the vision transformer to obtain fusion information, and adding the fusion information and the spatial dimension information element by element to obtain the comprehensive features of the image includes the following steps:

[0073] S121. Based on the self-attention mechanism of the vision transformer, divide the scene text image into several image patches, add position encoding to each image patch, and calculate the attention score of each image patch;

[0074] S122. Generate a self-attention mapping function based on all image patch attention scores, and combine the extracted text semantic information and spatial dimension information to obtain fused information.

[0075] Specifically, the expression of the self-attention mapping function is:

[0076] P fused = φ(P text + P structure );

[0077] In the formula, P fused represents the fused information; φ represents the self-attention mapping function; P text represents the text semantic information; P structure represents the spatial dimension information.

[0078] It should be noted that P fused represents the fused information, that is, the preliminary fusion prior obtained after the text semantic information and the spatial dimension information are fused by the vision transformer.

[0079] S123. Add the fused information and the spatial dimension information element by element to obtain the comprehensive features of the image.

[0080] It should be noted that as Figure 3 shown are the specific implementation steps of the bimodal prior refinement process. The process inputs the image from left to right and first performs feature extraction through a series of convolutional layers. These convolutional layers include downsampling (2x downsampling) and upsampling (2x upsampling) operations, as well as batch normalization, activation functions (such as SiLU), etc. During the feature extraction process, there is also a top-down path for fusing feature maps at different levels to enhance the model's ability to capture detailed information. After the feature extraction is completed, the feature map is further processed by the explicit visual center (LVC) module, and then the processed features are input into the classifier (Cls) and the regressor (Reg) for classification tasks and regression tasks respectively. In addition, the bimodal prior refinement process also includes a feature map fusion mechanism for optimizing feature representations. The entire process aims to improve the model's performance in image recognition and localization tasks through multi-scale feature fusion and the lightweight multi-perception machine (LVC) module.

[0081] The specific steps of dual-modal prior refinement are as follows: First, multi-scale feature extraction, that is, extracting multi-scale feature maps from different stages of the backbone network as inputs, and these feature maps contain different resolutions and semantic information; Second, hierarchical feature alignment, that is, using upsampling or downsampling operations to unify the resolutions of the multi-scale feature maps for feature map fusion in subsequent steps; among them, two-times upsampling is to upsample the low-resolution feature map to align it with the high-resolution feature map; two-times downsampling is to downsample the high-resolution feature map to align it with the low-resolution feature map; Third, centralized feature fusion, that is, concentrating the aligned multi-scale features onto a shared centralized feature map through a fusion module; the specific fusion process includes weighted fusion, convolutional batch normalization operation, and global context modeling; among them, weighted fusion is to assign different weights to the multi-scale features and sum them up; the convolutional operation is to perform convolutional processing on the fused features to enhance their expressive ability; global context modeling is to extract cross-scale context semantic information through a global information enhancement module; Fourth, feature enhancement, that is, using the attention mechanism to further process the centralized features to highlight key regions and channel information; among them, channel attention emphasizes channel features with significant semantics, and spatial attention focuses on the spatial region where the target is located; Fifth, distributing enhanced features, that is, redistributing the enhanced centralized features into feature maps of different scales to ensure that the multi-scale features contain unified enhanced semantic information, that is, using upsampling or downsampling operations for each scale to restore the original resolution of the feature map; Sixth, output multi-scale features, outputting the enhanced multi-scale feature maps, and these features will be used as subsequent inputs.

[0082] It should be noted that Figure 3 In it, LVC is an abbreviation. LVC is an encoder with an inherent dictionary, consisting of an inherent codebook and a set of learnable visual center scale factors; the specific process is as follows: First, use a set of convolutional layers to encode the input features and further process them using CBR blocks; Second, combine the encoded features with the inherent codebook through a set of learnable scale factors; then, use a fully connected layer and a 1×1 convolutional layer to predict prominent key class features; finally, perform channel multiplication and channel addition on the local angular region features of the input features and scale factor coefficients.

[0083] It should be noted that by performing global feature processing on image patches through the self-attention mechanism, the key details in the low-resolution image can be accurately restored.

[0084] It should be noted that by processing image patches through the self-attention mechanism, the key details in the low-resolution image can be accurately restored.

[0085] It should be noted that the global feature fusion processing processes the image blocks through the self-attention mechanism, which can efficiently capture the global features in multimodal information, that is, the global features of text semantic information and spatial dimension information, and then add the fused information to the spatial dimension information obtained a priori from the graphic structure element by element, so that the initial image super-resolution framework is further strengthened in terms of global structure and detail modeling.

[0086] It should be noted that the global feature fusion processing significantly improves the ability of the initial image super-resolution framework in restoring key details of low-resolution images, especially in the reconstruction effect of complex scenes and text areas, and the generated high-resolution images have better quality.

[0087] It should be noted that the use of the self-attention mechanism can minimize information loss and thus enhance image restoration capabilities in complex scenes.

[0088] It should be noted that the self-attention mechanism of the visual transformer is used to perform global feature fusion processing, and the steps of obtaining fusion information include: dividing the image into several image blocks; and performing embedding representation, that is, mapping each image block to a vector of fixed dimension through linear projection; then performing position encoding, that is, adding position encoding to each image block embedding to retain position information; then performing self-attention calculation, that is, generating query (Q), key (K), value (V) vectors, calculating attention scores, that is, calculating the similarity of Q and K through dot product, that is, normalizing the attention scores (Softmax), obtaining weights, and weighted summing to obtain the output value (the weighted sum of V); finally, aggregating the outputs, aggregating the attention outputs of all image blocks into an overall feature representation.

[0089] S13. Based on the comprehensive features of the image, several sequential residual blocks are cascaded and combined with sub-pixel convolution to reconstruct the image and obtain an initial image super-resolution framework.

[0090] It should be noted that the initial image super-resolution framework includes bimodal prior refinement processing, global feature fusion processing and reconstruction processing.

[0091] It should be noted that based on the comprehensive features of the image, the image is reconstructed by cascading several sequential residual blocks and combining them with sub-pixel convolution, which can deeply mine the image features and achieve the output of high-resolution images. At the same time, the sequence information is captured through the bidirectional long short-term memory network (BLSTM). Among them, the cascade operation avoids information loss and directly maps multi-scale features to the high-resolution space through upsampling operations, thereby improving the reconstruction accuracy and significantly reducing the generation of artifacts.

[0092] It should be noted that the reconstruction process is realized by cascading multiple sequential residual blocks (SRBs) and through sub-pixel convolution operations. The sequence information is captured by a bidirectional long short-term memory network (BLSTM), and the multi-scale feature maps are directly mapped to the high-resolution space through upsampling operations.

[0093] It should be noted that the sequential residual block includes a convolutional layer, an activation function (ReLU), element-wise addition, a bidirectional long short-term memory network, and a residual connection. For the convolutional layer, it extracts the local features of the image and generates feature maps by scanning the input image with a convolutional kernel. For the activation function, it introduces non-linearity to enhance the expression ability of the network and helps to learn complex features. For element-wise addition, it adds the prior obtained above element by element, adding the comprehensive features of the image obtained in the previous step to the structural prior extracted from the structural prior branch, that is, adding the comprehensive features of the image to the spatial dimension information to retain important information. For the bidirectional long short-term memory network (BLSTM), it captures the context information through the bidirectional long short-term memory network (BLSTM) to enhance the network's understanding of time and space sequences. For the residual connection, it directly adds the input to the output through a skip connection to alleviate the problem of gradient disappearance and accelerate training.

[0094] It should be noted that the SRB deeply excavates the image features, and the sub-pixel convolution directly maps the multi-scale features to the high-resolution space through an efficient upsampling operation. To further capture the sequence information, a bidirectional long short-term memory network (BLSTM) is also introduced, enabling the model to more accurately model the image structure information. At the same time, the cascade design avoids information loss, improves the reconstruction accuracy, reduces artifacts, and achieves a more refined image reconstruction effect.

[0095] S2. Optimize the initial image super-resolution framework through the structural similarity loss function and the connectionist temporal classification loss function, and combine the backpropagation mechanism to obtain the final image super-resolution framework.

[0096] Specifically, optimizing the initial image super-resolution framework through the structural similarity loss function and the connectionist temporal classification loss function, and combining the backpropagation mechanism to obtain the final image super-resolution framework includes the following steps:

[0097] S21. Generate the total loss function through the structural similarity loss function and the connectionist temporal classification loss function.

[0098] It should be noted that the loss function adopts a joint supervision mechanism of the structural similarity loss function (SSIM) and the connectionist temporal classification loss function (CTC) to provide low-level and high-level supervision respectively; as low-level supervision, SSIM not only focuses on the differences between pixels, but also considers features such as the structure, brightness, and contrast of the image, thus ensuring that the image is more in line with human visual perception; SSIM can optimize the overall quality of the generated image; while CTC ensures that the network can focus on the text area and restore more accurate text details.

[0099] It should be noted that CTC, as high-level supervision, focuses on the text recognition task and solves the problem of the mismatch between the lengths of the input sequence and the target sequence; through the automatic alignment mechanism, the CTC loss effectively improves the accuracy of text recognition.

[0100] It should be noted that a multi-level supervision mechanism combining the structural similarity loss function and the connectionist temporal classification loss function is adopted to achieve the joint optimization of low-level features and high-level semantics, thereby significantly improving the accuracy and robustness of scene text recognition, and combined with the backpropagation mechanism, it can effectively guide the image super-resolution reconstruction process.

[0101] It should be noted that the joint application of SSIM and CTC enhances the semantic accuracy of text recognition while providing guarantee for image structure details, provides comprehensive guidance for scene text image super-resolution, and this combination method significantly improves the overall performance of the final image super-resolution framework in restoring low-resolution images.

[0102] Specifically, the expression of the structural similarity loss function is:

[0103]

[0104] In the formula, SSIM(x,y) represents the structural similarity loss function; μ x represents the mean value of image x; μ y represents the mean value of image y; represents the variance of image x; represents the variance of image y; σ xy represents the covariance of images x and y; C 1 represents the calculation constant one; C 2 represents the calculation constant two.

[0105] Specifically, the expression of the connectionist temporal classification loss function is:

[0106]

[0107] In the formula, L CTC represents the negative logarithm average of probability; l irepresents the corresponding correct label of the input image; y i represents the given arrival path; i represents the i-th input image; N represents the batch size of the input images; p(l i |y i ) represents the probability of the target sequence.

[0108] It should be noted that y i represents the given arrival path, that is, the dynamic programming algorithm lists the given arrival paths, specifically the y i th input image; i represents the i-th input image, that is, the current is the i-th input image; N represents the batch size of the input images, that is, the batch size of the current batch of input images; p(l i |y i ) represents the probability of the target sequence, that is, the sum of the path probabilities of all mappings to the target sequence.

[0109] It should be noted that the derivation steps of p(l i |y i ) are as follows: calculate the output sequence of the initial image super-resolution framework, calculate the path probability according to the output sequence, and then calculate the probability of the target label according to the path probability to obtain p(l i |y i ).

[0110] The calculation formula for the output sequence of the initial image super-resolution framework is:

[0111]

[0112] In the formula, z represents the output sequence of the initial image super-resolution framework; z t represents the probability of each character category at time step t; represents the probability of each character category at time step t; C represents the number of character categories.

[0113] The calculation formula for the path probability is:

[0114]

[0115] In the formula, p(π|y) represents the path probability; represents the probability that the T-th frame is predicted as the character with the class index π t ; T represents the length of the sequence output; y represents the given arrival path of the dynamic programming algorithm, that is, the y-th input image; π represents the set of paths for all target labels l obtained through mapping; π t represents the character category of the path π at the t-th time step.

[0116] The calculation formula for the probability of the target label is:

[0117]

[0118] Where p(l|y) represents the probability of the target label; B represents a many-to-one mapping process for the input value; l represents the target label; B(π)=l represents the path converted to label l through B transformation.

[0119] Specifically, the expression of the total loss function is:

[0120]

[0121] Where L represents the total loss function; L rec represents the reconstruction loss; λ represents the hyperparameter that weighs the structural similarity loss and the connectionist temporal classification loss; SSIM(y i ,G(x i )) represents the structural similarity loss between the first input image and the reconstructed scene text image; p(l i |G(x i )) represents the connectionist temporal classification loss between the input image and the corresponding reconstructed scene text image; x i Indicates the xth i An input image; G(x i ) represents the reconstructed scene text image.

[0122] It should be noted that L represents the total loss function, which is rec Loss and L CTC The loss is composed of SSIM(y i ,G(x i )) represents the structural similarity loss between the yi-th input image and the reconstructed scene text image, that is, the current y-th i The structural similarity loss between the input image and the reconstructed scene text image; p(l i |G(x i )) represents the connectionist temporal classification loss between the input image and the corresponding reconstructed scene text image, that is, the connectionist temporal classification loss between the current input image and the corresponding reconstructed scene text image.

[0123] It should be noted that G(x i ) represents the reconstructed scene text image, that is, the generated high-resolution image.

[0124] S22, based on the initial image super-resolution framework and combined with the total loss function, obtaining a total loss value;

[0125] S23. According to the total loss value, the initial image super-resolution framework is optimized using the back-propagation mechanism to obtain the final image super-resolution framework.

[0126] S3. Input the pre-acquired scene text image into the final image super-resolution framework to obtain the reconstructed scene text image.

[0127] It should be noted that inputting the pre-acquired scene text image into the final image super-resolution framework to obtain the reconstructed scene text image specifically means: inputting the acquired low-resolution scene text image into the final image super-resolution framework to complete the reconstruction of the high-resolution image and improve the clarity and readability of the text area.

[0128] It should be noted that after the scene text image is processed by bimodal prior refinement, it is input into global feature fusion processing to extract global features of multimodal information. Subsequently, the fused information is added element by element to the spatial dimension information obtained from the structural prior to achieve comprehensive modeling and optimization of features.

[0129] For example, the effectiveness of the method provided by this solution is verified through simulation experiments. The specific content includes the real scene text super-resolution dataset (TextZoom). The TextZoom benchmark dataset is from two real-world scene text super-resolution datasets, which includes 21,740 pairs of low-resolution (LR) and high-resolution (HR) text images.

[0130] The final image super-resolution framework can be specifically constructed by software. For example, the final image super-resolution framework is trained and tested on the Pytorch 1.8 deep learning library of NVIDIA RTX 4060ti GPU; and the Adam optimizer is used to train a model with a batch size of 64, the learning rate is set to 10-3, and it is trained 500 times.

[0131] Table 1 Experimental results of TextZoom

[0132]

[0133]

[0134] As can be seen from Table 1, the final image super-resolution framework (DMP) is the most effective for simple, medium, and difficult tasks in the case of 50% high-resolution labels; compared with existing models, the recognition accuracy is also the best; in order to more intuitively show the role of the final image super-resolution framework in the scene text super-resolution task, a visualization study was conducted (as Figure 4 shown); the red letters in the figure represent the misrecognized text, the black letters represent the correct ones, and the last row represents the high-resolution image (HR).

[0135] The visual super-resolution effects of DMP, BICUBIC, TSRN, and TATT methods on the TextZoom test dataset were compared under the condition that 50% of the HR labels were available; the experimental results showed that DMP could effectively handle the intractable situation of insufficient or even incorrect information in the super-resolution model, significantly improve the visual quality of the super-resolution (SR) images, and enhance the readability of characters.

[0136] For 50% of the high-resolution labels, λ was selected through ablation experiments. Different λ values in the set {0, 0.01, 0.1, 1} were used in the experiments, and 50% of the high-resolution labels were used to optimize the performance of the final image super-resolution framework. The experimental results are summarized in Table 2.

[0137] Table 2 Ablation experiments on λ selection

[0138]

[0139] From Table 2, it can be observed that when λ was set to 0.1, the recognition accuracy was the highest; this configuration provided the best balance between the recognition accuracy and the robustness of the final image super-resolution framework. Therefore, λ = 0.1 was selected as the parameter for subsequent experiments to ensure continuous improvement in performance.

[0140] In summary, with the above technical solutions of the present invention, through the dual-modal prior refinement process, the capture of the global structure and semantic information of the scene text image is enhanced, thereby improving the reconstruction quality of the super-resolution image; at the same time, through the global feature fusion process, the overall quality of the generated high-resolution image is improved, especially in the restoration effect of complex scenes and text regions; and combined with the multi-level supervision strategy of the structural similarity loss function and the connectionist temporal classification loss function, the image structure and semantic information can be fully captured to help improve the text recognition performance; through the dual-modal prior refinement process, global feature fusion process, and reconstruction process, it is possible to effectively capture the multi-scale global and detailed features of the image under the condition of only using 50% of the high-resolution labels, significantly improving the average recognition accuracy of the scene text and optimizing the reconstruction quality of the image; through the text prior and structure prior branch processes, the text semantics and image structure features are combined, so as to be able to capture multi-scale global features and detailed information at the same time, effectively improving the reconstruction performance of the text region; and in the multi-level supervision mechanism of the structural similarity loss function and the connectionist temporal classification loss function, the joint optimization of low-level features and high-level semantics is realized, thereby significantly improving the accuracy and robustness of scene text recognition.

[0141] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A weakly supervised super-resolution method based on bimodal priors, characterized in that: The method comprises the following steps: S1, obtaining a scene text image, performing feature extraction on the scene text image, and constructing an initial image super-resolution framework based on the extracted image features; S2, optimizing the initial image super-resolution framework through the structural similarity loss function and the connectionist temporal classification loss function, combined with the back-propagation mechanism, to obtain the final image super-resolution framework; S3. Input the pre-acquired scene text image into the final image super-resolution framework to obtain a reconstructed scene text image.

2. The weakly supervised super-resolution method based on bimodal prior according to claim 1, characterized in that: The step of obtaining a scene text image, extracting features from the scene text image, and constructing an initial image super-resolution framework based on the extracted image features comprises the following steps: S11, obtaining a scene text image, and extracting text semantic information and spatial dimension information through bimodal prior refinement processing; S12, performing global feature fusion processing on the extracted text semantic information and spatial dimension information through the self-attention mechanism of the visual transformer to obtain fused information, and adding the fused information and the spatial dimension information element by element to obtain the comprehensive features of the image; S13. Based on the comprehensive features of the image, several sequential residual blocks are cascaded and combined with sub-pixel convolution to reconstruct the image and obtain the initial image super-resolution framework.

3. The weakly supervised super-resolution method based on bimodal prior according to claim 2, characterized in that: The method of extracting text semantic information and spatial dimension information through bimodal prior refinement processing includes the following steps: Extracting text semantic information in the scene text image through bimodal prior refined text prior branch processing; the text prior branch processing includes small sample text recognition processing and convolution processing; The spatial dimension information in the scene text image is extracted through the structural prior branch processing of bimodal prior refinement; the structural prior branch processing includes centralized feature pyramid processing and squeeze-excitation network processing.

4. The weakly supervised super-resolution method based on bimodal prior according to claim 3, characterized in that: The expression of the text prior branch processing is: P text =g text (I LR ); Where P text Represents text semantic information; g text Represents text prior; I LR Represents scene text image; The expression of the structural priori branch processing is: P structure =g structure (I LR ); Where P structure Represents spatial dimension information; g structure Represents structural prior.

5. The weakly supervised super-resolution method based on bimodal prior according to claim 2, characterized in that: The extracted text semantic information and spatial dimension information are subjected to global feature fusion processing through the self-attention mechanism of the visual transformer to obtain fused information, and the fused information is added element by element to the spatial dimension information to obtain the comprehensive features of the image, including the following steps: S121, based on the self-attention mechanism of the visual transformer, the scene text image is divided into several image blocks, and a position code is added to each image block, and an attention score of each image block is calculated; S122, generating a self-attention mapping function according to the attention scores of all image blocks, and combining the extracted text semantic information and spatial dimension information to obtain fusion information; S123. Add the fusion information and the spatial dimension information element by element to obtain the comprehensive features of the image.

6. The weakly supervised super-resolution method based on bimodal prior according to claim 5, characterized in that: The expression of the self-attention mapping function is: P fused =φ(P text +P structure ); Where P fused represents fusion information; φ represents the self-attention mapping function; P text Represents text semantic information; P structure Represents spatial dimension information.

7. The weakly supervised super-resolution method based on bimodal prior according to claim 1, characterized in that: The method of optimizing the initial image super-resolution framework by using the structural similarity loss function and the connectionist temporal classification loss function in combination with the back-propagation mechanism to obtain the final image super-resolution framework includes the following steps: S21, generate the total loss function through the structural similarity loss function and the connectionist temporal classification loss function; S22, based on the initial image super-resolution framework and combined with the total loss function, obtaining a total loss value; S23. According to the total loss value, the initial image super-resolution framework is optimized using the back-propagation mechanism to obtain the final image super-resolution framework.

8. The weakly supervised super-resolution method based on bimodal prior according to claim 7, characterized in that: The expression of the structural similarity loss function is: Where SSIM(x,y) represents the structural similarity loss function; μ x represents the mean of image x; μ y Represents the mean of image y; represents the variance of image x; represents the variance of image y; σ xy Represents the covariance of images x and y; C1 represents the calculation constant one; C2 represents the calculation constant two.

9. The weakly supervised super-resolution method based on bimodal prior according to claim 8, characterized in that: The expression of the connectionist temporal classification loss function is: Where, L CTC represents the negative logarithmic mean of probability; l i Represents the corresponding correct label of the input image; y i represents a given arrival path; i represents the i-th input image; N represents the batch size of the input image; p(l i |y i ) represents the probability of the target sequence.

10. The weakly supervised super-resolution method based on bimodal prior according to claim 9, characterized in that: The expression of the total loss function is: In the formula, L represents the total loss function; L rec represents reconstruction losses; λ represents a hyperparameter that weighs the structural similarity loss and the connectionist temporal classification loss; SSIM(y i ,G(x i )) represents the structural similarity loss between the first input image and the reconstructed scene text image; p(l i |G(x i )) represents the connectionist temporal classification loss between the input image and the corresponding reconstructed scene text image; x i Indicates the xth i An input image; G(x i ) represents the reconstructed scene text image.

Citation Information

Cited By

  • Scene text super-resolution method and system based on text prior and stationary wavelet domain transformation

    CN120807290A

  • A Scene Text Super-Resolution Method and System Based on Text Prior and Stationary Wavelet Domain Transform

    CN120807290B

  • Mangrove forest optical remote sensing image super-resolution reconstruction method, device, equipment, medium and product

    CN122115218A