Data-enhanced text image super-resolution reconstruction method and device

By combining hybrid and generative data augmentation methods, and utilizing CutBlur technology and generative adversarial networks to generate high-quality text image samples, the problem of insufficient training dataset size and diversity is solved, thereby improving the generalization ability and reconstruction quality of the text image super-resolution reconstruction model.

CN120807292AActive Publication Date: 2025-10-17AOKAI SPACE IMAGING TECHNOLOGY (NINGBO) CO LTD

Patent Information

Application Number
CN202510942930.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-17
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

Existing text image super-resolution reconstruction methods suffer from insufficient training dataset size and diversity, resulting in inadequate model generalization ability and difficulty in effectively recovering high-quality, clearly identifiable text content.

Method used

By combining hybrid data augmentation and generative data augmentation methods, CutBlur technology is used to achieve smooth blending of local regions, and generative adversarial networks are used to generate high-quality text image samples to expand the distribution space of training data.

Benefits of technology

It significantly improves the model's generalization ability and reconstruction quality, and the generated text image super-resolution effect is better than using hybrid or generative data augmentation methods alone, improving image clarity and recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807292A_ABST
    Figure CN120807292A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data-enhanced text image super-resolution reconstruction method and device. Through the weight mask mixing technology and data generation based on the generative adversarial network, the distribution space of training data is obviously expanded, the generalization ability and reconstruction quality of the model are improved, and the method has important application prospects in the field of text image super-resolution reconstruction. Smooth mixing of local areas is achieved through the CutBlur technology, then a high-quality text image sample is generated through the generative adversarial network, and finally the generalization ability and reconstruction performance of the model are remarkably improved through the synergistic effect of the two methods.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer vision and image technology, in particular to a data enhancement text image super-resolution reconstruction method and device. BACKGROUND

[0002] Text Image Super-Resolution Reconstruction (TISRR) is an important branch of computer vision, aiming to recover high-quality, clear and distinguishable text content from low-resolution (LR) text images. With the development of deep learning technology, super-resolution reconstruction methods based on deep neural networks have made significant progress. However, such methods usually require a large number of low-resolution-high-resolution (HR) image pairs as training data to learn complex non-linear mapping relationships.

[0003] In the task of text image super-resolution reconstruction, it is challenging to collect large-scale high-quality text image datasets due to the difficulty of collection and high annotation cost. Publicly available text image datasets are often limited in size and single in scene, making it difficult to cover diverse text degradation scenarios in the real world, such as blur, noise, distortion, etc. The limitations of such datasets lead to overfitting problems in the trained super-resolution reconstruction model, and the generalization ability is insufficient when facing unseen text images.

[0004] Data augmentation, as an effective strategy to alleviate the problem of data scarcity, has been widely used in multiple tasks of computer vision. Traditional data augmentation methods include geometric transformation, color adjustment, noise injection, etc., but these methods are too simple to effectively increase the diversity and complexity of the dataset [4]. In recent years, hybrid data augmentation methods such as CutMix generate new training samples by combining local regions of different images. However, such methods often produce unnatural pixel value mutations at the stitching boundary, affecting the model's learning of continuous textures.

[0005] On the other hand, generative data augmentation based on generative adversarial networks (GAN) can create new training samples, but its generation quality and stability still need to be improved. Especially in the field of text images, the generated images need to maintain the structural integrity and semantic consistency of the text, which puts higher requirements on the generation model.

[0006] Based on this, the present application proposes a hybrid and generative data augmentation method, which is specifically optimized for the characteristics of the text image super-resolution reconstruction task. SUMMARY

[0007] In view of the above problems, a text image super-resolution reconstruction technology is proposed to overcome the above problems or at least partially solve the above problems. In view of the problem of insufficient scale and diversity of training data set in the text image super-resolution reconstruction task, the application proposes a method and device combining mixed data enhancement and generative data enhancement. The method significantly expands the distribution space of the training data by using the weight mask mixing technology and the data generation based on the generative adversarial network, improves the generalization ability and reconstruction quality of the model, and has important application prospect in the field of text image super-resolution reconstruction.

[0008] A data enhanced text image super-resolution reconstruction method, the method comprising:

[0009] obtaining a training data set as a training sample from an original data set , using an improved CutBlur algorithm to process the sample, and obtaining a mixed enhanced data set ;

[0010] using the mixed enhanced data set to pre-train an ESRGAN generator, alternately training the generator and the discriminator, and optimizing through a loss function to obtain a trained generator;

[0011] using the trained generator to perform super-resolution reconstruction on an original low-resolution image to obtain new training image data , based on the original data set , the mixed enhanced data set , and the new training image data , combining image data based on a hyperparameter alpha using a hierarchical sampling strategy to obtain enhanced image data .

[0012] Further, obtaining a training data set as a training sample from an original data set , using an improved CutBlur algorithm to process the sample, and obtaining a mixed enhanced data set Before that, the original data set is divided into a training set, a validation set and a test set according to a ratio of 8:1:1.

[0013] Further, obtaining a training data set as a training sample from an original data set , using an improved CutBlur algorithm to process the sample, and obtaining a mixed enhanced data set ; comprising:

[0014] randomly selecting a rectangular mask area with a mask area ratio conforming to a uniform distribution ;

[0015] Computing adaptive smoothing parameters wherein an enhanced image is created using a smoothing weight mask; wherein is a base smoothing parameter, is a mask area, is a total image area;

[0016] The enhanced images are quality evaluated using Learned Perceptual Image Patch Similarity (LPIPS) and low quality samples with LPIPS score lower than a threshold of 0.3 are removed to obtain a mixed enhanced dataset wherein the LPIPS is calculated according to the following formula:

[0017] .

[0018] Further, the randomly selected mask area ratio conforms to a uniform distribution The rectangular mask area includes:

[0019] Two groups of low resolution-high resolution image pairs are randomly selected from the training dataset , denoted as , and the corresponding high definition image is , wherein the reconstruction magnification is denoted as s;

[0020] A rectangular area in the image is randomly selected as a base mask, wherein a distance change function is calculated , and the calculation formula is:

[0021] .

[0022] Further, the adaptive smoothing parameters are calculated wherein an enhanced image is created using a smoothing weight mask; including:

[0023] A bidirectional sigmoid weight mask is constructed to ensure that the weight is close to 1 in the central area inside the mask, and the weight is close to 0 in the area far away from the boundary outside the mask;

[0024] Near the mask boundary, the weight presents a smooth S-shaped transition, and the transition bandwidth can be controlled by the parameter Using this smooth weight mask, an enhanced generated image is created, wherein the weight mask calculation formula is:

[0025]

[0026] The enhanced image is generated using a smoothing weight mask, and the calculation formula is:

[0027]

[0028]

[0029] wherein, represents element-level multiplication, represents a downsampling operation, represents an upsampling operation, and respectively represent the pixel point values of the low-resolution image and the high-resolution image.

[0030] Further, the ESRGAN generator is pre-trained using the mixed enhanced dataset, the generator and the discriminator are alternately trained, and the trained generator is obtained through optimization by the loss function; comprising:

[0031] The learning rate is set to 0.0002, the training rounds are 50 rounds, the ESRGAN generator is pre-trained using the mixed enhanced dataset, and the generated image is obtained;

[0032] The parameters are set by the loss function, including the feature modulation coefficient , The generator learning rate is 0.001 and the discriminator learning rate is 0.004 using the Adam optimizer setting, and the total training total rounds are 200 rounds; according to the set parameters, the generated image is used as input data, and the generator is trained once and the discriminator is trained once in the alternating training mode, and the trained generator is obtained through the adversarial training.

[0033] Further, the learning rate is set to 0.0002, the training rounds are 50 rounds, the ESRGAN generator is pre-trained using the mixed enhanced dataset, and the generated image is obtained; comprising a feature extraction front end and an information processing main function module, and the following processing is performed:

[0034] In the front-end processing stage, the initial feature acquisition and feature dimension expansion are realized through single-layer convolution operation, specifically, the shallow feature is obtained through the first convolution layer, which contains visual features, as shown in the following formula:

[0035]

[0036] wherein, refers to the initial convolution operation unit for processing the low-resolution input image, and the extracted feature is then transmitted to the network main part for deep feature mining;

[0037] The core backbone network adopts a cascaded residual feature aggregation unit to construct the backbone network to enhance the feature extraction capability and information transmission efficiency, and includes N series of residual calculation units and a feature modulation coefficient group; the operation formula is:

[0038]

[0039] Wherein, represents the mapping function of the i-th residual calculation unit, represents the information flow input to the i-th feature fusion module, is the output feature generated by the module;

[0040] The extracted depth features are reconstructed by the reconstruction part First, a convolutional layer is passed, and then the resolution is expanded to generate an image, as shown in the following formula:

[0041]

[0042] Wherein, represents the super-resolution reconstructed image, represents the up-sampling operation, respectively, the convolutional layer after the main part and the convolutional layer after the up-sampling.

[0043] Further, the parameter setting by the loss function includes:

[0044] The loss function of the generator Combining the adversarial loss , the content loss and the perception loss , the calculation formula is:

[0045]

[0046] Wherein, the adversarial loss , the content loss ensures the consistency of the generated image and the real high-resolution image at the pixel level, and the calculation formula is , the perception loss is based on the pre-trained VGG network, and the calculation formula is , represents the feature extraction function of the pre-trained VGG network.

[0047] Further, the trained generator is used to perform super-resolution reconstruction on the original low-resolution image to obtain new training image data According to the original data set , the mixed enhanced data set And new training image data , based on the hyperparameter α, a layered sampling strategy is used to combine image data to obtain enhanced image data that can be dynamically adjusted with the training results. ,include:

[0048] Use the trained generator to reconstruct the original image through the reconstruction formula Super-resolution reconstruction of low-resolution images in the image is performed to generate new training image data :

[0049]

[0050] Based on the hyperparameter α that can be dynamically adjusted according to the training results, a stratified sampling strategy is used to sample new training image data. , original image And the mixed enhanced image Three types of data are combined to obtain enhanced image data. , where the formula for layered sampling combined image data is:

[0051]

[0052] in, is a hyperparameter.

[0053] A data-enhanced text image super-resolution reconstruction device, the device comprising:

[0054] Hybrid data preprocessing module for raw data sets Get the training data set as training samples, use the improved CutBlur algorithm to process the samples, and obtain the hybrid enhanced data set ;

[0055] The generative model training module is used to pre-train the ESRGAN generator using a mixed augmented dataset, alternately train the generator and discriminator, and optimize the loss function to obtain a trained generator;

[0056] The dataset fusion module is used to use the trained generator to perform super-resolution reconstruction on the original low-resolution image to obtain new training image data. , based on the original dataset , mixed augmented dataset And new training image data , based on the hyperparameter α, the image data is combined using a layered sampling strategy to obtain enhanced image data .

[0057] An electronic device comprises a processor, a memory, and a computer program stored on the memory and capable of running on the processor, the computer program being implemented when executed by the processor to realize a data-enhanced text image super-resolution reconstruction method.

[0058] A computer-readable storage medium stores a computer program, the computer program being implemented when executed by a processor to realize a data-enhanced text image super-resolution reconstruction method.

[0059] A computer program product comprises a computer program, the computer program being implemented when executed by a processor to realize a data-enhanced text image super-resolution reconstruction method.

[0060] Embodiments of the present application have the following advantages: by using the weight mask mixing technology and the data generation based on the generative adversarial network, the distribution space of the training data is significantly expanded, the generalization ability and the reconstruction quality of the model are improved, and the model has important application prospects in the field of text image super-resolution reconstruction. By using the CutBlur technology, the local area is mixed smoothly, then high-quality text image samples are generated by using the generative adversarial network, and finally the generalization ability and the reconstruction performance of the model are significantly improved through the synergistic effect of the two methods. BRIEF DESCRIPTION OF DRAWINGS

[0061] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the description of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0062] Figure 1 is a step flow chart of a data-enhanced text image super-resolution reconstruction method provided by some embodiments of the present application;

[0063] Figure 2 is a combined mixing and generative text image data enhancement flow chart of a data-enhanced text image super-resolution reconstruction method provided by some embodiments of the present application;

[0064] Figure 3 is a generator overall network structure schematic diagram of a data-enhanced text image super-resolution reconstruction method and device provided by some embodiments of the present application. DETAILED DESCRIPTION

[0065] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below in combination with the drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0066] The core idea of the present application is that based on the characteristics of the text image super-resolution reconstruction task, a hybrid and generative data enhancement method is proposed. The method first realizes the smooth mixing of the local area through the CutBlur technology, then generates high-quality text image samples by using the generative adversarial network, and finally significantly improves the generalization ability and reconstruction performance of the model through the synergistic effect of the two methods.

[0067] Referring to Figure 1 and Figure 2 , a method for data-enhanced text image super-resolution reconstruction is shown, which can specifically include the following steps:

[0068] S1, obtaining a training data set as a training sample from an original data set , processing the sample by using an improved CutBlur algorithm to obtain a hybrid enhanced data set ;

[0069] S2, pre-training the ESRGAN generator using the hybrid enhanced data set, alternately training the generator and the discriminator, and optimizing by using a loss function to obtain a trained generator;

[0070] S3, performing super-resolution reconstruction on the original low-resolution image by using the trained generator to obtain new training image data , combining image data based on a hyperparameter alpha by using a hierarchical sampling strategy according to the original data set , the hybrid enhanced data set and the new training image data , to obtain enhanced image data .

[0071] First, the CutBlur technology is used to realize the smooth mixing of the local area, and then high-quality text image samples are generated by using the generative adversarial network, and finally the generalization ability and reconstruction performance of the model are significantly improved through the synergistic effect of the two methods.

[0072] The improved hybrid data enhancement and generative data enhancement are organically combined in the present application to form a synergistic enhancement strategy. The strategy is based on the following core theory:

[0073] Complementarity principle: hybrid data augmentation mainly increases data diversity by recombining local regions of existing data, while generative data augmentation can create new data samples. Both have natural complementarity in data distribution space coverage. Hybrid methods maintain the authenticity and local feature consistency of original data, while generative methods expand the boundaries of data distribution and introduce new feature combination patterns.

[0074] Progressive learning mechanism: first expand the initial training set through hybrid data augmentation to provide a richer training basis for the generative model, and then use the trained generative model to further expand the data distribution, forming a progressive data augmentation method.

[0075] In some embodiments of the application, the step S1 is performed on the original data set The training data set is obtained as a training sample, and the improved CutBlur algorithm is used to process the sample to obtain a hybrid augmented data set In some embodiments of the application, the step S1 is performed on the original data set , and the original data set is divided into a training set, a validation set and a test set according to a ratio of 8:1:1.

[0076] The step S1 is performed on the original data set The training data set is obtained as a training sample, and the improved CutBlur algorithm is used to process the sample to obtain a hybrid augmented data set , comprising:

[0077] Randomly selecting a rectangular mask area with a mask area ratio conforming to a uniform distribution ;

[0078] Calculating an adaptive smoothing parameter , wherein , a smoothing weight mask is used to create an augmented image; wherein is a basic smoothing parameter, is the area of the mask region, is the total area of the image. In this way, the smoothing strength can be adaptively adjusted according to the size of the mask region.

[0079] The quality of the augmented image is evaluated using a learned perceptual image patch similarity (LPIPS) index, and low-quality samples with an LPIPS score below a threshold of 0.3 are removed to obtain a hybrid augmented data set , wherein the calculation formula of the LPIPS is: .

[0080] The above step S2, pre-training the ESRGAN generator using the mixed augmented dataset, alternately training the generator and the discriminator, and optimizing them through the loss function to obtain a trained generator, includes: setting the learning rate to 0.0002, training 50 rounds, pre-training the ESRGAN generator using the mixed augmented dataset, and obtaining generated images;

[0081] Parameter setting through loss function, including feature modulation coefficient , The Adam optimizer is used to set the generator learning rate to 0.001, the discriminator learning rate to 0.004, and the total number of training rounds to 200 rounds. According to the set parameters, the generated image is used as input data, and adversarial training is performed in an alternating training method of training the generator once and the discriminator once to obtain a trained generator.

[0082] In the above step S3, the trained generator is used to perform super-resolution reconstruction on the original low-resolution image to obtain new training image data. , based on the original dataset , mixed augmented dataset And new training image data , based on the hyperparameter α, a layered sampling strategy is used to combine image data to obtain enhanced image data that can be dynamically adjusted with the training results. ,include:

[0083] Use the trained generator to reconstruct the original image through the reconstruction formula Super-resolution reconstruction of low-resolution images in the image is performed to generate new training image data : ;

[0084] Based on the hyperparameter α that can be dynamically adjusted according to the training results, a stratified sampling strategy is used to sample new training image data. , original image And the mixed enhanced image Three types of data are combined to obtain enhanced image data. , where the formula for layered sampling combined image data is:

[0085] ;in, is a hyperparameter.

[0086] As an example, a data-enhanced text image super-resolution reconstruction method adopts a three-stage training strategy by combining hybrid and generative data enhancement methods. The relevant algorithm pseudo code is shown in Table 1:

[0087] 1. Phase one: mixed data preprocessing; first, data set segmentation: the original data set is divided into training set, validation set and test set in the ratio of 8:1:1. Then, the CutBlur enhancement is performed, that is, the improved CutBlur method is applied to each image pair in the training set to generate enhanced samples. The specific process is as follows: first, a rectangular mask area is randomly selected, and the mask area ratio conforms to a uniform distribution ; then, the adaptive smoothing parameter is calculated, wherein , and finally, the enhanced image pair is created using the smoothing weight mask; in order to screen the enhanced images, the learned perceptual image patch similarity (LPIPS) index is used in the present application to evaluate the quality of the enhanced images, and low-quality samples with an LPIPS score lower than a threshold value of 0.3 are removed, and the calculation formula of LPIPS is as follows:

[0088]

[0089] 2. Phase two: training of the generative model; first, the pre-training phase, the mixed enhanced data set is used to pre-train the ESRGAN generator, the training round number is 50 rounds, and the learning rate is set to 0.0002. Then, the adversarial training phase, that is, the generator and the discriminator are alternately trained, the generator is trained once, and the discriminator is trained once. The loss function weight setting is as follows: , The Adam optimizer is adopted, the generator learning rate is 0.004, and the total training round number is 200 rounds.

[0090] 3. Phase three: data set fusion; the trained generator is used to perform super-resolution reconstruction on the original low-resolution image to generate new training data pairs: ; then, a hierarchical sampling strategy is adopted to combine the original image , the mixed enhanced image , and the image generated by using the mixed enhanced image for training three types of data: ; wherein is a hyperparameter, which can be manually initially set and dynamically adjusted according to the training result.

[0091] In some embodiments of the present application, the randomly selected mask area ratio conforms to a uniform distribution of the rectangular mask area, including: randomly selecting two groups of low-resolution-high-resolution image pairs from the training data set, denoted as its corresponding high-definition image is wherein the reconstruction magnification is denoted by s;

[0092] randomly select a rectangular region in the image as the base mask, wherein the distance change function is calculated , the calculation formula is: .

[0093] the adaptive smoothing parameter is calculated wherein an enhanced image is created using the smoothing weight mask; comprising:

[0094] a bidirectional sigmoid weight mask is constructed to ensure that the weight is close to 1 in the central region inside the mask, while the weight is close to 0 in the region far away from the boundary outside the mask;

[0095] near the mask boundary, the weight presents a smooth S-shaped transition, and the transition bandwidth can be controlled by the parameter using this smooth weight mask, an enhanced image (new training sample) is created, wherein the weight mask calculation formula is: ; using the smooth weight mask to generate an enhanced image (new training sample), the calculation formula is:

[0096] ;

[0097] ; wherein, element-level multiplication is denoted by downsampling operation is denoted by up-sampling operation is denoted by and respectively represent the pixel point values of the low-resolution image and the high-resolution image.

[0098] The learning rate is set to 0.0002, the number of training rounds is 50 rounds, and the ESRGAN generator is pre-trained using a mixed enhanced data set to obtain a generated image; comprising three functional modules of feature extraction front end, information processing backbone and image reconstruction back end, which are processed as follows:

[0099] In the front-end processing stage, the initial feature acquisition and feature dimension expansion are realized through single-layer convolution operation, specifically, the shallow feature is obtained through the first convolution layer, which contains visual features, as shown in the following formula: ; wherein, denotes the initial convolution operation unit for processing the low-resolution input image, and the extracted feature is then transmitted to the network backbone part for deep feature mining;

[0100] The core backbone network adopts a cascaded residual feature aggregation unit to construct the backbone network to enhance the feature extraction capability and information transmission efficiency, and includes N series of residual calculation units and a feature modulation coefficient group; the operation formula is: ; wherein, represents the mapping function of the i-th residual calculation unit, represents the information flow input to the i-th feature fusion module, is the output feature generated by the module;

[0101] The extracted depth features are reconstructed by the reconstruction part First, a convolutional layer is passed, and then the resolution is expanded to obtain the generated image, as shown in the following formula: ; wherein, represents the super-resolution reconstructed image, represents the up-sampling operation, are the convolutional layer after the main part and the convolutional layer after the up-sampling, respectively.

[0102] The loss function is used to set the parameters, including the loss function of the generator combined with the adversarial loss , the content loss and the perception loss , and the calculation formula is:

[0103] ;

[0104] Among them, the adversarial loss , the content loss ensures the consistency of the generated image and the real high-resolution image at the pixel level, and the calculation formula is , and the perception loss is based on the pre-trained VGG network, and the calculation formula is , represents the feature extraction function of the pre-trained VGG network.

[0105] It should be noted that in the formulas of the present application, some parameter terms appearing in the formulas, unless otherwise specified, are generally general parameters, that is, the same parameter symbols listed in the above formulas or subsequent formulas should be understood as the same parameters, and the same parameters appearing in each formula can be referred to each other, unless the parameter has another parameter specification. The one specified is the one that should be followed.

[0106] Exemplarily, based on the characteristics of the text image super-resolution reconstruction task, a hybrid and generative data enhancement method is proposed. The method first realizes the smooth mixing of the local area through the CutBlur technology, and then generates high-quality text image samples by using the generative adversarial network. Finally, the generalization ability and reconstruction performance of the model are significantly improved through the synergistic effect of the two methods. The overall process is shown in Figure 2 .

[0107] By improving the hybrid data enhancement method of CutBlur, specifically, for the pixel value mutation problem caused by the traditional CutMix method at the image splicing place, the hybrid data enhancement method of the literature CutBlur is adopted, and a weight mask for the TISRR task is designed to realize the natural transition of the local area. The specific steps are as follows:

[0108] First, randomly select two sets of low-resolution-high-resolution image pairs from the training data set , and mark the low-resolution image as , and the corresponding high-resolution image as , wherein the reconstruction magnification is represented by s, and then randomly select a rectangular area in the image as the basic mask. First, calculate the distance change function , and the calculation formula is: ; Then calculate the adaptive smoothing parameter , and the calculation formula is: ; wherein is the basic smoothing parameter, is the area of the mask region, is the total area of the image. In this way, the smoothing strength can be adaptively adjusted according to the size of the mask region.

[0109] Finally, construct a bidirectional sigmoid weight mask . This design ensures that the weight is close to 1 in the center area inside the mask, and the weight is close to 0 in the area far from the boundary outside the mask. Secondly, near the mask boundary, the weight presents a smooth S-shaped transition, and the transition bandwidth can be controlled by the parameter . The weight mask calculation formula is: ;

[0110] Using this smooth weight mask to generate new training samples, the calculation formula is:

[0111] ;

[0112] ; wherein, represents element-level multiplication, denotes a down-sampling operation, denotes an up-sampling operation, and denote pixel values of low-resolution images and high-resolution images, respectively.

[0113] The application discloses a generative data enhancement method based on a generative adversarial network.

[0114] The overall network structure of the generator is shown in Figure 3 The overall structure of the model is derived from ESRGAN. The generation network architecture includes three functional modules: a feature extraction front end, an information processing backbone and an image reconstruction back end. In the front end processing stage, initial feature acquisition and feature dimension expansion are achieved through single-layer convolution operation. The core backbone network is responsible for complex feature representation learning. This part is flexible and variable, and different network structures can be selected according to specific requirements to maximize feature utilization efficiency. In the application, a cascaded residual feature aggregation unit is used to construct the backbone network to enhance feature extraction capability and information transmission efficiency. The loss function of the generator combines the adversarial loss , the content loss and the perceptual loss , and the calculation formula is:

[0115]

[0116] The adversarial loss is used to ensure that the generated super-resolution image has a realistic visual effect. The content loss ensures the consistency of the generated image and the real high-resolution image at the pixel level, and the calculation formula is The perceptual loss is based on a pre-trained VGG network and constrains the quality of the generated image from the perspective of high-level semantic features, and the calculation formula is , denotes a feature extraction function of the pre-trained VGG network.

[0117] Detailed process description of the generator network model: shallow features are obtained through the first convolutional layer and contain visual features, as shown in the following formula:

[0118] wherein, refers to an initial convolution operation unit for processing a low-resolution input image, and the extracted feature is then transmitted to the network backbone part for deep feature mining. The backbone architecture is composed of N cascaded residual calculation units and feature modulation coefficients composition,

[0119] The operation formula is:

[0120] wherein, represents a mapping function of the i-th residual calculation unit, represents an information flow input to the i-th feature fusion module, is an output feature generated by the module. The input of the residual unit is derived from and The element-level accumulation result, which effectively reduces the training complexity and improves the stability of the optimization process, is designed by using a global residual connection. Finally, the deep features extracted by the reconstruction part are first passed through a convolution layer, and then the resolution is expanded to generate an image,

[0121] as shown in the following formula:

[0122] wherein, represents a super-resolution reconstructed image, represents an up-sampling operation, respectively, are the convolution layer after the main part and the convolution layer after the up-sampling. The image restoration terminal is composed of three key links: front-end feature integration convolution, center resolution enhancement module and rear-end detail optimization convolution. Among them, the resolution enhancement component, as the core unit of the reconstruction process, is responsible for converting the low-dimensional feature space to the target output scale, realizing the accurate enlargement and detail recovery of visual content.

[0123] In a specific example, the performance test of the present application is based on the TextZoom dataset, which contains 17367 pairs of training images and 4373 pairs of test images. The experiment uses SRCNN and TSRN as the benchmark super-resolution reconstruction network, uses PSNR and SSIM as the image quality evaluation index, and uses the accuracy of ASTER, MORAN and CRNN three text recognizers as the text recognition performance index.

[0124] The higher the PSNR value, the smaller the difference between the two images, indicating higher quality. Given a noisy image and an image of size m x n,

[0125] The specific calculation formula of PSNR and MSE is:

[0126]

[0127] wherein, ​​​​L is the maximum possible value of a pixel in an image, for RGB images, L is generally 255, assuming that there are N pixels in the image, and respectively represent the gray value or color value of the i-th pixel in the original image and the enhanced image.

[0128] The structural similarity index is a measure of comparing image quality by measuring the degree of similarity between two images, rather than simply calculating the error between two images. The SSIM index considers the brightness, contrast and structural information of the image in the calculation process,

[0129] The SSIM calculation formula is as follows: ;

[0130] wherein, and respectively represent the mean value of the pixel value of the original image and the enhanced image, represent the pixel value covariance of the original image and the enhanced image, and the constant and are used for stability. The SSIM index takes a value between -1 and 1.

[0131] The recognition accuracy ACC refers to the proportion of correct prediction results of a given classification model for a given sample set. In recent years, research on text image super-resolution reconstruction usually uses open source models such as CRNN, ASTER and MORAN to identify the restored SR image and calculate the recognition accuracy.

[0132] The ACC calculation formula is as follows: ;

[0133] wherein, is the number of correctly identified samples, is the total number of samples in the test set.

[0134] As shown in Figure 3 , in the generator of the above embodiment, an original image which is relatively blurred and has a small resolution is taken as input data of the shallow feature convolutional layer, and is sequentially processed through the shallow feature, the main part and the reconstruction part (that is, the above generation network architecture includes three functional modules: feature extraction front end, information processing main body and image reconstruction back end). The processed image is output by the convolutional layer of the reconstruction part. By comparing the input image and the output image, it can be seen that the processed image is superior to the original image with small resolution in terms of clarity and resolution.

[0135] The experimental results are shown in Table 2, and the improved hybrid data enhancement can improve the PSNR of SRCNN by 0.07dB, and the text recognition accuracy by 0.9%. The method combining hybrid and generative data enhancement achieves a 0.09dB PSNR improvement and a 0.4% recognition accuracy improvement on SRCNN. In addition, on the more complex TSRN network, the combination of data enhancement methods improves the PSNR by 0.06dB and the recognition accuracy by 0.8%. In the ablation experiment, the experimental results verify the complementarity of hybrid and generative enhancement, and the combination of the two methods can produce a synergistic effect. In summary, the data enhancement method proposed in the application can effectively improve the performance of the text image super-resolution reconstruction model.

[0136] Table 1 Data enhancement algorithm pseudocode

[0137] Algorithm: Data augmentation algorithm Input: LR image #timg#, HR image #timg#, GAN generator #timg#, discriminator #timg# Output: GAN generated image #timg# Hybrid data augmentation: 1: Randomly select image region to create mask #timg# 2: Create edge-smoothed weight mask #timg# 3: Generate hybrid LR image: #timg# 4: Construct hybrid dataset #timg# Generative data augmentation: 5: for #timg# do 6: Randomly select LR, HR image batch #timg# Generator forward propagation: 7: Generate SR image: #timg# Discriminator forward propagation: 8: #timg# #timg# 9: #timg# 10: #timg# #timg# 11: Update generator: #timg# Discriminator: #timg#

[0138] Table 2 Experimental results of generative data enhancement

[0139] Method PSNR↑ SSIM↑ ACC↑ SRCNN 20.78 0.7227 47.2 SRCNN* 20.84 0.7232 47.5 SRCNN** 20.87 0.7235 47.6 TSRN 21.42 0.7690 54.8 TSRN* 21.45 0.7694 55.2 TSRN** 21.48 0.7696 55.6

[0140] It should be noted that, for the method embodiment, in order to simply describe, it is expressed as a series of action combinations, but those skilled in the art should know that the embodiments of the application are not limited by the order of the described actions, because according to the embodiments of the application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily necessary for the embodiments of the application.

[0141] Some embodiments of the application also provide a data enhanced text image super-resolution reconstruction device for implementing the above-mentioned data enhanced text image super-resolution reconstruction method, which can specifically include the following modules:

[0142] The hybrid data preprocessing module is used to obtain a training data set from an original data set as a training sample, and process the sample using an improved CutBlur algorithm to obtain a hybrid enhanced data set.

[0143] The generative model training module is used to pre-train the ESRGAN generator using the hybrid enhanced data set, alternately train the generator and the discriminator, and optimize through a loss function to obtain a trained generator.

[0144] The data set fusion module is configured to perform super-resolution reconstruction on the original low-resolution image by using the trained generator to obtain new training image data, and combine image data based on the hyperparameter alpha by using a hierarchical sampling strategy according to the original data set, the mixed enhanced data set, and the new training image data to obtain enhanced image data.

[0145] The present application aims at the problem of insufficient scale and diversity of training data set in the text image super-resolution reconstruction task, and proposes a method and device combining mixed data enhancement and generative data enhancement. By using the weight mask mixing technology and the data generation based on the generative adversarial network, the distribution space of the training data is significantly expanded, the generalization ability and the reconstruction quality of the model are improved, and the method has important application prospect in the field of text image super-resolution reconstruction.

[0146] Some embodiments of the present application also provide an electronic device, including a processor, a memory, and a computer program stored on the memory and capable of running on the processor, and the computer program is executed by the processor to implement the method as above.

[0147] Some embodiments of the present application also provide a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method as above.

[0148] Some embodiments of the present application also provide a computer program product, including a computer program, and the computer program is executed by the processor to implement the method as above.

[0149] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts are referred to the part of the method embodiment.

[0150] Each embodiment in the specification is described in a progressive manner, and each embodiment mainly describes the difference from other embodiments, and the same and similar parts of each embodiment are referred to each other.

[0151] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, device, or computer program product. Therefore, the embodiments of the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can be in the form of a computer program product implemented on one or more computer usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0152] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0153] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0155] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0156] Finally, it is to be understood that the phraseology or terminology such as "first" and "second" etc. used herein is merely intended to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any actual such relationship or order between such entities or operations. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the aforesaid element.

[0157] The above provides a detailed description of the method and device for text image super-resolution reconstruction of data enhancement. The principles and implementation modes of the present application are described by using specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed. In summary, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A data-enhanced text image super-resolution reconstruction method, characterized in that: The method comprises: From the original dataset Get the training data set as training samples, use the improved CutBlur algorithm to process the samples, and obtain the hybrid enhanced data set ; The ESRGAN generator is pre-trained using a mixed augmented dataset, and the generator and discriminator are alternately trained. The trained generator is then optimized using the loss function. The trained generator will be used to perform super-resolution reconstruction on the original low-resolution image to obtain new training image data , based on the original dataset , mixed augmented dataset And new training image data , based on the hyperparameter α, the image data is combined using a layered sampling strategy to obtain enhanced image data .

2. The method according to claim 1, characterized in that The original data set Get the training data set as training samples, use the improved CutBlur algorithm to process the samples, and obtain the hybrid enhanced data set Previously, the original dataset was also included , divided into training set, validation set and test set in the ratio of 8:1:

1.

3. The method according to claim 1, characterized in that The original data set Get the training data set as training samples, use the improved CutBlur algorithm to process the samples, and obtain the hybrid enhanced data set ;include: Randomly select the mask area ratio to meet the uniform distribution The rectangular mask area of ​​​​ Calculate adaptive smoothing parameters ,in , creates an enhanced image using a smooth weight mask; where is the basic smoothing parameter, is the area of ​​the mask region, is the total area of ​​the image; The quality of the enhanced images was evaluated using the learning-aware image patch similarity metric, and low-quality samples with LPIPS scores below the threshold of 0.3 were removed to obtain a mixed enhancement dataset. The calculation formula for the above LPIPS is: 。 4. The method according to claim 3, characterized in that The randomly selected mask area ratio conforms to the uniform distribution The rectangular mask area includes: Randomly select two sets of low-resolution-high-resolution image pairs from the training dataset ( ), the low-resolution image is recorded as , and its corresponding high-resolution image is , where the reconstruction magnification is represented by s; Randomly select a rectangular area in the image As the base mask, where the distance change function is calculated , the calculation formula is: 。 5. The method according to claim 3, characterized in that The calculation of the adaptive smoothing parameter ,in , creating enhanced images using smooth weight masks; include: Constructing a bidirectional sigmoid weight mask , ensuring that the weight is close to 1 in the central area inside the mask, and close to 0 in the area outside the mask away from the boundary; Near the mask boundary, the weights show a smooth S-shaped transition, and the transition band width can be adjusted by the parameter Control, using this smooth weight mask, creates an enhanced image where, The weight mask calculation formula is: , Generate enhanced images using smooth weight masks, The calculation formula is: , , in, represents element-wise multiplication, represents the downsampling operation, represents the upsampling operation, and Represent the pixel values ​​of low-resolution images and high-resolution images respectively.

6. The method according to claim 1, characterized in that The method sequentially uses a mixed augmented dataset to pre-train the ESRGAN generator, alternately trains the generator and the discriminator, and optimizes the training through a loss function to obtain a trained generator; including: The learning rate is set to 0.0002, the number of training rounds is 50, and the ESRGAN generator is pre-trained using the mixed augmentation dataset to obtain generated images; Parameter setting through loss function, including feature modulation coefficient , The Adam optimizer is used to set the generator learning rate to 0.001, the discriminator learning rate to 0.004, and the total number of training rounds to 200 rounds. According to the set parameters, the generated image is used as input data, and adversarial training is performed in an alternating training method of training the generator once and the discriminator once to obtain a trained generator.

7. The method according to claim 6, characterized in that The learning rate is set to 0.0002, the number of training rounds is 50, and the ESRGAN generator is pre-trained using the mixed enhanced dataset to obtain generated images; This includes the following processing through the feature extraction front-end and information processing backbone functional modules: In the front-end processing stage, the initial feature acquisition and feature dimension expansion are achieved through a single-layer convolution operation. Specifically, the shallow features Obtained by the first convolutional layer, it contains visual features, As shown in the following formula: , in, Refers to the initial convolution operation unit that processes the low-resolution input image and extracts the representation features It is then passed into the backbone of the network for deep feature mining; The core backbone network uses a cascaded residual feature aggregation unit to build a backbone network, which consists of N series-connected residual calculation units and feature modulation coefficients. The calculation formula is: , in, represents the mapping function of the i-th residual calculation unit, represents the information flow input to the i-th feature fusion module, is the output feature generated by the module; By reconstructing the deep features extracted from the partial First, it passes through a convolution layer, and then the resolution is expanded to obtain the generated image, as shown in the following formula: ; in, represents the super-resolution reconstructed image, represents the upsampling operation, They are the convolutional layer after the main part and the convolutional layer after upsampling.

8. The method according to claim 6, characterized in that The parameter setting by the loss function includes: Generator loss function Combined with adversarial loss , content loss and perceptual loss ,in, The calculation formula is: ; The adversarial loss , content loss To ensure the consistency of the generated image with the real high-resolution image at the pixel level, the calculation formula is: , perceptual loss Based on the pre-trained VGG network, the calculation formula is , Represents the feature extraction function of the pre-trained VGG network.

9. The method according to claim 1, characterized in that The trained generator will be used to perform super-resolution reconstruction on the original low-resolution image to obtain new training image data , based on the original dataset , mixed augmented dataset And new training image data , based on the hyperparameter α, a layered sampling strategy is used to combine image data to obtain enhanced image data that can be dynamically adjusted with the training results. ,include: Use the trained generator to reconstruct the original image through the reconstruction formula Super-resolution reconstruction of low-resolution images in the image is performed to generate new training image data : ; Based on the hyperparameter α that can be dynamically adjusted according to the training results, a stratified sampling strategy is used to sample new training image data. , original image And the mixed enhanced image Three types of data are combined to obtain enhanced image data. , where the formula for layered sampling combined image data is: ,in, is a hyperparameter.

10. A data-enhanced text image super-resolution reconstruction device, characterized in that: The device comprises: Hybrid data preprocessing module for raw data sets Get the training data set as training samples, use the improved CutBlur algorithm to process the samples, and obtain the hybrid enhanced data set ; The generative model training module is used to pre-train the ESRGAN generator using a mixed augmented dataset, alternately train the generator and discriminator, and optimize the loss function to obtain a trained generator; The dataset fusion module is used to use the trained generator to perform super-resolution reconstruction on the original low-resolution image to obtain new training image data. , based on the original dataset , mixed augmented dataset And new training image data , based on the hyperparameter α, the image data is combined using a layered sampling strategy to obtain enhanced image data .

Citation Information

Patent Citations

  • Heart magnetic resonance image data enhancement method based on evolutionary GAN

    CN111861924A

  • Novel super-resolution reconstruction method based on convolutional neural network

    CN113674149A

  • Infrared image super-resolution reconstruction method based on deep neural network

    CN114913069A

  • Weak supervision change detection method and device based on background mixed data expansion technology

    CN115641316A

  • Real world image super-resolution method based on stable diffusion

    CN118918009A

Cited By

  • Testing method and device for calculating robustness of imaging model

    CN121582717A