Scene character image super-resolution method for generative image prior

Through the scene text image super-resolution method of generative image priors, the multimodal diffusion model and ITPGDM model are used to solve the problems of missing semantic information and insufficient prior diversity in low-resolution images, and high-quality image reconstruction and high-accuracy text recognition are achieved.

CN119941509APending Publication Date: 2025-05-06CHINA UNIV OF MINING & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510014159.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the super-resolution task of scene text image, the semantic information of low-resolution images is missing and the prior diversity is insufficient, resulting in low text recognition accuracy.

Method used

Using the generative image prior method, high-resolution image priors and super-resolution text images are generated through multimodal diffusion model and ITPGDM model, and the semantic information and text recognition accuracy are enhanced.

Benefits of technology

It significantly improves the quality of low-resolution images, enhances the accuracy of text recognition, and solves the problems of missing semantic information and insufficient prior diversity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941509A_ABST
    Figure CN119941509A_ABST
Patent Text Reader

Abstract

The invention discloses a scene character image super-resolution method for generative image prior. The method comprises two stages. In the first stage, a diffusion model based on multiple modes is constructed, a GPT model is used for obtaining specific text information from a low-resolution character image, and high-resolution image prior is generated; in the second stage, an ITPGDM model is constructed, a high-resolution character image is reconstructed through high-resolution image prior and character recognition prior, the ITPGDM model comprises a PSAB module and a CFAB module, the PSAB module is used for aligning different prior information, and the CFAB module is used for refining character-level features; the ITPGDM model represents a scene text picture super-resolution diffusion model based on image and text prior guidance, the PSAB module represents a prior semantic alignment module, and the CFAB module represents a character attention module. According to the method, the powerful advantages of the diffusion model and the GPT model are fully utilized, and the super-resolution capability of the scene character image is enhanced by using the multi-prior semantic alignment module and the character attention module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a generative image prior scene text image super-resolution method, belonging to the scene text image super-resolution technology. Background Art

[0002] Scene text recognition (STR) focuses on extracting text from images and has made significant progress and is widely used in autonomous driving, document scanning, image retrieval and other fields. However, scene text images often experience multiple quality degradations when they are captured, resulting in low resolution and blur, which seriously affects the effect of text recognition models. To solve this problem, researchers began to study scene text image super-resolution (STISR).

[0003] As the preprocessing stage of STR, STISR aims to reconstruct the missing text details in low-resolution images to improve the accuracy of text recognition. Currently, the methods used by STISR to improve the quality of text images are mainly divided into two categories: general super-resolution and super-resolution guided by text recognition priors. Early STISR tasks mainly used traditional super-resolution techniques to improve the resolution of images. These methods did not fully focus on inherent text and structural features and showed obvious limitations. There is also a category of super-resolution that uses text-guided super-resolution to improve fidelity and recognition accuracy, such as using character probability sequences as text priors to guide.

[0004] In summary, there are still two major challenges for the STISR task:

[0005] 1. Inaccurate semantic information: Blurred text and missing details on low-resolution scene text images lead to the loss of semantic information, which prevents STISR from fully reconstructing the original content.

[0006] 2. Insufficient diversity: STISR based on text prior uses low-resolution text recognition results, resulting in insufficient diversity of priors. Summary of the invention

[0007] Purpose of the invention: Scene text image super-resolution aims to reconstruct low-resolution text images lacking text details into high-resolution text images with high fidelity and semantic accuracy, which is a key task to improve the accuracy of text recognition; although the existing STISR has made progress, due to degradation such as low resolution and jitter, the existing methods face challenges in effectively reconstructing the original text details and accurate semantic information. In order to overcome these challenges, the present invention provides a generative image prior scene text image super-resolution method, which can significantly improve the quality of low-resolution images and enhance the accuracy of text recognition.

[0008] Technical solution: To achieve the above purpose, the technical solution adopted by the present invention is:

[0009] A method for super-resolution of scene text images based on generative image priors, the method comprising two stages; in the first stage, a diffusion model based on multimodality is constructed, and a GPT model is used to obtain specific text information from low-resolution text images, thereby generating a high-resolution image prior, wherein the generated high-resolution image prior is close to the original scene, and thus can retain and enhance semantic information; in the second stage, an ITPGDM model is constructed, and a high-resolution text image is reconstructed through high-resolution image priors and text recognition priors, wherein the ITPGDM model comprises a PSAB module and a CFAB module, wherein the PSAB module is used to align different prior information, and the CFAB module is used to refine character-level features; the GPT model represents a multimodal generative pre-training model, the ITPGDM model represents a scene text image super-resolution diffusion model guided by image and text priors, the PSAB module represents a multi-prior semantic alignment module, and the CFAB module represents a character attention module.

[0010] Specifically, the method comprises the following steps:

[0011] (1) Using a multimodal diffusion model to build an IP generator for generating high-resolution image priors IP ; First, the high-resolution image I HR Input it into the pre-trained encoder to get the original hidden code h0, add noise to the original hidden code h0 to generate the noisy hidden code h at time t t ; At the same time, use the text label Label to draw a low-resolution text image I LR Standard text image I ST ; Then concatenate the noisy hidden code h in the channel dimension t and standard text image I ST Get f0 and input f0 into the denoising network Unet IP Then, the low-resolution text image I is obtained through the GPT model and the pre-trained CLIP text encoder LR The embedded text vector f p , embed the text vector f p and standard text image I ST Input to the denoising network Unet IP The MResAtt block in performs multimodal alignment to guide the denoising network Unet IP Perform the denoising process to generate the noisy hidden code h at time t-1 t-1 ; The denoising process is repeated until the original latent code h0 is obtained, and the original latent code h0 is decoded using the pre-trained decoder to finally generate a high-resolution image prior I IP ; The MResAtt block represents a multimodal residual attention block;

[0012] (2) Use the training set to train the IP generator to obtain a trained IP generator;

[0013] (3) Constructing the ITPGDM model to generate high-resolution text images I SR ; First, the high-resolution image I HR Input into the pre-trained encoder and add noise to get the noisy hidden code z at time t t ; Then concatenate the noisy hidden code z in the channel dimension t and low-resolution text image I LR Get g0 and input g0 into the denoising network Unet SR Then, the low-resolution text image I is obtained through the TPG generator, clustering operation and trained IP generator. LR The text recognition prior f tp , low resolution image mask I LRM and high-resolution image prior I IP , the text recognition prior f tp , low resolution image mask I LRM and high-resolution image prior I IP First, it is input into the PSAB module for multi-prior alignment, and then input into the CFAB module for character feature refinement, and finally a super-resolution text image I is generated. SR ;

[0014] (4) Use the training set to train the ITPGDM model, and use the trained ITPGDM model to test on the test set to obtain super-resolution text images.

[0015] Specifically, in step (1), the pre-trained encoder converts the high-resolution image Project from pixel space to latent space and add a series of Gaussian noise to generate a noisy latent code that satisfies the Gaussian distribution Use TPG generator to obtain low-resolution text image I LR Use the text label Label to draw a low-resolution text image I LR Standard text image

[0016] I ST =Draw(Label)

[0017] h t =Diffusion(ε(I HR ), t)

[0018] Wherein: Draw represents the image drawing process, Diffusion represents the forward diffusion process, ε represents the pre-trained encoder; and the TPG generator represents the text prior generator.

[0019] Specifically, in step (1), in order to avoid semantic loss, the GPT model is used to accurately answer the low-resolution text image I according to the text prompt information. LR characteristics (such as font size, color, and background), and convert the low-resolution text image into LR The characteristics of are creatively described as text information prompts; the text information prompts are input into the pre-trained CLIP text encoder to obtain the embedded text vector f p ; Embed the text vector f p and standard text image I ST Input to the denoising network Unet IP The MResAtt block in performs multimodal alignment to guide the denoising network Unet IP The noisy hidden code h t Gradually restore to the original hidden code h0; including the following steps:

[0020] (11) The low-resolution text image I LR The text information Prompts generated by the GPT model is input into the pre-trained CLIP text encoder to obtain the embedded text vector f p ;

[0021] f p =Clip(Prompts)

[0022] Where: Clip represents the pre-trained CLIP text encoder;

[0023] (12) Concatenate noisy hidden codes h in the channel dimension t and standard text image I ST Get f0, embed f0 into text vector f p and standard text image I ST Input to the denoising network Unet IP The denoising process is performed in the process to obtain the noisy hidden code h at time t-1 t-1 :

[0024] f0=Concat(h t , I ST )

[0025] f1=Res(Res(CA(I ST ,Res(f0))+CA(f p ,Res(f0))))

[0026] f2=CA(I ST ,Res(f1))+CA(f p ,Res(f1))

[0027] h t-1 =f2

[0028] Among them: Concat represents channel concatenation, Res represents residual block, and CA represents cross attention mechanism, which is used to dynamically adjust Unet m The intermediate features of the network are f0, f1 and f2, t = 1, 2, ..., T; denoising network Unet IP It includes two MResAtt blocks, the input and output of the first MResAtt block are f0 and f1 respectively, and the input and output of the second MResAtt block are f1 and f2 respectively;

[0029] (13) t = t-1, repeat step (12), and cyclically perform the denoising process until the original latent code h0 is obtained. The original latent code h0 is decoded to obtain a high-resolution image prior I with diversity. IP :

[0030] I IP =D′(h0)

[0031] Where: D′ represents the pre-trained decoder.

[0032] Specifically, in step (2), the IP generator is trained using the training set until convergence; and the parameters of the trained IP generator are saved for training the ITPGDM model.

[0033] Specifically, in step (3), the pre-trained encoder converts the high-resolution image Project from pixel space to latent space and add a series of Gaussian noise to generate a noisy latent code that satisfies the Gaussian distribution

[0034]

[0035] z t =Diffusion(ε(I HR ), t)

[0036] Where: Diffusion represents the forward diffusion process, and ε represents the pre-trained encoder.

[0037] Specifically, in step (3), the low-resolution text image The text recognition prior f obtained in tp , low resolution image mask I LRM and high-resolution image prior I IPInput to the denoising network Unet SR The ITRes block in the image guides the ITPGDM model to generate super-resolution text images with high fidelity and accurate semantic information. SR , including the following steps:

[0038] (31) Obtain low-resolution image I through TPG generator, clustering operation and trained IP generator respectively LR The text recognition prior f tp , low resolution image mask I LRM and high-resolution image prior I IP :

[0039] f tp =TPG(I LR )

[0040] I LRM =Clustering(I LR )

[0041] I IP =IPG(I LR )

[0042] Where: TPG represents TPG generator, Clustering represents clustering operation, and IPG represents IP generator;

[0043] (32) Concatenate noisy hidden codes z in the channel dimension t and low-resolution text image I LR Get g0:

[0044] g0=Concat(z t , I LR )

[0045] Among them: Concat means channel concatenation;

[0046] (33) Input g0 into the denoising network Unet SR The denoising process is performed in , and the noisy hidden code z at time t-1 is obtained t-1 :

[0047] G1=ITRes(g0,f tp , I LRM , I IP )

[0048] g1=Res(Res(G1))

[0049] G2=ITRes(g1,f tp , I LRM , I IP )

[0050] z t-1 =G2

[0051] Among them: Res represents the residual block, ITRes represents the ITRes block, and the denoising network Unet SR There are two ITRes blocks in it, the input and output of the first ITRes block are g0 and G1 respectively, and the input and output of the second ITRes block are g1 and G2 respectively; the ITRes block represents the image text residual block;

[0052] (34) t = t-1, repeat steps (32) and (33), and cyclically perform the denoising process until the original latent code z0 is obtained. The original latent code z0 is decoded to generate a super-resolution text image

[0053] I SR =D′(z0)

[0054] Where: D′ represents the pre-trained decoder.

[0055] Specifically, in step (33), each ITRes block includes a residual block and an ITResAtt block, each ITResAtt block includes a PSAB module and a CFAB module, and the ITResAtt block represents an image text residual attention block; the data processing flow of the ITRes block includes the following steps:

[0056] (331) Different prior information is first aligned through the PSAB module to produce a complementary effect, and then the advantages of different prior information are fused by adding in the pixel direction to obtain the complementary feature f. pSAB ;The input of the ITRes block is g i-1 , g i-1 First, we pass a residual block to obtain g′ i-1 , g i-1 The three cross-attention mechanisms are used to fuse different prior information: the input of the first cross-attention mechanism is the low-resolution image mask I LRM and text recognition prior f tp , used to obtain semantic features; the input of the second cross attention mechanism is g′ i-1 , high-resolution image prior I IP and text recognition prior f tp , used to obtain HR-IP features; since semantic features and HR-IP features inevitably show differences, in order to solve this difference, semantic features and HR-IP features are aligned through pixel-wise addition and self-attention mechanisms; the input of the third cross-attention mechanism is g′ i-1 and text recognition prior f tp, the output of the third cross attention mechanism and the output of the self-attention mechanism are added in the pixel direction to obtain the complementary feature f PSAB ; Extract complementary features f through Intra-Strip operation PSAB The important features in Intra , extract important features f through Inter-Strip-H operation Intra Important text features in Inter :

[0057] g′ i-1 =Res(g i-1 )

[0058] f PSAB =CA(g′ i-1 , f tp )+SA(CA(g′ i-1 , f tp , I IP )+CA(I LRM , f tp ))

[0059] f Intra =Intra(f PsAB )

[0060] f Inter =Inter(f intra )

[0061] Where: i = 1, 2, Res represents the residual block, CA represents the cross attention mechanism, SA represents the self-attention mechanism, Intra represents the Intra-Strip operation (intra-strip attention operation), and Inter represents the Inter-Strip-H operation (inter-strip attention operation in the horizontal direction);

[0062] (332) The CFAB module is used to refine the character-level features. The CFAB module consists of two parts: one part is Inter-Strip-V, which focuses on the relationship between pixel columns, and the second part is CharFocus-V, which enhances the character-level features. Inter-Strip-V represents the attention operation between strips in the vertical direction, and CharFocus-V represents the character focus module in the vertical direction. The input of the CFAB module is important text features. Where: H′, W′ and C′ represent f Inter The height, f Inter The width and f Inter The number of channels; exchange f Inter The height and width of X are obtained by W′, H′ and C′ respectively; the dimensions of X obtained by linear projection are The query Q′, key K′, and value V′ are then reshaped into 2D tensors Q, K, and V of size W′×D:

[0063] X=Swap(f Inter )

[0064] (Q′, K′, V′) = (XP Q , XP K , XP V )

[0065] (Q, K, V) = reshape (Q′, K′, V′)

[0066] Among them: Swap means exchanging the height and width of the input feature; P Q , P K and P V denotes the linear layers of query projection, key-value projection, and value projection respectively; reshape denotes the reshaping operation; D = H′×C′;

[0067] (333) Divide X into W′ non-overlapping vertical strips And perform attention operations on these vertical strips to obtain the aggregate column features InterV:

[0068]

[0069] InterV=S×V

[0070] Wherein: Sigmoid represents the Sigmoid function; Indicates the output of Inter-Strip-V, represents the attention weight matrix of Inter-Strip-V, k = 1, 2, ..., W′;

[0071] (334) Obtain the forward boundary relative position prediction of the attention weight matrix S through one or more linear layers Transpose the attention weight matrix S to S T , obtain the reverse boundary relative position prediction of the attention weight matrix S through one or more linear layers For b and The elements at the same position are averaged to obtain the relative boundary

[0072] b=Linears(S)

[0073]

[0074] Among them: Linears represents multi-linear layer, mean represents the mean operation of elements at the same position;

[0075] (335) Divide the relative boundary B into the left relative position boundary along the column direction and right relative position border Two parts, and use the range operation to limit the relative position boundary to [0, 1]; P l Multiply by width W' to get the left border P r Multiply by width W' to get the right border Keep S and S according to the obtained left boundary L and right boundary R T The attention score on the boundary is obtained and Perform character-level enhancement on the aggregate column feature InterV to obtain the character-level enhanced feature E:

[0076] (P l , P r )=Split(B)

[0077] L=Clamp(P l )×W′

[0078] R = Clamp(Sigmoid(Clamp(P r )+Clamp(P l )))×W′

[0079] S [L,R] =Cut(S,L,R)

[0080] S T [L,R] =Cut(S T , L, R)

[0081] E=(S [L,R] +S T [L,R] )×InterV+InterV

[0082] Among them: Split represents the split operation along the column direction; Clamp represents the range limit operation; Sigmoid represents the Sigmoid function; Cut represents the operation of cutting according to the left and right boundaries;

[0083] (37) The character-level enhanced features E and important text features f Inter Concatenate in the channel dimension to get the output of the ITRes block:

[0084] G i =Concat(f Inter , E)

[0085] Among them: Concat means channel concatenation.

[0086] Beneficial effect: The generative image prior scene text image super-resolution method provided by the present invention has the following advantages over the prior art: 1. The IP generator in the present invention can directly generate diverse high-resolution image priors from low-resolution text images, effectively reducing the problem of prior loss; 2. The MResAtt block in the present invention can perform multimodal feature fusion, guiding the IP generator to generate accurate and diverse high-resolution image priors; 3. The multi-prior semantic alignment module designed by the present invention can effectively align multiple prior information; 4. The character attention module designed by the present invention can effectively enhance character-level features. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] Figure 1 1 is a flow chart for implementing the method of the present invention;

[0088] Figure 2 It is a structural schematic diagram of the IP generator in the method of the present invention;

[0089] Figure 3 It is a schematic diagram of the structure of the ITPGDM model in the method of the present invention;

[0090] Figure 4 It is a structural block diagram of the CFAB module in the method of the present invention. DETAILED DESCRIPTION

[0091] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0092] like Figure 1 The figure shows a method for super-resolution of scene text images based on generative image priors. The method includes two stages, each of which uses an implicit diffusion model. In the first stage, a diffusion model based on multimodality is constructed, and the GPT model is used to obtain specific text information from low-resolution text images, thereby generating high-resolution image priors. The generated high-resolution image priors tend to the original scene, so that semantic information can be retained and enhanced. In the second stage, an ITPGDM model is constructed to reconstruct high-resolution text images through high-resolution image priors and text recognition priors. The ITPGDM model includes a PSAB module and a CFAB module. The PSAB module is used to align different prior information, and the CFAB module is used to refine character-level features. The GPT model represents a multimodal generative pre-training model, and the ITPGDM model represents a scene text image super-resolution diffusion model guided by image and text priors. The PSAB module represents a multi-prior semantic alignment module, and the CFAB module represents a character attention module. The following is a detailed description of each stage.

[0093] Phase 1 01: Design IP Generator

[0094] The IP generator is constructed using a multimodal diffusion model to generate high-resolution image priors. IP First, the high-resolution image I HR Input it into the pre-trained encoder to get the original hidden code h0, add noise to the original hidden code h0 to generate the noisy hidden code h at time t t ; At the same time, use the text label Label to draw a low-resolution text image I LR Standard text image I ST ; Then concatenate the noisy hidden code h in the channel dimension t and standard text image I ST Get f0 and input f0 into the denoising network Unet IP Then, the low-resolution text image I is obtained through the GPT model and the pre-trained CLIP text encoder LR The embedded text vector f p , embed the text vector f p and standard text image I ST Input to the denoising network Unet IP The MResAtt block in performs multimodal alignment to guide the denoising network Unet IP Perform the denoising process to generate the noisy hidden code h at time t-1 t-1 ; The denoising process is repeated until the original latent code h0 is obtained, and the original latent code h0 is decoded using the pre-trained decoder to finally generate a high-resolution image prior I IP ; The MResAtt block represents a multimodal residual attention block.

[0095] Generate high-resolution image priors using IP Generator IP The process is as follows Figure 2 As shown, the steps are as follows:

[0096] Step S11: Use the pre-trained encoder to convert the high-resolution image Project from pixel space to latent space and add a series of Gaussian noise to generate a noisy latent code that satisfies the Gaussian distribution

[0097] h t =Diffusion(ε(I HR ), t)

[0098] Wherein: Diffusion represents the forward diffusion process, ε represents the pre-trained encoder; and the TPG generator represents the text prior generator.

[0099] Step S12: Use the TPG generator to obtain a low-resolution text image I LR Use the text label Label to draw a low-resolution text image I LR Standard text image

[0100] I ST =Draw(Label)

[0101] Wherein: Draw represents the image drawing process; the TPG generator represents the text prior generator.

[0102] Step S13: In order to avoid semantic loss, the GPT model is used to accurately answer the low-resolution text image I according to the text prompt information. LR characteristics (such as font size, color, and background), and convert the low-resolution text image into LR Features are creatively described as text message prompts.

[0103] The text prompt information consists of two parts, one part focuses on the text features and the other part focuses on the background features; Example of the format of the text prompt information: Please use a 40-word phrase to describe the text image in the following order: (1) Text features: font weight, all uppercase or all lowercase or first letter capitalized, position, font style, color of text, other useful features; (2) Background features: background color, other useful features.

[0104] Step S14: Convert the low-resolution text image I LR The text information Prompts generated by the GPT model is input into the pre-trained CLIP text encoder to obtain the embedded text vector f p ;

[0105] f p =Clip(Prompts)

[0106] Where: Clip represents the pre-trained CLIP text encoder.

[0107] Step S15: Concatenate the noisy hidden code h in the channel dimension t and standard text image I ST Get f0:

[0108] f0=Concat(h t , I ST )

[0109] Among them: Concat means channel concatenation.

[0110] Step S16: embed f0, text vector fp and standard text image I ST Input to the denoising network Unet IP The denoising process is performed in the process to obtain the noisy hidden code h at time t-1 t-1 :

[0111] f1=Res(Res(CA(I ST ,Res(f0))+CA(f p ,Res(f0))))

[0112] f2=CA(I ST ,Res(f1))+CA(f p ,Res(f1))

[0113] h t-1 =f2

[0114] Among them: Res represents the residual block, CA represents the cross attention mechanism, which is used to dynamically adjust Unet IP The intermediate features of the network are f0, f1 and f2, t = 1, 2, ..., T; denoising network Unet IP It includes two MResAtt blocks, the input and output of the first MResAtt block are f0 and f1 respectively, and the input and output of the second MResAtt block are f1 and f2 respectively.

[0115] Step S17: t = t-1, repeat steps S15 and S16, and execute the denoising process cyclically until the original latent code h0 is obtained, and decode the original latent code h0 to obtain a high-resolution image prior I with diversity. IP :

[0116] I IP =D′(h0)

[0117] Where: D′ represents the pre-trained decoder.

[0118] Phase 1 02: Training the IP Generator

[0119] Use the training set to train the IP generator until convergence to obtain the trained IP generator; save the parameters of the trained IP generator for training the ITPGDM model.

[0120] Phase 2 01: Designing the ITPGDM model

[0121] Constructing ITPGDM model to generate high-resolution text images SR ; First, the high-resolution image I HR Input into the pre-trained encoder and add noise to get the noisy hidden code z at time t t; Then concatenate the noisy hidden code z in the channel dimension t and low-resolution text image I LR Get g0 and input g0 into the denoising network Unet SR Then, the low-resolution text image I is obtained through the TPG generator, clustering operation and trained IP generator. LR The text recognition prior f tp , low resolution image mask I LRM and high-resolution image prior I IP , the text recognition prior f tp , low resolution image mask I LRM and high-resolution image prior I IP First, it is input into the prior semantic alignment module for multiple prior alignment, and then input into the character attention module for character feature refinement, and finally a super-resolution text image I is generated. SR .

[0122] Generate super-resolution text images using ITPGDM module I SR The flow chart is as follows Figure 3 As shown, the following steps are included:

[0123] Step S21: Use the pre-trained encoder to convert the high-resolution image Project from pixel space to latent space and add a series of Gaussian noise to generate a noisy latent code that satisfies the Gaussian distribution

[0124] z t =Diffusion(ε(I HR ), t)

[0125] Where: Diffusion represents the forward diffusion process, and ε represents the pre-trained encoder.

[0126] Step S22: Obtain low-resolution images through the TPG generator, clustering operation and trained IP generator respectively The text recognition prior f tp , low resolution image mask I LRM and high-resolution image prior I IP :

[0127] f tp =TPG(I LR )

[0128] I LRM =Clustering(I LR )

[0129] I IP =IPG(I LR )

[0130] Wherein: TPG represents TPG generator, Clustering represents clustering operation, and IPG represents IP generator.

[0131] Step S23: Concatenate the noisy hidden code z in the channel dimension t and low-resolution text image I LR Get g0:

[0132] g0=Concat(z t , I LR )

[0133] Among them: Concat means channel concatenation.

[0134] Step S24: g0, text recognition priori f tp , low resolution image mask I LRM and high-resolution image prior I IP Input denoising network Unet SR The denoising process is performed in , and the noisy hidden code z at time t-1 is obtained t-1 :

[0135] G1=ITRes(g0,f tp , I LRM , I IP )

[0136] g1=Res(Res(G1))

[0137] G2=ITRes(g1,f tp , I LRM , I IP )

[0138] z t-1 =G2

[0139] Among them: Res represents the residual block, ITRes represents the ITRes block, and the denoising network Unet SR There are two ITRes blocks in it, G1 and G2 represent the outputs of the first ITRes block and the second ITRes block respectively, and the ITRes block represents the image text residual block.

[0140] Each ITRes block includes a residual block and an ITResAtt block. Each ITResAtt block includes a PSAB module and a CFAB module. The ITResAtt block represents the image text residual attention block. The data processing flow of the ITRes block includes the following steps:

[0141] Step 241: Different prior information is first aligned through the PSAB module to produce a complementary effect, and then the advantages of different prior information are fused by addition in the pixel direction to obtain complementary features f PSAB The input of the ITRes block (and also the input of the PSAB module) is g i-1 , g i-1 First, we pass a residual block to obtain g′ i-1 , g′ i-1 The three cross-attention mechanisms are used to fuse different prior information: the input of the first cross-attention mechanism is the low-resolution image mask I LRM and text recognition prior f tp , used to obtain semantic features; the input of the second cross attention mechanism is g′ i-1 , high-resolution image prior I IP and text recognition prior f tp , used to obtain HR-IP features; since semantic features and HR-IP features inevitably show differences, in order to solve this difference, semantic features and HR-IP features are aligned through pixel-wise addition and self-attention mechanisms; the input of the third cross-attention mechanism is g′ i-1 and text recognition prior f tp , the output of the third cross attention mechanism and the output of the self-attention mechanism are added in the pixel direction to obtain the complementary feature f PSAB :

[0142] g′ i-1 =Res(g i-1 )

[0143] f PSAB =CA(g′ i-1 , f tp )+SA(CA(g′ i-1 , f tp , I IP )+CA(I LRM , f tp ))

[0144] Where: i = 1, 2, Res represents the residual block, CA represents the cross attention mechanism, and SA represents the self-attention mechanism.

[0145] Step 242: Extract complementary features f through Intra-Strip operation PSAB The important features in Intra , extract important features f through Inter-Strip-H operation Intra Important text features in Inter :

[0146] f Intra=Intra(f PSAB )

[0147] f Inter =Inter(f intra )

[0148] Among them: Intra represents Intra-Strip operation (intra-strip attention operation), Inter represents Inter-Strip-H operation (inter-strip attention operation in the horizontal direction).

[0149] Step 243: Use the CFAB module to refine the character-level features, such as Figure 4 As shown in the figure, the CFAB module consists of two parts, one part is Inter-Strip-V which focuses on the relationship between pixel columns, and the second part is CharFocus-V which enhances the character-level features. Inter-Strip-V represents the attention operation between strips in the vertical direction, and CharFocus-V represents the character focus module in the vertical direction.

[0150] The input of the CFAB module is important text features Where: H′, W′ and C′ represent f Inter The height, f Inter The width and f Inter The number of channels; exchange f Inter The height and width of X are obtained by W′, H′ and C′ respectively; the dimensions of X obtained by linear projection are The query Q′, key K′, and value V′ are then reshaped into 2D tensors Q, K, and V of size W′×D:

[0151] X=Swap(f Inter )

[0152] (Q′, K′, V′) = (XP Q , XP K , XP V )

[0153] (Q, K, V) = reshape (Q′, K′, V′)

[0154] Among them: Swap means exchanging the height and width of the input feature; P Q , P K and P V They represent the linear layers of query projection, key-value projection, and value projection respectively; reshape represents the reshaping operation, D = H′×C′.

[0155] Step 244: Divide X into W′ non-overlapping vertical strips And perform attention operations on these vertical strips to obtain the aggregate column features InterV:

[0156]

[0157] InterV=S×V

[0158] Wherein: Sigmoid represents the Sigmoid function; Indicates the output of Inter-Strip-V, Represents the attention weight matrix of Inter-Strip-V. In the attention weight matrix S, the weight of the rth row represents the vertical strip X of the rth column. r and all other columns have vertical strips X k The relationship is, r, k = 1, 2,…, W′.

[0159] Step 245: Obtain the forward boundary relative position prediction of the attention weight matrix S through one or more linear layers The rth row of b contains the vertical strip X r The relative position of the left border of b r,1 The relative position of the right border is b r,2 ; Transpose the attention weight matrix S to S T , obtain the reverse boundary relative position prediction of the attention weight matrix S through one or more linear layers For b and The elements at the same position are averaged to obtain the relative boundary

[0160] b=Linears(S)

[0161]

[0162] Among them: Linears represents multi-linear layer, mean represents the mean operation of elements at the same position;

[0163] Step 246: After obtaining the relative boundary B, project it onto the absolute attention weight matrix. Divide the relative boundary B into the left relative position boundary along the column direction: and right relative position border Two parts, and use the range operation to limit the relative position boundary to [0, 1]; P l Multiply by width W' to get the left border P r Multiply by width W' to get the right border Keep S and S according to the obtained left boundary L and right boundary R T The attention score in the boundary is obtained and Perform character-level enhancement on the aggregate column feature InterV to obtain the character-level enhanced feature E:

[0164] (P l , P r )=Split(B)

[0165] L=Clamp(P l )×W′

[0166] R = Clamp(Sigmoid(Clamp(P r )+Clamp(P l )))×W′

[0167] S [L,R] =Cut(S,L,R)

[0168] S T [L,R] =Cut(S T , L, R)

[0169] E=(S [L,R] +S T [L,R] )×InterV+InterV

[0170] Among them: Split represents the splitting operation along the column direction; Clamp represents the range limiting operation; Sigmoid represents the Sigmoid function; Cut represents the cutting operation according to the left and right boundaries.

[0171] Step 247: Combine the character-level enhanced features E and the important text features f Inter Concatenate in the channel dimension to get the output of the ITRes block:

[0172] G i =Concat(f Inter , E)

[0173] Among them: Concat means channel concatenation.

[0174] Step S25: t = t-1, repeat steps S23 and S24, and cyclically perform the denoising process until the original latent code z0 is obtained. The original latent code z0 is decoded using the pre-trained decoder to generate a super-resolution text image with high fidelity and accurate semantic information.

[0175] I SR =D′(z0)

[0176] Where: D′ represents the pre-trained decoder.

[0177] Second stage 02: Training ITPGDM model

[0178] The ITPGDM model is trained using the training set, and the trained ITPGDM model is tested on the test set to obtain super-resolution text images.

[0179] A generative image prior scene text image super-resolution device, including an IP generator and an ITPGDM model designed using a latent diffusion model. The IP generator mainly includes a pre-trained encoder, a pre-trained CLIP text encoder, and a denoising network Unet. IP . And pre-trained decoder, used to generate high-resolution image priors; ITPGDM model mainly includes pre-trained encoder, TPG generator, clustering mechanism, trained IP generator, denoising network Unet SR and a pre-trained decoder to generate super-resolution text images I SR .

[0180] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any form, and any technical solution obtained by equivalent replacement or equivalent transformation falls within the protection scope of the present invention.

Claims

1. A method for super-resolution of scene text images based on generative image priors, characterized by: The method includes two stages. In the first stage, a diffusion model based on multimodality is constructed, and the GPT model is used to obtain text information from low-resolution text images, thereby generating high-resolution image priors. In the second stage, an ITPGDM model is constructed to reconstruct high-resolution text images through high-resolution image priors and text recognition priors. The ITPGDM model includes a PSAB module and a CFAB module. The PSAB module is used to align different prior information, and the CFAB module is used to refine character-level features. The GPT model represents a multimodal generative pre-training model, the ITPGDM model represents a scene text image super-resolution diffusion model guided by image and text priors, the PSAB module represents a multi-prior semantic alignment module, and the CFAB module represents a character attention module.

2. The method for super-resolution of scene text images based on generative image prior according to claim 1, characterized in that: The steps include: (1) Using a multimodal diffusion model to build an IP generator for generating high-resolution image priors IP ; First, the high-resolution image I HR Input it into the pre-trained encoder to get the original hidden code h0, add noise to the original hidden code h0 to generate the noisy hidden code h at time t t ; At the same time, use the text label Label to draw a low-resolution text image I LR Standard text image I ST ; Then concatenate the noisy hidden code h in the channel dimension t and standard text image I ST Get f0 and input f0 into the denoising network Unet IP Then, the low-resolution text image I is obtained through the GPT model and the pre-trained CLIP text encoder LR The embedded text vector f p , embed the text vector f p and standard text image I ST Input to the denoising network Unet IP The MResAtt block in performs multimodal alignment to guide the denoising network Unet IP Perform the denoising process to generate the noisy hidden code h at time t-1 t-1 ; The denoising process is repeated until the original latent code h0 is obtained, and the original latent code h0 is decoded using the pre-trained decoder to finally generate a high-resolution image prior I IP ; The MResAtt block represents a multimodal residual attention block; (2) Use the training set to train the IP generator to obtain a trained IP generator; (3) Constructing the ITPGDM model to generate high-resolution text images I SR ; First, the high-resolution image I HR Input into the pre-trained encoder and add noise to get the noisy hidden code z at time t t ; Then concatenate the noisy hidden code z in the channel dimension t and low-resolution text image I LR Get g0 and input g0 into the denoising network Unet SR Then, the low-resolution text image I is obtained through the TPG generator, clustering operation and trained IP generator. LR The text recognition prior f tp , low resolution image mask I LRM and high-resolution image prior I IP , the text recognition prior f tp , low resolution image mask I LRM and high-resolution image prior I IP First, it is input into the PSAB module for multi-prior alignment, and then input into the CFAB module for character feature refinement, and finally a super-resolution text image I is generated. SR ; (4) Use the training set to train the ITPGDM model, and use the trained ITPGDM model to test on the test set to obtain super-resolution text images.

3. The method for scene text image super-resolution based on generative image prior according to claim 1, characterized in that: In step (1), the pre-trained encoder converts the high-resolution image Project from pixel space to latent space and add a series of Gaussian noise to generate a noisy latent code that satisfies the Gaussian distribution Use TPG generator to obtain low-resolution text image I LR Use the text label Label to draw a low-resolution text image I LR Standard text image 4. The method for super-resolution of scene text images based on generative image prior according to claim 1, characterized in that: In step (1), the GPT model is used to answer the low-resolution text image I according to the text prompt information. LR The low-resolution text image I LR The characteristics of the text are described as text information Prompts; the text information Prompts is input into the pre-trained CLIP text encoder to obtain the embedded text vector f p ; Embed the text vector f p and standard text image I ST Input to the denoising network Unet IP The MResAtt block in performs multimodal alignment to guide the denoising network Unet IP The noisy hidden code h t Gradually restore to the original hidden code h0; including the following steps: (11) The low-resolution text image I LR The text information Prompts generated by the GPT model is input into the pre-trained CLIP text encoder to obtain the embedded text vector f p ; (12) Concatenate noisy hidden codes h in the channel dimension t and standard text image I ST Get f0, embed f0 into text vector f p and standard text image I ST Input to the denoising network Unet IP The denoising process is performed in the process to obtain the noisy hidden code h at time t-1 t-1 : f0=Concat(h t ,I ST ) f1=Res(Res(CA(I ST ,Res(f0))+CA(f p ,Res(f0)))) f2=CA(I ST ,Res(f1))+CA(f p ,Res(f1)) h t-1 =f2 Where: t = 1, 2, ..., T, Concat represents channel concatenation, Res represents residual block, CA represents cross attention mechanism; denoising network Unet IP It includes two MResAtt blocks, the input and output of the first MResAtt block are f0 and f1 respectively, and the input and output of the second MResAtt block are f1 and f2 respectively; (13) t = t-1, repeat step (12), and cyclically perform the denoising process until the original latent code h0 is obtained. The original latent code h0 is decoded to obtain a high-resolution image prior I with diversity. IP .

5. The method for scene text image super-resolution based on generative image prior according to claim 1, characterized in that: In the step (2), the IP generator is trained using the training set until convergence; and the parameters of the trained IP generator are saved for training the ITPGDM model.

6. The method for scene text image super-resolution based on generative image prior according to claim 1, characterized in that: In step (3), the pre-trained encoder converts the high-resolution image Project from pixel space to latent space and add a series of Gaussian noise to generate a noisy latent code that satisfies the Gaussian distribution 7. The method for scene text image super-resolution based on generative image prior according to claim 1, characterized in that: In the step (3), the low-resolution text image The text recognition prior f obtained in tp , low resolution image mask I LRM and high-resolution image prior I IP Input to the denoising network Unet SR The ITRes block in the image guides the ITPGDM model to generate super-resolution text images. SR , including the following steps: (31) Obtain low-resolution image I through TPG generator, clustering operation and trained IP generator respectively LR The text recognition prior f tp , low resolution image mask I LRM and high-resolution image prior I IP ; (32) Concatenate noisy hidden codes z in the channel dimension t and low-resolution text image I LR Get g0: g0=Concat(z t ,I LR ) Among them: Concat means channel concatenation; (33) Input g0 into the denoising network Unet SR The denoising process is performed in , and the noisy hidden code z at time t-1 is obtained t-1 : G1=ITRes(g0,f tp ,I LRM ,I IP ) g1=Res(Res(G1)) G2=ITRes(g1,f tp ,I LRM ,I IP ) z t-1 =G2 Among them: Res represents the residual block, ITRes represents the ITRes block, and the denoising network Unet SR There are two ITRes blocks in it, the input and output of the first ITRes block are g0 and G1 respectively, and the input and output of the second ITRes block are g1 and G2 respectively; the ITRes block represents the image text residual block; (34) t = t-1, repeat steps (32) and (33), and cyclically perform the denoising process until the original latent code z0 is obtained. The original latent code z0 is decoded to generate a super-resolution text image 8. The method for scene text image super-resolution based on generative image prior according to claim 1, characterized in that: In the step (33), each ITRes block includes a residual block and an ITResAtt block, each ITResAtt block includes a PSAB module and a CFAB module, and the ITResAtt block represents an image text residual attention block; the data processing flow of the ITRes block includes the following steps: (331) Different prior information is first aligned through the PSAB module to produce a complementary effect, and then the advantages of different prior information are fused by adding in the pixel direction to obtain the complementary feature f. PSAB ; Extract complementary features f through Intra-Strip operation PSAB The important features in Intra , extract important features f through Inter-Strip-H operation Intra Important text features in Inter : g′ i-1 =Res(g i-1 ) f PSAB =CA(g′ i-1 ,f tp )+SA(CA(g′ i-1 ,f tp ,I IP )+CA(I LRM ,f tp )) Where: i = 1, 2, Res represents the residual block, CA represents the cross attention mechanism, SA represents the self-attention mechanism; Intra-Strip operation represents the intra-strip attention operation, and Inter-Strip-H operation represents the inter-strip attention operation in the horizontal direction; (332) The input of the CFAB module is important text features Where: H′, W′ and C′ represent f Inter The height, f Inter The width and f Inter The number of channels; exchange f Inter The height and width of X are obtained by W′, H′ and C′ respectively; the dimensions of X obtained by linear projection are The query Q′, key value K′ and value V′ are then reshaped into 2D tensors Q, K and V of size W′×D, where D=H′×C′; (333) Divide X into W′ non-overlapping vertical strips And perform attention operations on these vertical strips to obtain the aggregate column features InterV: InterV=S×V Wherein: Sigmoid represents the Sigmoid function; (334) Obtain the forward boundary relative position prediction of the attention weight matrix S through one or more linear layers Transpose the attention weight matrix S to S T , obtain the reverse boundary relative position prediction of the attention weight matrix S through one or more linear layers For b and The elements at the same position are averaged to obtain the relative boundary (335) Divide the relative boundary B into the left relative position boundary along the column direction and right relative position border Two parts, and use the range operation to limit the relative position boundary to [0, 1]; P l Multiply by width W' to get the left border P r Multiply by width W' to get the right border Keep S and S according to the obtained left boundary L and right boundary R T The attention score on the boundary is obtained and Perform character-level enhancement on the aggregate column feature InterV to obtain the character-level enhanced feature E: S [L,R] =Cut(S,L,R) S T [L,R] =Cut(S T ,L,R) E=(S [L,R] +S T [L,R] )×InterV+InterV Among them: Cut means the operation of cutting according to the left and right boundaries; (37) The character-level enhanced features E and important text features f Inter Concatenate in the channel dimension to get the output of the ITRes block: G i =Concat(f Inter ,E) Among them: Concat means channel concatenation.

Citation Information

Cited By

  • Image preprocessing method, system and equipment based on OCR (Optical Character Recognition), medium and product

    CN120236285A