A text image reconstruction method and device

By extracting and fusing feature and probabilistic information from text images, the problems of deformation and blurring in super-resolution reconstruction of text images are solved, thereby improving image quality and the accuracy of text information.

CN116416135BActive Publication Date: 2026-07-21UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UNIV OF SCI & TECH OF CHINA
Filing Date
2023-04-11
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, super-resolution reconstruction of text images suffers from distortion and blurring issues, resulting in low-quality reconstructed super-resolution images.

Method used

By extracting the horizontal, vertical, and channel features of the text image, fusing them, and combining them with text probability information, the contextual relationships between characters are optimized to construct a super-resolution image.

Benefits of technology

It improves the quality of super-resolution images, reduces image blur and distortion, and presents text information more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116416135B_ABST
    Figure CN116416135B_ABST
Patent Text Reader

Abstract

The application provides a text image reconstruction method and device. In the application, transverse features, longitudinal features and channel features in a text image to be reconstructed are extracted to fully obtain image detail features of the text image, so that the fused image features fully reflect various details in the image, thereby solving the problem of blurring after reconstruction. On the other hand, text probability information in the text image is determined and optimized, so that the reconstructed text information is more accurate. By using the method, various details of the text image can be reconstructed, the blurring problem of the reconstructed text image can be improved, and more accurate text information can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and more specifically, to a method and apparatus for text image reconstruction. Background Technology

[0002] Super-resolution reconstruction of text images has been widely applied in fields such as text recognition, signal recognition, and autonomous driving. Super-resolution reconstruction of text images refers to reconstructing a low-resolution text image into a high-resolution text image while preserving and restoring as much detail and texture as possible.

[0003] In real-world scenarios, low-resolution text images commonly suffer from problems such as blurring, distortion, or text edges blending with other objects, making super-resolution reconstruction of text images quite difficult. As a result, super-resolution images reconstructed from text images still generally suffer from distortion and blurring, resulting in low quality. Summary of the Invention

[0004] In view of this, this application provides a text image reconstruction method and apparatus to improve the quality of the reconstructed super-resolution image.

[0005] To achieve the above objectives, the following solution is proposed:

[0006] On the one hand, this application provides a text image reconstruction method, including:

[0007] Extract the horizontal, vertical, and channel features of the text image to be reconstructed;

[0008] The horizontal, vertical, and channel features are fused together to obtain the image features of the text image;

[0009] Determine the text probability information corresponding to the text image, wherein the text probability information characterizes at least one text region in the text image and the probability distribution of each character contained within the text region;

[0010] Based on the contextual relationship between characters within the text region reflected by the text probability information, the text probability information is optimized to obtain optimized text probability information;

[0011] Based on the image features and the optimized text probability information, a super-resolution image of the text image is constructed.

[0012] In one possible implementation, constructing a super-resolution image of the text image based on the image features and the optimized text probability information includes:

[0013] The text probability information and the optimized text probability information are fused to obtain the text features of the text image;

[0014] Based on the image features and the text features, a super-resolution image of the text image is constructed.

[0015] In another possible implementation, the extraction of the horizontal, vertical, and channel features of the text image to be reconstructed includes:

[0016] Feature extraction is performed on the text image to be reconstructed to obtain a basic feature map;

[0017] The base feature map is convolved to obtain the enhanced feature map;

[0018] The horizontal, vertical, and channel features of the enhanced feature map are extracted respectively.

[0019] In another possible implementation, before extracting the horizontal, vertical, and channel features of the text image to be reconstructed, the following is also included:

[0020] The reconstructed text image is then corrected.

[0021] In another possible implementation, the fusion processing of the text probability information and the optimized text probability information to obtain the text features of the text image includes:

[0022] Determine the fusion weights corresponding to the text probability information and the optimized text probability information respectively;

[0023] Based on the fusion weights corresponding to the text probability information and the optimized text probability information, the text probability information and the optimized text probability information are fused to obtain the text features of the text image.

[0024] In another possible implementation, constructing a super-resolution image of the text image based on the image features and the text features includes:

[0025] The image features and text features are aligned according to their dimensions, and the aligned image features and text features are concatenated to obtain a joint feature map;

[0026] Extract the lateral features from the joint feature map to obtain a preliminary fused feature map;

[0027] Extract the vertical features from the preliminary fused feature map to obtain the target fused feature map;

[0028] The target fusion feature map is upsampled to obtain a super-resolution image.

[0029] In another possible implementation, the process of fusing the horizontal features, vertical features, and channel features to obtain the image features of the text image includes:

[0030] Determine the fusion weights corresponding to the horizontal features, vertical features, and channel features respectively;

[0031] Based on the fusion weights corresponding to the horizontal, vertical, and channel features, the horizontal, vertical, and channel features are fused to obtain image features.

[0032] Furthermore, this application also provides a text image reconstruction apparatus, comprising:

[0033] The feature extraction unit is used to extract the horizontal, vertical, and channel features of the text image to be reconstructed.

[0034] The feature fusion unit is used to fuse the horizontal features, vertical features, and channel features to obtain the image features of the text image;

[0035] An information determination unit is used to determine the text probability information corresponding to the text image, wherein the text probability information characterizes at least one text region in the text image and the probability distribution of each character contained within the text region;

[0036] The information optimization unit is used to optimize the text probability information based on the contextual relationship between each character in the text region reflected by the text probability information, so as to obtain optimized text probability information;

[0037] An image reconstruction unit is used to construct a super-resolution image of the text image based on the image features and the optimized text probability information.

[0038] In one possible implementation, the image reconstruction unit includes:

[0039] The information fusion subunit is used to fuse the text probability information and the optimized text probability information to obtain the text features of the text image.

[0040] The image reconstruction subunit is used to construct a super-resolution image of the text image based on the image features and the text features.

[0041] In yet another possible implementation, the feature extraction unit includes:

[0042] The feature extraction subunit is used to extract features from the text image to be reconstructed to obtain a basic feature map;

[0043] The convolution processing subunit is used to perform convolution processing on the basic feature map to obtain the enhanced feature map;

[0044] The dimensional feature extraction subunit is used to extract the horizontal, vertical, and channel features of the enhanced feature map, respectively.

[0045] As can be seen from the above, in this embodiment, on the one hand, the horizontal, vertical, and channel features of the text image to be reconstructed are extracted, thereby fully obtaining the image detail features of the text image. This allows the fused image features to more fully reflect the various image details in the text image, reducing image blurring after reconstruction. On the other hand, after extracting the text probability information of the text image, this application further optimizes the text probability information based on the contextual relationships between characters in each text region reflected by the text probability information. This optimizes the text probability information to more accurately represent the text feature information in the text image, reducing text information errors in the reconstructed image. Based on this, using the image features of the text image determined by this application and the optimized text probability information to reconstruct the text image not only allows the reconstructed super-resolution image to fully reflect the various image details in the text image and reduce image blurring, but also more accurately presents the text information in the text image, improving the image quality of the reconstructed super-resolution image. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0047] Figure 1 This illustration shows a flowchart of a text image reconstruction method provided in an embodiment of this application;

[0048] Figure 2 This illustration shows another flowchart of a text image reconstruction method provided in an embodiment of this application;

[0049] Figure 3 This illustration shows a schematic diagram of the implementation principle framework of the text image reconstruction method provided in an embodiment of this application;

[0050] Figure 4 This illustration shows a structural schematic diagram of a text image reconstruction apparatus provided in an embodiment of this application. Detailed Implementation

[0051] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0052] The following describes a text image reconstruction method provided by an embodiment of this application.

[0053] like Figure 1 The illustration shows a flowchart of a text image reconstruction method provided in an embodiment of this application. This application can be applied to computer devices that perform text image reconstruction without limitation.

[0054] The method in this embodiment includes:

[0055] Step S100: Extract the horizontal features, vertical features, and channel features of the text image to be reconstructed.

[0056] Text images refer to images that contain text information.

[0057] Because some text images have low resolution, they need to be reconstructed to obtain super-resolution text images. Therefore, the text images to be reconstructed in this application refer to low-resolution text images.

[0058] The reconstruction of text images can be divided into two parts: image reconstruction and text reconstruction. Therefore, feature extraction is required for the text image to be reconstructed, including extracting image features representing image information and extracting text features representing text information.

[0059] In this application, image features are extracted from different dimensions, namely, channel features, horizontal features, and vertical features. Channel features refer to the color channels in the image, namely the three RGB channels. Horizontal features refer to the features in the width direction of the image, and vertical features refer to the features in the vertical direction of the image.

[0060] Step S110: The horizontal features, vertical features, and channel features are fused to obtain the image features of the text image.

[0061] After obtaining the horizontal, vertical, and channel features of the text image to be reconstructed, these three features need to be fused to obtain the image features in order to analyze the features of the image information as a whole.

[0062] There are several ways to perform fusion processing; let's take one possible approach as an example.

[0063] Horizontal, vertical, and channel features can be input into the dynamic fusion module for feature fusion processing. This dynamic fusion module uses a gated feature fusion module to fuse these three dimensions of features. The gating mechanism controls the contribution of each feature, specifically determining the fusion weights for the horizontal, vertical, and channel features respectively.

[0064] For example, Formula 1 can be used to calculate the fusion weight of each feature, where i represents a feature; Gating(i) represents the fusion weight corresponding to feature i; W and b are two parameters of the gating function; σ(·) is the normalized sigmoid function, which is used to map the input value to the range of 0 to 1.

[0065] Gating(i)=σ(Wi+b) (Formula 1).

[0066] Furthermore, the horizontal features, vertical features, and channel features, along with their corresponding fusion weights, are fused. For example, the fused feature obtained by fusing the horizontal features, vertical features, and channel features can be obtained using the following formula:

[0067] f df =Gating(i h )⊙i h +Gating(i v )⊙i v +Gating(i c )⊙i c (Formula 2);

[0068] Among them, f df This represents the fused feature, which can be used as an image feature of the text image. ⊙ represents the Hadamard multiplication representation of element-wise multiplication. h Indicates lateral features, i v Representing vertical features, i c Indicates channel characteristics.

[0069] Furthermore, to make the fused features more prominent, this application uses a Multi-Layer Perceptron (MLP) to enhance the fused features. The MLP(·) can consist of two 1x1 convolutions and multiple normalization layers, as shown in Equation 3:

[0070] i lr' =MLP(f df (i h iv i c (Formula 3);

[0071] Among them, i lr' The enhanced fusion feature in the middle can be used as the final extracted image feature. The image feature is the result of fusing the three features mentioned earlier: horizontal, vertical, and channel features, thus representing various detailed features of the text image to be reconstructed from multiple perspectives. df (·) is a gated fusion module.

[0072] Step S120: Determine the text probability information corresponding to the text image.

[0073] Text images contain not only image information but also text information. Therefore, it is necessary to extract the features of the text information from the text image, thereby extracting the text probability information. For example, a pre-trained text recognition device can be used to extract the text probability information of the text image.

[0074] The text probability information includes: information about the text regions contained in the text image and the probability distribution of each character contained in each text region. There is no restriction on whether there is one or more text regions. Each text region may include a string of one or more characters.

[0075] The text probability information can be represented by a feature matrix. Each row of the feature matrix corresponds to a text region in the text image. The data in this row represents the probability distribution of each character in the text region. Therefore, the feature matrix can reflect the probability distribution of each character in the strings that may be contained in each text region.

[0076] Understandably, the probability distribution of a character is essentially an estimation of what that character is, thus outputting the probability that the character belongs to any character in a given character list. The probability distribution of each character reflects the recognition result. For example, after determining the probability of the current character belonging to each of the 26 English characters, if the probability of the character being "a" is the highest, then that character will be recognized as "a". Based on this, text probability information can characterize the probability distribution of each character in the string contained within each text region of a text image.

[0077] Step S130: Based on the contextual relationships between characters within the text region reflected by the text probability information, optimize the text probability information to obtain optimized text probability information. It is noted that some errors may occur when using a text recognizer to recognize text images; therefore, the obtained text probability information can be optimized to correct the probabilities of characters with deviations.

[0078] It is understandable that text probability information can characterize the text regions contained in a text image and the probability distribution of the characters contained in each text region. Based on this, the text probability information can be used to obtain the possible strings contained in each text region of the text image, and naturally, the contextual relationship between each character in the string can also be obtained.

[0079] The contextual relationships between characters within a text area can indicate whether there are any recognition errors. For example, the semantic relationships between characters within a text area can be determined by combining the contextual relationships between characters within the text area, and the text probability information can be optimized by combining the semantic relationships.

[0080] For example, if the contextual relationships between characters within a text area indicate that there are semantic errors among the characters, it means that the probability distribution of the characters within the text area is incorrect. In this case, the probability distribution of the characters within the text area needs to be adjusted until it is possible to determine that there are no abnormalities in the semantic relationships between the characters within the text area by combining the probability distribution of the characters within the text area.

[0081] In one possible implementation, a trained attention-based Transformer model is used to optimize the text probability information, thereby obtaining optimized text probability information.

[0082] Step S140: Construct a super-resolution image of the text image based on image features and optimized text probability information.

[0083] After obtaining the image features and the optimized text probability information, the super-resolution image can be constructed using a general method for reconstructing super-resolution images. In this application, no restrictions are placed on this method.

[0084] In this application, on the one hand, horizontal, vertical, and channel features are extracted from the text image to be reconstructed, thereby fully obtaining the image detail features of the text image. This allows the fused image features to more fully reflect the various image details in the text image, reducing image blurring after reconstruction. On the other hand, after extracting the text probability information of the text image, this application further optimizes the text probability information based on the contextual relationships between characters in each text region reflected by the text probability information. This optimizes the text probability information to more accurately represent the text feature information in the text image, reducing the possibility of text information errors in the reconstructed image. Based on this, using the image features of the text image determined by this application and the optimized text probability information to reconstruct the text image not only allows the reconstructed super-resolution image to fully reflect the various image details in the text image and reduce image blurring, but also more accurately presents the text information in the text image, improving the image quality of the reconstructed super-resolution image.

[0085] It is understood that there may be some deformation in the text image to be reconstructed. In order to reduce the impact of image deformation on the quality of the reconstructed image, this application can first perform a correction process on the text image before extracting various features. Subsequently, image features and text probability information can be extracted based on the corrected text image.

[0086] To facilitate the correction of text image distortion, a trained spatial deformation network can be used to correct the text image. For example, the correction of a text image can be expressed as Formula 4:

[0087] I LR' =STN(I LR (Formula 4);

[0088] Among them, I LR ∈R C×H×W This represents the text image to be corrected, where H and W are the height and width of the text image, and C is the number of channels in the text image; I LR' For the corrected text image, STN() indicates that the text image is corrected using a spatial deformation network.

[0089] To facilitate understanding, the following section introduces one possible method for extracting the horizontal, vertical, and channel features of the text image to be reconstructed.

[0090] First, after obtaining the text image to be reconstructed, feature extraction can be performed on it. For example, a convolutional neural network (CNN) model with a single layer of large convolutional kernels can be used to extract the basic features of the corrected text image, resulting in a basic feature map. For example, obtaining the basic feature map i... lr The process can be represented by the following formula five:

[0091] i lr =CNN(I LR' (Formula 5);

[0092] Here, CNN(·) represents a convolutional neural network model.

[0093] Secondly, a convolutional neural network model is used to perform convolution processing on the features in the basic feature map to further enhance the features and obtain an enhanced feature map. For example, a convolutional neural network model containing three sets of 1×1 convolutional kernels can be used to perform convolution processing on the basic features.

[0094] Finally, the horizontal, vertical, and channel features of the enhanced feature map are extracted respectively. For example, convolutional neural networks with kernels of 1×7, 7×1, and 1×1 are used to extract image features from the horizontal, vertical, and channel dimensions, respectively, thus obtaining three-dimensional image features. As shown in Formula 6,

[0095] i h i v i c =f axis (f trans (i lr (Formula Six);

[0096] Among them, i h Indicates lateral features, i v Representing vertical features, i c Indicates channel characteristics. f axis and f trans These are two convolutional neural network models, one for extracting three types of features and the other for performing convolutional processing on the basic features.

[0097] To make the text information in the obtained text image more accurate, the text probability information and the optimized text probability information can be fused to further enrich the text information and thus obtain text features.

[0098] After obtaining the text probability information and the optimized text probability information, the dimensions of the text probability information and the optimized text probability information are aligned with the dimensions of the image features. Then, the aligned text probability information and the optimized text probability information are fused using a gated fusion module, as shown in Formula 7.

[0099] h t =f df (f align (f rec (I LR )),f aligh (f opt (f rec (I LR (Formula 7);

[0100] Among them, I LR For the text image to be reconstructed, f rec For text recognition, h rec =f rec (I LR ), where h rec ∈R L ×|A| L represents the length of the recognized text, and |A| represents the length of the preset alphabet.

[0101] Among them, f opt This is the text probability information optimization module, used to optimize text probability information. It uses Transformer to optimize the text probability information h. rec This module uses an attention mechanism to model the contextual relationships between words, optimizing the probability distribution predicted by the text recognizer, and obtaining the optimized text probability information h. opt ∈R L×|A| .

[0102] Among them, f aligh (·) is the alignment module, which first uses deconvolution to reduce the feature size of the text probability information from R. L ×|A| Transform into R C'×N Where C' is the feature dimension after deconvolution, and N is the feature size after deconvolution. Then, interpolation is used to convert the feature size into the image feature size, i.e., from R... C'×N Convert to R C'×H×W .

[0103] Among them, f df (·) is a gated fusion module, similar to the aforementioned fusion modules, but differing in that it fuses text probability information and optimized text probability information to obtain the final text feature h. t ∈RC'×H×W .

[0104] It is understandable that when using a gated fusion module to fuse the two, the fusion method is similar to the fusion method of the three features mentioned above. That is, it is necessary to determine the fusion weights corresponding to the text probability information and the optimized text probability information respectively, and then perform fusion processing on the two to obtain the text features.

[0105] Furthermore, to gain a clearer understanding of the process of constructing a super-resolution image from the text image to be reconstructed, another implementation of the text image reconstruction method is introduced below. For example... Figure 2 The diagram shows another flowchart of a text image reconstruction method.

[0106] Step S200: Obtain the text image to be reconstructed.

[0107] Step S210: Correct the text image to obtain the corrected text image.

[0108] The correction process for the text image to be reconstructed has been described above and will not be repeated here.

[0109] Step S220: Extract the horizontal, vertical, and channel features of the corrected text image.

[0110] Step S230: The horizontal features, vertical features, and channel features are fused to obtain the image features of the text image.

[0111] Step S240: Determine the text probability information corresponding to the corrected text image.

[0112] Step S250: Based on the contextual relationship between characters within the text region reflected by the text probability information, optimize the text probability information to obtain optimized text probability information.

[0113] Steps S220-250 correspond to steps S100-130. For a detailed explanation, please refer to the previous description. They will not be repeated here.

[0114] Step S260: The text probability information and the optimized text probability information are fused to obtain text features.

[0115] The fused text probability information and the optimized text probability information in step S260 have been described in detail above and will not be repeated here.

[0116] Step S270: Construct a super-resolution image of the text image based on image features and text features.

[0117] In one possible implementation, when constructing a super-resolution image of the text image using image and text features, the image and text features can first be aligned according to their dimensions to ensure they are in the same dimension. Then, the aligned image and text features are concatenated to obtain a joint feature map. Based on this, features can be extracted from the joint feature map horizontally to obtain a preliminary fused feature map; further features are extracted from the preliminary fused feature map to obtain a target fused feature map, thus ensuring that the target fused feature map fully reflects the characteristics of the joint feature map in both the horizontal and vertical directions. Finally, upsampling processing is performed on the target fused feature map to obtain the super-resolution image.

[0118] For example, after concatenating image features and text features to obtain a joint feature map, the joint feature map can be input into a bidirectional recurrent neural network. This recurrent neural network first extracts features horizontally, and then extracts features vertically based on the horizontally extracted features to obtain a target fused feature map. Finally, this target fused feature map is input into an upsampling module, which is used to construct a super-resolution image, thereby obtaining a super-resolution image of the text image.

[0119] For example, constructing a super-resolution image based on image features and text features can be represented by the following formula:

[0120] I SR =f ps (f b ([i lr' ,h t (Formula 8);

[0121] Among them, i lr' h t These are image features and text features, respectively, f b It is a bidirectional long short-term memory artificial neural network LSTM, f ps It is the upsampling module, which mainly consists of a pixel-shuffle module and a CNN model; I SR This refers to a super-resolution image.

[0122] In this application, by further fusing the probabilities of textual information with optimized textual information, more accurate textual features are obtained. Then, the super-resolution image is constructed using both textual and image features. This method emphasizes the extraction of detailed features to make the reconstructed super-resolution image clearer and the accuracy of textual information in the text image higher.

[0123] It is understandable that, prior to text image reconstruction, this application can also train the various models involved in the preceding embodiments. The following describes one implementation of the text image reconstruction method. For example... Figure 3 The diagram illustrates a schematic representation of the implementation principle framework of a text image reconstruction method. Figure 3 The image shows an example of text image reconstruction using a combination of a model or a template. The function of the corresponding modules has already been explained in the preceding introduction to text image reconstruction methods.

[0124] like Figure 3 After obtaining the text image to be reconstructed, the deformation of the text image is first corrected by a spatial deformation network to obtain the corrected text image.

[0125] Then, the feature transformation module is used to extract the enhanced features of the corrected text image. Based on this, the axial feature module is used to extract the horizontal, vertical, and channel features from the enhanced feature map containing the enhanced features. Based on this, the horizontal, vertical, and channel features are input into the fusion module to obtain the image features of the text image.

[0126] Simultaneously, the text recognition module extracts text probability information from the text image, and the text probability information optimization module optimizes the obtained text probability information. Then, the alignment module aligns the dimensions of the text probability information and the optimized text probability information with the dimensions of the image features. Finally, the dimension-aligned text probability information and the dimension-aligned optimized text probability information are fused together to obtain the text features.

[0127] Finally, after aligning and concatenating the image features and text features, the concatenated joint feature map is input into a bidirectional recurrent neural network to obtain the target fusion feature map. Then, the target fusion feature map is input into the upsampling module for upsampling processing to obtain the constructed super-resolution image.

[0128] It is understandable that when using Figure 3 The modules or network models mentioned above need to be trained continuously beforehand.

[0129] Based on this, this application can also obtain multiple text image samples, each of which is labeled with an actual high-resolution image. This high-resolution image can be a pre-obtained image with a resolution that meets the requirements.

[0130] For example, multiple high-resolution images with sufficient resolution and high image quality can be obtained in advance. Based on these, text image samples with relatively lower resolution can be obtained by downsampling the high-resolution images. Alternatively, text image samples can be obtained first, and then a super-resolution image corresponding to the text image samples can be constructed as the high-resolution image using some more traditional methods or manual methods. Of course, in this application, there are no restrictions on the specific method of obtaining text image samples labeled with high-resolution images.

[0131] After obtaining multiple text image samples, the text image samples can be used to analyze... Figure 3 The various modules or network models involved are trained. During training, the loss function can be used to continuously optimize each module or network model until the training conditions are met.

[0132] For example, for each text image sample, it can be done according to... Figure 3 For example, the text image is sequentially processed... Figure 3 The processing of each module or network model in the process, thereby obtaining a result based on Figure 3 The model reconstructs super-resolution images from text image samples.

[0133] Based on this, for each text image sample, the histogram of oriented gradients (HOGs) is calculated using the histogram of oriented gradients (HOGs) operator for both the high-resolution and super-resolution images corresponding to that text image sample. The minimum absolute value deviation of the HOGs between the high-resolution and super-resolution images is then calculated. Accordingly, the loss function value can be determined based on the minimum absolute value deviation of the HOGs between the high-resolution and super-resolution images corresponding to each text image sample.

[0134] Based on this, during the training of each module and network model, the minimum absolute value deviation obtained by the loss function is used to determine whether the training requirements have been met. If the training requirements have not been met, the gradient of the loss function with respect to the parameters of each module and network model can be calculated by the backpropagation algorithm, and the parameters of each module and network model can be updated by the optimization algorithm to minimize the loss function. If the training requirements have been met, the training of each module and network model is terminated.

[0135] In this application, during the training of each module and network model, a loss function is used to reduce the gap between high-resolution images and super-resolution images, so that the reconstructed super-resolution images are less different from the real text images, and thus the reconstructed super-resolution images are clearer.

[0136] The text image reconstruction apparatus provided in the embodiments of this application will be described below. The text image reconstruction apparatus described below can be referred to in correspondence with the text image reconstruction method described above.

[0137] like Figure 4 This application introduces the text image reconstruction apparatus. For example... Figure 4 As shown, the device may include:

[0138] The feature extraction unit 400 is used to extract the horizontal features, vertical features, and channel features of the text image to be reconstructed;

[0139] The feature fusion unit 410 is used to fuse horizontal features, vertical features and channel features to obtain the image features of the text image;

[0140] The information determination unit 420 is used to determine the text probability information corresponding to the text image. The text probability information represents at least one text region in the text image and the probability distribution of each character contained in the text region.

[0141] The information optimization unit 430 is used to optimize the text probability information based on the contextual relationship between each character in the text region reflected by the text probability information, and obtain the final text probability information.

[0142] Image reconstruction unit 440 is used to construct a super-resolution image of the text image based on image features and optimized text probability information.

[0143] Optionally, the image reconstruction unit includes:

[0144] The information fusion subunit is used to fuse text probability information and optimized text probability information to obtain text features of the text image.

[0145] The image reconstruction subunit is used to construct a super-resolution image of a text image based on image features and text features.

[0146] Optionally, the feature extraction unit includes:

[0147] The feature extraction subunit is used to extract features from the text image to be reconstructed, and obtain the basic feature map;

[0148] The convolution processing subunit is used to perform convolution processing on the basic feature map to obtain the enhanced feature map;

[0149] The dimensional feature extraction subunit is used to extract the horizontal, vertical, and channel features of the enhanced feature map, respectively.

[0150] Optionally, the text image reconstruction apparatus of this application further includes:

[0151] The correction unit is used to perform correction processing on the text image to be reconstructed.

[0152] Optionally, the information fusion unit includes:

[0153] The first unit, weight determination, is used to determine the fusion weights corresponding to the text probability information and the optimized text probability information, respectively.

[0154] The first feature fusion unit is used to fuse the text probability information and the optimized text probability information based on their respective fusion weights to obtain the text features of the text image.

[0155] Optionally, the image reconstruction subunit includes:

[0156] The stitching processing unit is used to align image features and text features according to dimensions, and then stitch the aligned image features and text features together to obtain a joint feature map;

[0157] The lateral feature extraction unit is used to extract lateral features from the joint feature map to obtain a preliminary fused feature map;

[0158] The vertical feature extraction unit is used to extract the vertical features of the preliminary fused feature map to obtain the target fused feature map;

[0159] The upsampling processing unit is used to upsample the target fusion feature map to obtain a super-resolution image.

[0160] Optionally, the feature fusion unit includes:

[0161] The second unit for weight determination is used to determine the fusion weights corresponding to horizontal features, vertical features, and channel features, respectively.

[0162] The second feature fusion unit is used to fuse horizontal, vertical and channel features based on the fusion weights corresponding to the horizontal features, vertical features and channel features to obtain image features.

[0163] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. Furthermore, the features described in the various embodiments of this specification can be substituted or combined with each other, enabling those skilled in the art to implement or use this application. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0164] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0165] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0166] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A text image reconstruction method, characterized in that, include: Extract the horizontal, vertical, and channel features of the text image to be reconstructed; The horizontal, vertical, and channel features are fused together to obtain the image features of the text image; Determine the text probability information corresponding to the text image, wherein the text probability information characterizes at least one text region in the text image and the probability distribution of each character contained within the text region; Based on the contextual relationship between characters within the text region reflected by the text probability information, the text probability information is optimized to obtain optimized text probability information; The text probability information and the optimized text probability information are fused to obtain the text features of the text image; Based on the image features and the text features, a super-resolution image of the text image is constructed; The extraction of horizontal, vertical, and channel features from the text image to be reconstructed includes: Feature extraction is performed on the text image to be reconstructed to obtain a basic feature map; The base feature map is convolved to obtain the enhanced feature map; Extract the horizontal, vertical, and channel features of the enhanced feature map respectively; The step of constructing a super-resolution image of the text image based on the image features and the text features includes: The image features and text features are aligned according to their dimensions, and the aligned image features and text features are concatenated to obtain a joint feature map; Extract the lateral features from the joint feature map to obtain a preliminary fused feature map; Extract the vertical features from the preliminary fused feature map to obtain the target fused feature map; The target fusion feature map is upsampled to obtain a super-resolution image.

2. The method according to claim 1, characterized in that, Before extracting the horizontal, vertical, and channel features of the text image to be reconstructed, the following steps are also included: The reconstructed text image is then corrected.

3. The method according to claim 1, characterized in that, The process of fusing the text probability information and the optimized text probability information to obtain the text features of the text image includes: Determine the fusion weights corresponding to the text probability information and the optimized text probability information respectively; Based on the fusion weights corresponding to the text probability information and the optimized text probability information, the text probability information and the optimized text probability information are fused to obtain the text features of the text image.

4. The method according to claim 1, characterized in that, The process of fusing the horizontal features, vertical features, and channel features to obtain the image features of the text image includes: Determine the fusion weights corresponding to the horizontal features, vertical features, and channel features respectively; Based on the fusion weights corresponding to the horizontal, vertical, and channel features, the horizontal, vertical, and channel features are fused to obtain image features.

5. A text image reconstruction apparatus, characterized in that, include: The feature extraction unit is used to extract the horizontal, vertical, and channel features of the text image to be reconstructed. The feature fusion unit is used to fuse the horizontal features, vertical features, and channel features to obtain the image features of the text image; An information determination unit is used to determine the text probability information corresponding to the text image, wherein the text probability information characterizes at least one text region in the text image and the probability distribution of each character contained within the text region; The information optimization unit is used to optimize the text probability information based on the contextual relationship between each character in the text region reflected by the text probability information, so as to obtain optimized text probability information; An image reconstruction unit is used to construct a super-resolution image of the text image based on the image features and the optimized text probability information; The image reconstruction unit includes: The information fusion subunit is used to fuse the text probability information and the optimized text probability information to obtain the text features of the text image. An image reconstruction subunit is used to construct a super-resolution image of the text image based on the image features and the text features; The image reconstruction subunit includes: The stitching processing unit is used to align image features and text features according to dimensions, and then stitch the aligned image features and text features together to obtain a joint feature map; The lateral feature extraction unit is used to extract lateral features from the joint feature map to obtain a preliminary fused feature map; The vertical feature extraction unit is used to extract the vertical features of the preliminary fused feature map to obtain the target fused feature map; The upsampling processing unit is used to upsample the target fusion feature map to obtain a super-resolution image; The feature extraction unit includes: The feature extraction subunit is used to extract features from the text image to be reconstructed to obtain a basic feature map; The convolution processing subunit is used to perform convolution processing on the basic feature map to obtain the enhanced feature map; The dimensional feature extraction subunit is used to extract the horizontal, vertical, and channel features of the enhanced feature map, respectively.