Image preprocessing method, system and equipment based on OCR (Optical Character Recognition), medium and product
By standardizing and enhancing the OCR images, combining image diffusion and feature fusion of OCR and TDSR models, the problem of inaccurate text extraction in low-quality pictures is solved, and the recognition accuracy and stability of OCR are improved.
Patent Information
- Application Number
- CN202510702890.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Existing OCR technology is difficult to accurately extract text information in low-quality pictures, mainly due to the inappropriate image size, jitter, and exposure problems, resulting in blurred text edges and difficult character features, which affects the recognition accuracy.
By standardizing OCR images, rotation correction, resolution enhancement, color space conversion and morphological optimization, combining OCR model and TDSR model for image diffusion and text recognition, image feature vectors that conform to the text image structure, and feature fusion is performed to improve the super-resolution of text images.
It improves the accuracy and model stability of OCR text extraction under low-quality pictures, especially when stroke adhesion and character features are difficult to extract, so it can accurately identify and extract extracted text.
Smart Images

Figure CN120236285A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to an OCR-based image preprocessing method, system, device, medium, and product. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] A high-quality input image is the basis and key for OCR to correctly extract text information in the image. Currently, most OCR methods simply adjust the size of the image to a size suitable for the input of the model. However, the client neither knows the input format requirements of OCR nor can simply adjust the image to a state suitable for OCR model processing. Therefore, in most cases, the images input to OCR are not suitable for direct calculation by the OCR model, which results in the OCR model being unable to accurately extract text information from many images. Therefore, it is necessary to preprocess the input image to conform to the OCR input format.
[0004] Currently, due to factors such as shooting jitter, overexposure / underexposure, and poor scanning quality, the input images to OCR have blurred text edges and stroke adhesion, making it difficult to extract character features, which directly affects the recognition accuracy. Summary of the Invention
[0005] In order to solve the technical problems existing in the above background art, the present invention provides an OCR-based image preprocessing method, system, device, medium, and product. After standardizing and image enhancing the OCR image, the present invention uses an OCR model to extract text blocks. If the text block is larger than a set threshold, the processed OCR image is output; otherwise, based on the fused image, a TDSR model is used to obtain the processed OCR image. For the situation of stroke adhesion and difficult character feature extraction, it can accurately recognize and extract text, improve the quality of the OCR input image, increase the text extraction accuracy in the case of low input image resolution and incorrect orientation, and improve the performance and stability of the OCR model in applications.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions: The first aspect of the present invention provides an OCR-based image preprocessing method.
[0007] An OCR-based image preprocessing method includes: Obtain an OCR image to be processed, perform standardization, and perform rotation correction on the standardized OCR image; Perform resolution enhancement, color space conversion, edge blurring, and morphological optimization on the rotation-corrected OCR image in sequence, and then perform weighted fusion with the resolution-enhanced image and the color space-converted image to obtain a fused image; Based on the fused image, use the OCR model to extract text blocks. If the text block is larger than the set threshold, output the processed OCR image; otherwise, based on the fused image, use the TDSR model to perform image diffusion, text recognition, and fusion processing; among them, image diffusion is performed under the condition of text to generate image feature vectors that conform to the text image structure; text recognition is conditional on the image feature vectors of image diffusion to obtain text features; fusion processing fuses the image feature vectors and text features to enhance the super-resolution of the text image and obtain the processed OCR image.
[0008] Further, perform standardization; the method includes: calculating the mean and standard deviation of the three channels of the OCR image to be processed to perform standardization processing on the three channels.
[0009] Further, perform rotation correction on the standardized OCR image; the method includes: inputting the standardized OCR image into a trained image rotation angle classifier to obtain the rotation angle of the OCR image; performing rotation correction on the standardized OCR image based on the rotation angle of the OCR image.
[0010] Further, perform resolution enhancement, color space conversion, edge blurring, and morphological optimization on the rotation-corrected OCR image in sequence, and then perform weighted fusion with the resolution-enhanced image and the color space-converted image to obtain a fused image; the method includes: Use the bicubic interpolation algorithm to perform resolution enhancement on the rotation-corrected OCR image to obtain a first enhanced image; Perform color space conversion on the first enhanced image to obtain a second enhanced image; Set the upper and lower limits of the threshold, perform edge blurring feature selection and enhancement on the second enhanced image to obtain a third enhanced image; Perform morphological optimization on the edge region in the third enhanced image to obtain a fourth enhanced image; Perform weighted fusion on the first enhanced image, the second enhanced image, and the fourth enhanced image to obtain a fused image.
[0011] Further, use the TDSR model, the method includes: Perform text encoding and variational auto-encoding on the fused image respectively to obtain a first text feature and a first image feature; Encode the first text feature to obtain a second text feature; Perform cross-attention calculation on the second text feature and the first image feature to obtain the first image feature vector; Based on the first image feature vector, perform multi-level decoding on the first text feature to obtain the text feature of each layer; Encode the text feature of the previous layer, and then perform cross-attention calculation with the image feature vector of the previous layer until the penultimate layer is processed to obtain the final image feature vector, which is decoded to obtain the processed OCR image.
[0012] Further, in the process of training the TDSR model, loss functions for the image diffusion and text recognition processes are respectively constructed, weighted and fused to construct the total loss function to optimize the parameters in the TDSR model training process.
[0013] The second aspect of the present invention provides an OCR-based image preprocessing system.
[0014] An OCR-based image preprocessing system includes: A normalization and rotation module configured to: obtain the OCR image to be processed, perform normalization, and perform rotation correction on the normalized OCR image; An image enhancement module configured to: sequentially perform resolution enhancement, color space conversion, edge blurring processing, and morphological optimization on the rotation-corrected OCR image, and then perform weighted fusion with the resolution-enhanced image and the color space-converted image to obtain a fused image; A threshold judgment and super-resolution processing module configured to: based on the fused image, extract text blocks using an OCR model. If the text block is greater than a set threshold, output the processed OCR image; otherwise, based on the fused image, use the TDSR model to perform image diffusion, text recognition, and fusion processing; wherein, the image diffusion is performed under the condition of the text to generate an image feature vector conforming to the text image structure; the text recognition is conditional on the image feature vector of the image diffusion to obtain the text feature; the fusion processing fuses the image feature vector and the text feature to enhance the super-resolution of the text image to obtain the processed OCR image.
[0015] The third aspect of the present invention provides a computer device, which includes: A processor suitable for executing a computer program; A computer-readable storage medium storing a computer program, which when executed by the processor, implements the steps in the OCR-based image preprocessing method described in the first aspect above.
[0016] The fourth aspect of the present invention provides a computer-readable storage medium storing a computer program, which is adapted to be loaded and executed by a processor to perform the steps in the OCR-based image preprocessing method as described in the first aspect above.
[0017] The fifth aspect of the present invention provides a computer program product or a computer program.
[0018] The present invention provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in the OCR-based image preprocessing method as described in the first aspect above.
[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention provides an OCR-based image preprocessing method, system, device, medium and product. The OCR-based image preprocessing method includes: obtaining an OCR image to be processed and performing normalization, and performing rotation correction on the normalized OCR image; sequentially performing resolution enhancement, color space conversion, edge blurring processing and morphological optimization on the rotation-corrected OCR image, and then performing weighted fusion with the resolution-enhanced image and the color space-converted image to obtain a fused image; based on the fused image, using an OCR model to extract text blocks, and if the text block is larger than a set threshold, outputting the processed OCR image; otherwise, based on the fused image, using a TDSR model to perform image diffusion, text recognition and fusion processing; wherein, the image diffusion is performed under the condition of text to generate an image feature vector conforming to the text image structure; the text recognition is conditional on the image feature vector of the image diffusion to obtain text features; the fusion processing fuses the image feature vector with the text features to enhance the super-resolution of the text image and obtain the processed OCR image. The present invention adopts a TDSR model, and through the fusion process of image diffusion, text recognition and cross-attention calculation, can accurately extract text in the case of stroke adhesion and difficult character feature extraction, and improves the accuracy of text extraction.
[0020] The present invention effectively solves the problem of low accuracy of text information extraction by OCR in low-quality pictures. According to the calculation characteristics of OCR, a picture preprocessing method is proposed. Among them, a picture direction classifier and a corrector are proposed based on the Transformer-Resnet model, and a size preprocessing method is proposed based on the picture characteristics, which improves the quality of the OCR input picture and effectively improves the accuracy of OCR text information extraction. Description of the Drawings
[0021] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments and descriptions thereof of the invention are used to explain the invention and do not unduly limit the invention.
[0022] Figure 1 is a flowchart of the OCR-based image preprocessing method shown in an embodiment of the present invention; Figure 2 is a specific flowchart of the OCR-based image preprocessing method shown in an embodiment of the present invention; Figure 3 is a structural diagram of the picture orientation classification model shown in an embodiment of the present invention; Figure 4 is a flowchart of the low-resolution feature enhancement shown in an embodiment of the present invention; Figure 5 is a flowchart of the OCR preprocessing bicubic interpolation calculation shown in an embodiment of the present invention; Figure 6(a) is an example diagram of the character edge before low-resolution feature enhancement shown in an embodiment of the present invention; Figure 6(b) is an example diagram of the character edge after low-resolution feature enhancement shown in an embodiment of the present invention; Figure 7(a) is an example diagram before the feature enhancement fusion shown in an embodiment of the present invention; Figure 7(b) is an example diagram after the feature enhancement fusion shown in an embodiment of the present invention; Figure 8 is a structural diagram of the TDSR model shown in an embodiment of the present invention; Figure 9 is a schematic diagram of the text feature controlled diffusion process shown in an embodiment of the present invention; Figure 10 is a schematic diagram of the diffusion process image features affecting text recognition shown in an embodiment of the present invention; Figure 11(a) is an example diagram of the low-resolution text block image shown in an embodiment of the present invention; Figure 11(b) is an example diagram of the TDSR calculation result shown in an embodiment of the present invention; Figure 11(c) is an example diagram of the high-definition original image shown in an embodiment of the present invention; Figure 12 is a structural diagram of the OCR-based image preprocessing system shown in an embodiment of the present invention; Figure 13 is a structural diagram of the computer device shown in an embodiment of the present invention. Detailed implementation manners
[0023] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0024] It should be noted that the following detailed description is illustrative and aims to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention pertains.
[0025] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0026] To facilitate understanding of the technical solution of the present invention, some technical terms related to the present invention will be introduced below.
[0027] Bicubic Interpolation is a technique for two-dimensional image or data interpolation, which is an extension of Cubic Interpolation in two dimensions. Bicubic Interpolation has higher interpolation accuracy than bilinear interpolation because it considers a larger pixel neighborhood during calculation and generates a smoother result.
[0028] The Text block Diffusion Super-Resolution algorithm is used to restore and extract information from text images with complex strokes and severe blurring in real-world scenarios. This is a super-resolution algorithm that mutually controls the text recognition and image diffusion processes. Text recognition can guide the image diffusion to generate a text image with the correct structure, and the process of reconstructing the text image can also continuously correct the recognition result of the text.
[0029] Based on this, the present invention provides an OCR-based image preprocessing method, system, device, medium, and product. The solution of the present invention will be described in detail through several embodiments below: Figure 1 is a flowchart of the OCR-based image preprocessing method shown in the embodiments of the present invention; referring to Figure 1 , the method includes: Obtain the OCR image to be processed, perform normalization, and perform rotation correction on the normalized OCR image; Perform resolution enhancement, color space conversion, edge blurring processing, and morphological optimization on the rotation-corrected OCR image in sequence, and then perform weighted fusion with the resolution-enhanced image and the color space-converted image to obtain a fused image; Based on the fused image, use the OCR model to extract text blocks. If the text block is larger than the set threshold, output the processed OCR image; otherwise, based on the fused image, use the TDSR model to perform image diffusion, text recognition, and fusion processing. Among them, image diffusion is carried out under the condition of text to generate image feature vectors that conform to the text image structure. Text recognition is conditional on the image feature vectors of image diffusion to obtain text features. Fusion processing fuses the image feature vectors and text features to enhance the super-resolution of the text image and obtain the processed OCR image.
[0030] After the present invention performs standardization and image enhancement processing on the OCR image, use the OCR model to extract text blocks. If the text block is larger than the set threshold, output the processed OCR image; otherwise, based on the fused image, use the TDSR model to obtain the processed OCR image. For the situation where strokes are adhered and character features are difficult to extract, it can accurately identify and extract text, improve the quality of the OCR input picture, increase the text extraction accuracy rate in the case of low input picture resolution and incorrect orientation, and improve the performance and stability of the OCR model in applications.
[0031] Figure 2 is the specific flowchart of the OCR-based image preprocessing method shown in the embodiments of the present invention; refer to Figure 2 , the method includes: Analyze the OCR image, calculate the mean and standard deviation of each channel of the image, and standardize it; Use the picture direction classification model to correct the picture direction; Perform low-resolution feature extraction and enhancement on the corrected picture; When the text block is too blurred and cannot be compensated by traditional computer vision algorithms, extract the text block; The confidence level of the output text block (that is, the degree of certainty, the lower the confidence level, the more likely to make mistakes) will decrease. At this time, set a reasonable threshold. When the confidence level of the text block is too low, perform text diffusion super-resolution (Text blockDiffusion Super-Resolution, TDSR) calculation on the text block to obtain a clearer text block picture and output the text in the picture.
[0032] In some embodiments, the analyzing the OCR image, calculating the mean and standard deviation of each channel of the image, and standardizing it; the method includes: Collect image files such as business licenses, permits, certificates, documents, etc. commonly seen in daily operations, calculate the mean and standard deviation of the three channels of the image, and then perform standardization.
[0033] The standardization formula is:
[0034] Among them, is the mean value of the RGB values in one channel of the image, and is the standard deviation of the RGB values in this channel.
[0035] After standardization, the distribution of the dataset will be adjusted to a distribution with a mean of 0 and a standard deviation of 1. The standardized image is more suitable for model inference calculation. During the training process of the model, the values of the standardized data are more stable, which is also beneficial to model convergence and reduces the occurrence of gradient explosion.
[0036] In some embodiments, the method of correcting the picture direction by using the picture direction classification model includes: using 4478 OCR images commonly seen in daily business collected, and respectively rotating them by 90 、180 、270 processing and marking the rotation angle as a label, and then building a TransformerEncoder-Resnet model with reduced width as a lightweight picture rotation angle classifier for training. After the trained model inputs the standardized image, it outputs the rotation angle of the image.
[0037] Through a large number of business practices, it is found that most OCR models are very sensitive to the resolution of the picture input. When the input picture resolution is too low or the clarity is too low, the performance of the OCR model is often poor. The reason for the problem is that generally, the OCR model directly enlarges the picture with a small resolution and inputs it into the model. This operation will lose some image details and make the key text features become blurred. Here, the bicubic interpolation algorithm is used to first expand the picture resolution, supplement sufficient picture details, and then reduce the picture size to the size required by the model, ensuring the rich features of the picture input into the model.
[0038] Figure 3 is the structural diagram of the picture direction classification model shown in the embodiments of the present invention; referring to Figure 3 , the picture direction classification model includes: a Resnet-18 model, a Transformer encoder, and a Softmax function. Among them, the Transformer encoder is a part of the Transformer model. Here, the Resnet-18 model with adjusted width is first used to perform convolution and downsampling calculations on the images in the dataset. When the size of the feature map is suitable for Transformer calculation, it is input into the Transformer encoder for calculation. The output result is input into the Softmax function to calculate the classification result, and the cross-entropy loss function is used to calculate the classification loss.
[0039] The proposed lightweight Transformer-Resnet model can quickly and accurately identify the rotation angle of an image, process the image into a format more suitable for OCR input, and improve the accuracy of subsequent text information extraction.
[0040] The training process of the image orientation classification model is introduced as follows: The preparation of the dataset includes four parts: collection, cleaning, preprocessing, and partitioning. Here, a total of 4,478 images containing invoices, licenses, and documents in business and on the Internet were collected, duplicates and abnormal images containing a large amount of irrelevant information were removed, the images were resized and normalized, and the images were randomly divided into a training dataset and a validation dataset at a ratio of 8:2 after being rotated 90° three times respectively.
[0041] The backbone network uses a simplified version of the Resnet18 model. The original model contains four convolutional modules, and the number of channels of the output feature maps of each convolutional module is 64, 128, 256, and 512 respectively, and the image is downsampled 16 times. Here, in order to ensure the lightweight of the model and avoid overfitting, the number of channels is reduced to one-fourth of the original, and the downsampling factor is adjusted to 8 to retain more image details. The backbone network receives the input image as input, extracts a low-resolution feature, and encodes the high-level spatial representation of the image, which is a sequence of length × width × number of channels.
[0042] The Transformer-Encoder, that is, the Transformer encoder, has two standard layers, which are connected in sequence. Each layer contains a multi-head attention module and a fully connected layer and adds position encoding to enable the model to receive the position information in the picture. The Transformer-Encoder can learn rich feature representations, which are very suitable for various computer vision tasks such as image classification, object detection, and semantic segmentation. Through multi-layer stacking and the multi-head attention mechanism, the model can extract features from different angles.
[0043] The output of the Transformer-Encoder is input into the softmax module through a fully connected layer to calculate the classification probability, and the cross-entropy loss function is used to calculate the classification loss for optimizing the model parameters by the gradient descent algorithm.
[0044] After receiving the image to be processed, the above image classification model will perform a series of inference calculations. In the inference calculations, there is no update of the model parameters. The model only needs to calculate the image according to the trained parameters and weights. After the calculation, the model will generate the orientation recognition result of the image, which are 0 、90 、180 、270 , the subsequent OpenCV library or PIL library can correct the image orientation by rotating the image according to the angle.
[0045] Experiments have found that in the samples with incorrect OCR results, most of them contain the phenomenon of blurred pictures. Low-quality images will seriously affect the model's understanding of picture information, especially for the image features of text and numbers.
[0046] Figure 4 is the flowchart of low-resolution feature enhancement shown in the embodiments of the present invention; refer to Figure 4 , extract and enhance the low-resolution features of the corrected picture; the method includes: Input the corrected original image, and use the bicubic interpolation algorithm to enhance the resolution of the rotated and corrected OCR image to obtain the first enhanced image; perform color space conversion on the first enhanced image to obtain the second enhanced image; set the upper and lower limits of the threshold, and perform edge blur feature selection and enhancement on the second enhanced image to obtain the third enhanced image; perform morphological optimization on the edge region in the third enhanced image to obtain the fourth enhanced image; perform weighted fusion on the first enhanced image, the second enhanced image, and the fourth enhanced image to obtain the fused image.
[0047] Figure 5 is the flowchart of bicubic interpolation calculation for OCR preprocessing shown in the embodiments of the present invention; refer to Figure 5 , the bicubic interpolation algorithm includes: performing cubic interpolation in the x direction of the OCR image, and performing cubic interpolation on each row among the four rows of pixels around the target point to generate four intermediate interpolation results; performing cubic interpolation on the four difference results in the y direction of the OCR image.
[0048] The bicubic interpolation formula can be expressed as:
[0049] where is the newly interpolated pixel value, are the 16 pixel values within the neighborhood, is the weight function, which determines the weight based on the distance and usually uses a cubic function.
[0050] The picture preprocessing method for improving OCR recognition accuracy proposed by the present invention can effectively improve the accuracy and stability of OCR in different situations by using deep learning methods and the bicubic interpolation algorithm.
[0051] Regarding the Lab color space conversion, there are many cases of black text on a white background in texts and document images. The Lab color space significantly improves the robustness and accuracy of color-related tasks (such as segmentation, detection, and enhancement) by separating luminance from color and providing a perceptually uniform color gamut, and has greater advantages in cases where color differences are obvious.
[0052] Regarding the threshold segmentation of text digital features: By calculating the threshold mean of the dataset, a threshold upper limit and a threshold lower limit are set, and the images in the Lab color space are selected and enhanced for edge blur features according to the upper and lower limits. Figure 6(a) is an example diagram of the character edge low-resolution feature before enhancement shown in the embodiment of the present invention; Figure 6(b) is an example diagram of the character edge low-resolution feature after enhancement shown in the embodiment of the present invention; Edge feature enhancement can be achieved by enhancing the luminance channel of the selected area.
[0053] The method for enhancing the low-resolution features of character edges proposed by the present invention greatly improves the text area quality of OCR images and enhances the performance of OCR in processing variable-resolution images.
[0054] After optimizing the edge area through morphology, it is possible to eliminate small parts, shrink the boundary, and connect breaks. Then, the original image after interpolation super-resolution, the Lab channel image, and the enhanced character edge features are weighted and added together to achieve the enhancement and fusion of specific features. Figure 7(a) is an example diagram before the feature enhancement and fusion shown in the embodiment of the present invention; Figure 7(b) is an example diagram after the feature enhancement and fusion shown in the embodiment of the present invention.
[0055] In some embodiments, the confidence level of the calculated output text block (i.e., the degree of certainty, the lower the confidence level, the more likely it is to be incorrect) will decrease. At this time, a reasonable threshold is set. When the confidence level of the text block is too low, text diffusion super-resolution (TDSR) calculation is performed on the text block to obtain a clearer text block image and the text in the output image; The method includes: After the low-resolution feature enhancement, the first-pass OCR calculation is performed, and the position and confidence level of each text block are obtained. To ensure the correctness of the result, the present invention proposes a text block diffusion super-resolution algorithm (TDSR) to perform TDSR calculation on text blocks with a confidence level lower than the threshold to obtain a clear text block image and the text information therein.
[0056] Figure 8 is the structural diagram of the TDSR model shown in the embodiment of the present invention; Refer to Figure 8, the fused image is encoded using a text encoder and a variational autoencoder respectively to obtain the first text feature and the first image feature; the first text feature is encoded using a self-attention encoder (Transformer-Encoder) to obtain the second text feature; a diffusion model (DiffBlock) is used to perform cross-attention calculation on the second text feature and the first image feature to obtain the first image feature vector; based on the first image feature vector, a self-attention decoder (Transformer-Decoder) is used to perform multi-level decoding on the first text feature to obtain the text feature of each layer; the text feature of the previous layer is encoded using a self-attention encoder and then cross-attention calculation is performed with the image feature vector of the previous layer until the penultimate layer is processed to obtain the final image feature vector, which is decoded to obtain the processed OCR image.
[0057] Specifically, in the TDSR model, it includes three parts: an image diffusion process, a text recognition process, and a fusion process; the image diffusion process is carried out under the condition of text to generate high-quality pictures that conform to the text image structure; the text recognition process is conditional on the feature vector of image diffusion to achieve more accurate text recognition and correction; the output of the self-attention encoder and the output of each DiffBlock encode and fuse the text and image features in the previous step, and finally achieve high-fidelity and high-realistic text image super-resolution through the cooperation of the three parts.
[0058] The above process can be expressed by the following formula:
[0059]
[0060]
[0061]
[0062]
[0063]
[0064] Among them, t represents the t-th time step, represents the control condition required for the diffusion process at the t-th time step (in the text recognition process, this control condition is the control condition provided in the image diffusion process; in the image diffusion process, this control condition is the control condition provided in the text recognition process), represents the text sequence feature generated by the text recognition process at the t-th time step, represents the calculation of the self-attention encoder. represents the condition required for text recognition at the t-th time step, Indicates DiffBlock calculation Indicates the features extracted by the variational autoencoder from the low-resolution image Indicates the sampling in the Gaussian distribution at the t-th time step Indicates the result of the image diffusion process at the t-th time step Indicates the calculation of the U-shaped network structure of DiffBlock Indicates the condition for controlling the image diffusion process at the t-th time step Indicates the noise schedule that can be artificially controlled during the image diffusion process, i.e., the noise addition process Indicates the prediction result of text recognition at the t-th time step Indicates the calculation of the self-attention decoder
[0065] The image diffusion process uses the Latent Diffusion Model architecture. It should be noted that during the image diffusion process, the diffusion process needs to be controlled by text information, and the control is achieved through the text feature vector output during the recognition process. The diffusion process receives an input image of 512×128 (long text). In the variational autoencoder, f = 4 is set, and the input image is downsampled by four times while the number of channels of the feature map is expanded to 320. That is, after passing through the variational autoencoder, the input image of 512×128×3 is transformed into a feature vector of 128×32×320; then it is a standard diffusion model, where each DiffBlock is a Unet structure in a standard diffusion process.
[0066] The text block diffusion super-resolution algorithm proposed by the present invention can incorporate a specific text image structure into the image diffusion super-resolution process to achieve the image reconstruction of complex structure numbers and letters.
[0067] Figure 9 Is a schematic diagram showing the text feature control diffusion process in the embodiment of the present invention; referring to Figure 9 , the image feature vector is input through the Encoder structure of the Transformer, the text feature vector is input through the Decoder structure of the Traansformer, and finally, after cross-attention calculation, an image feature vector containing text features is output, downsampled, and then continues the standard diffusion process calculation in the Unet.
[0068] During training, n = 200 is used, that is, it contains 200 time steps. The number of time steps determines the discrete number of time steps used in the forward process (from data to noise) and the reverse process (from noise to recover data). Each time step is associated with a specific variance value, and these variance values define how much noise to add or remove in each step.
[0069] During the text recognition process, the Text Encoder uses an image encoder with a self-attention encoder structure to process a 512×128×3 image into a 128×32×320 feature vector, and inputs it into the Encoder and Decoder respectively. The encoder encodes it into the control vector received by the DiffBlock during the diffusion process, and the decoder learns the text features therein so that the text information in the image can be finally output according to the image features.
[0070] During the text recognition process, the self-attention decoder receives the image features generated during the diffusion process to enable mutual communication between the image denoising and text recognition processes, so that the denoised image retains the structural information of the text, and the processed image will be more reasonable and clear.
[0071] Figure 10 It is a schematic diagram showing the influence of the diffusion process image features on text recognition in the embodiments of the present invention; referring to Figure 10 , after the image feature vector and the diffusion process image feature vector are calculated by the self-attention decoder at the last time step, an identification vector corresponding to each character is generated. The final self-attention decoder can also be regarded as the Text Decoder, corresponding to the Text Encoder at the beginning of the process.
[0072] During training: The VAE loads the pre-trained weights on Open-Image, and the weights of the other modules in the diffusion process and the text recognition process are all trained from scratch. It is trained for 100,000 iter on the CTR dataset, where batch_size = 16. This model calculates the low-confidence and low-resolution text blocks passed in from the previous step. The size of the low-resolution image block input to the model needs to be adjusted to 128×512, and it can only process single-line text blocks.
[0073] The training includes three parts. First, train the self-attention decoder in the text recognition process. At this time, the loss function is the KL divergence between the predicted text and the prior of the actual text in the dataset. The second part is to train the image diffusion model, where the loss function includes the L2 loss and the OCR loss. The L2 loss is the loss function of the traditional diffusion model that enables the model to have the noise estimation ability, and the OCR loss is the loss of the text extraction result after connecting the super-resolved image to the OCR model. The weights of the diffusion model are updated by freezing the weights of the OCR model, and the weight parameters are used to limit the proportion of the OCR loss. Finally, train the TDSR model as a whole. At this time, it is necessary to freeze the weights of the above two parts and train the weights of the self-attention encoder part. The formula is as follows:
[0074]
[0075]
[0076] Among them, represents the image diffusion loss, represents the L2 loss, represents the OCR loss, represents the weight parameter, represents the text recognition loss, represents the th string result predicted at the th time step, represents the string feature vector generated at the th time step, represents the initial string feature vector, represents the probability distribution of two strings, represents the total loss,
[0077] FIG. 11(a) is an example diagram of a low-resolution text block image shown in an embodiment of the present invention; FIG. 11(b) is an example diagram of the TDSR calculation result shown in an embodiment of the present invention; FIG. 11(c) is an example diagram of a high-definition original image shown in an embodiment of the present invention; the text block after being processed by the TDSR model is shown in FIG. 11(b). Here, a sample picture with a low OCR text information extraction accuracy rate is selected for testing, and it can be seen that there is a large difference between the results with and without preprocessing.
[0078] The above Figure 1 has introduced in detail the OCR-based image preprocessing method provided by the embodiments of the present invention. Next, the OCR-based image preprocessing system provided by the embodiments of the present invention will be introduced with reference to the accompanying drawings.
[0079] Figure 12 is a schematic structural diagram of an OCR-based image preprocessing system shown in an embodiment of the present invention. Referring to Figure 12 , the system described in the present invention includes: A normalization and rotation module, which is configured to: obtain an OCR image to be processed, perform normalization, and perform rotation correction on the normalized OCR image; An image enhancement module, which is configured to: sequentially perform resolution enhancement, color space conversion, edge blurring processing, and morphological optimization on the rotation-corrected OCR image, and then perform weighted fusion with the resolution-enhanced image and the color space-converted image to obtain a fused image; The threshold judgment and super-resolution processing module is configured to: extract text blocks using an OCR model based on the fused image. If the text block is larger than the set threshold, output the processed OCR image; otherwise, based on the fused image, use the TDSR model to perform image diffusion, text recognition, and fusion processing. Among them, image diffusion is performed under the condition of text to generate image feature vectors that conform to the text image structure; text recognition is conditional on the image feature vectors of image diffusion to obtain text features; fusion processing fuses the image feature vectors and text features to enhance the super-resolution of the text image and obtain the processed OCR image.
[0080] In some embodiments, the normalization and rotation module is specifically configured to: calculate the mean and standard deviation of the three channels of the OCR image to be processed to perform normalization processing on the three channels.
[0081] Input the normalized OCR image into the trained image rotation angle classifier to obtain the rotation angle of the OCR image; perform rotation correction on the normalized OCR image based on the rotation angle of the OCR image.
[0082] In some embodiments, the image enhancement module is specifically configured to: use the bicubic interpolation algorithm to enhance the resolution of the rotation-corrected OCR image to obtain the first enhanced image; perform color space conversion on the first enhanced image to obtain the second enhanced image; set the upper and lower limits of the threshold to perform edge blur feature selection and enhancement on the second enhanced image to obtain the third enhanced image; perform morphological optimization on the edge region in the third enhanced image to obtain the fourth enhanced image; perform weighted fusion on the first enhanced image, the second enhanced image, and the fourth enhanced image to obtain the fused image.
[0083] In some embodiments, the threshold judgment and super-resolution processing module is specifically configured to: perform text encoding and variational auto-encoding on the fused image respectively to obtain the first text feature and the first image feature; encode the first text feature to obtain the second text feature; perform cross-attention calculation on the second text feature and the first image feature to obtain the first image feature vector; based on the first image feature vector, perform multi-level decoding on the first text feature to obtain the text features of each layer; encode the text features of the previous layer and perform cross-attention calculation with the image feature vector of the previous layer until the penultimate layer is processed to obtain the final image feature vector, and after decoding, obtain the processed OCR image.
[0084] In some embodiments, during the training of the TDSR model, loss functions for the image diffusion and text recognition processes are constructed respectively, and weighted fusion is performed to construct the total loss function to optimize the parameters in the training process of the TDSR model.
[0085] According to an embodiment of the present invention, the OCR-based image preprocessing system may correspond to executing the methods described in the embodiments of the present invention, and the above and other operations and / or functions of each module of the OCR-based image preprocessing system are respectively for implementing Figure 1 the corresponding processes of each method in
[0086] See Figure 13 the structural diagram of the computer device shown in
[0087] This computer device includes a processor, a communication interface, and a computer-readable storage medium. Among them, the processor, the communication interface, and the computer-readable storage medium can be connected through a bus or other means. Among them, the communication interface is used to receive and send data. The computer-readable storage medium can be stored in the memory of the computer device. The computer-readable storage medium is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer-readable storage medium. The processor (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the computer device, and is adapted to implement one or more instructions. Specifically, it is adapted to load and execute one or more instructions to implement the corresponding steps in the embodiment of the OCR-based image preprocessing method.
[0087] This embodiment provides a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the processing system of the computer device. Moreover, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.
[0088] In one embodiment, one or more instructions are stored in the computer-readable storage medium; the processor loads and executes one or more instructions stored in the computer-readable storage medium to implement the corresponding steps in the embodiment of the above OCR-based image preprocessing method.
[0089] This embodiment provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the corresponding steps in the above-mentioned embodiment of the OCR-based image preprocessing method.
[0090] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.
[0091] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0092] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the specified functions in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0093] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0094] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0095] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An OCR-based image preprocessing method, characterized in that Including: Obtain the OCR image to be processed, perform normalization, and perform rotation correction on the normalized OCR image; Successively perform resolution enhancement, color space conversion, edge blurring processing, and morphological optimization on the rotation-corrected OCR image, and then perform weighted fusion with the resolution-enhanced image and the color space-converted image to obtain a fused image; Based on the fused image, use the OCR model to extract text blocks. If the text block is larger than the set threshold, output the processed OCR image; Otherwise, based on the fused image, use the TDSR model to perform image diffusion, text recognition, and fusion processing; among them, image diffusion is performed under the condition of text to generate an image feature vector that conforms to the text image structure; text recognition is based on the image feature vector of image diffusion to obtain text features; fusion processing fuses the image feature vector with the text features to enhance the super-resolution of the text image and obtain the processed OCR image.
2. The OCR-based image preprocessing method according to claim 1, wherein The said performing normalization; the method includes: calculating the mean and standard deviation of the three channels of the OCR image to be processed to perform normalization processing on the three channels.
3. The OCR-based image preprocessing method according to claim 1, wherein The said performing rotation correction on the normalized OCR image; the method includes: inputting the normalized OCR image into the trained image rotation angle classifier to obtain the rotation angle of the OCR image; performing rotation correction on the normalized OCR image based on the rotation angle of the OCR image.
4. The OCR-based image preprocessing method according to claim 1, wherein, The said successively performing resolution enhancement, color space conversion, edge blurring processing, and morphological optimization on the rotation-corrected OCR image, and then performing weighted fusion with the resolution-enhanced image and the color space-converted image to obtain a fused image; the method includes: Adopt the bicubic interpolation algorithm to perform resolution enhancement on the rotation-corrected OCR image to obtain the first enhanced image; Perform color space conversion on the first enhanced image to obtain the second enhanced image; Set the upper and lower limits of the threshold, perform edge blurring feature selection and enhancement on the second enhanced image to obtain the third enhanced image; Perform morphological optimization on the edge region in the third enhanced image to obtain the fourth enhanced image; Perform weighted fusion on the first enhanced image, the second enhanced image, and the fourth enhanced image to obtain a fused image.
5. The OCR-based image preprocessing method according to claim 1, characterized in that The said using the TDSR model; the method includes: Respectively perform text encoding and variational auto-encoding on the fused image to obtain the first text feature and the first image feature; Encode the first text feature to obtain the second text feature; Perform cross-attention calculation on the second text feature and the first image feature to obtain the first image feature vector; Based on the first image feature vector, perform multi-level decoding on the first text feature to obtain the text features of each layer; Encode the text feature of the previous layer, and then perform cross-attention calculation with the image feature vector of the previous layer until the penultimate layer is processed to completion, obtain the final image feature vector, and after decoding, obtain the processed OCR image.
6. The OCR-based image preprocessing method according to claim 1, wherein, In the process of training the TDSR model, respectively construct the loss functions of the image diffusion and text recognition processes, and perform weighted fusion to construct the total loss function to optimize the parameters in the training process of the TDSR model.
7. An OCR-based image preprocessing system, characterized in that, Including: A normalization and rotation module, which is configured to: obtain an OCR image to be processed, perform normalization, and perform rotation correction on the normalized OCR image; An image enhancement module, which is configured to: sequentially perform resolution enhancement, color space conversion, edge blurring processing, and morphological optimization on the rotation-corrected OCR image, and then perform weighted fusion with the resolution-enhanced image and the color space-converted image to obtain a fused image; A threshold judgment and super-resolution processing module, which is configured to: based on the fused image, extract text blocks using an OCR model, and if the text blocks are larger than a set threshold, output the processed OCR image; Otherwise, based on the fused image, use a TDSR model to perform image diffusion, text recognition, and fusion processing; wherein, the image diffusion is performed under the condition of text to generate image feature vectors that conform to the text image structure; the text recognition is conditional on the image feature vectors of the image diffusion to obtain text features; the fusion processing fuses the image feature vectors with the text features to enhance the super-resolution of the text image and obtain the processed OCR image.
8. A computer device, characterized in that a processor, adapted to execute a computer program; a computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the steps in the OCR-based image preprocessing method according to any one of claims 1-6 are implemented.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps in the OCR-based image preprocessing method according to any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by the processor, the steps in the OCR-based image preprocessing method according to any one of claims 1-6 are implemented.
Citation Information
Patent Citations
Bill correction method based on key point positioning for OCR identification
CN111126382A
OCR (optical character recognition) preprocessing method and system for certificate photo shot by mobile phone and storage medium
CN117173712A
Global character image restoration method and device, and medium
CN118154476A
Image restoration method and device, equipment and storage medium
CN118172292A
Scene character image super-resolution method for generative image prior
CN119941509A