OCR Training Data Generation Method, Device, Computer Equipment, and Storage Medium

By generating and enhancing training images with specific character placement and image processing techniques, the method addresses the poor recognition accuracy of OCR models for scanned and fax documents, significantly improving their performance.

CN113361512BActive Publication Date: 2025-07-15WAVEFAX TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110620906.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-03
Publication Date
2025-07-15
Estimated Expiration
2041-07-15

AI Technical Summary

Technical Problem

The existing OCR training image generation methods cannot effectively deal with the missing character strokes and low resolution problems in scanning electronic documents and fax documents, resulting in the low recognition rate of OCR models for identifying these documents.

Method used

By obtaining the training corpus, creating blank images, writing characters according to preset parameter information, image enhancement processing is performed, training images that simulate fax and scanning effects are generated, and training data of the OCR model is optimized.

Benefits of technology

The generated training data can significantly improve the recognition rate of the OCR model for fax and scanned images, and improve the recognition accuracy of the OCR model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113361512B_ABST
    Figure CN113361512B_ABST
Patent Text Reader

Abstract

The present application relates to an OCR training data generation method, apparatus, computer device, and storage medium. The method includes: obtaining training corpus, obtaining image parameter information, establishing a blank image according to the image parameter information, extracting preset strings from the training corpus, writing the characters in the extracted strings into the blank image according to the preset generated character parameter information to generate an initial image, performing image enhancement processing on the initial image to obtain a training image, and generating training data for training an OCR deep learning engine according to the training image. This solution is used to optimize the recognition of fax images and scanned images. The pictures in the generated training data have the characteristics of real fax images or scanned images. When the OCR model trained using this training data is used to recognize fax images and scanned images, the recognition rate is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of OCR recognition, and particularly to an OCR training data generation method, device, computer device, and storage medium. Background Art

[0002] Neural network models currently achieve better accuracy effects than traditional models in most academic disciplines and field applications, and also have good application generalization. The currently commonly used neural network models mainly include three types: convolutional neural network (CNN), recurrent neural network (RNN), and transformer network; at the same time, the developing graph neural network (GNN) also has certain applications in fields such as biology and chemistry.

[0003] Any neural network model cannot do without model training, which requires constructing relevant training data and performing weight training with the neural network model. For the training task of the OCR model, it is necessary to generate text line images of a fixed size and record the text line characters corresponding to each image; input the digital images into the neural network model, and the model makes predictions on the input image data; calculate the predicted text line characters output with the true text line characters to obtain the error value of the model prediction; update the model parameters with the error value.

[0004] The currently commonly used OCR training image generation method is to use the random algorithm to generate characterless content images with random grayscale or color backgrounds. Use digital image processing libraries such as opencv to write character images of random scales into the images. And combine simple algorithms such as Gaussian filtering, affine transformation, thickening, and cropping to add noise to the images to obtain training images.

[0005] The above-mentioned training image generation method is for character recognition in natural scenes, and many noise processing methods are to increase the complexity of the training scenario. However, in the OCR recognition tasks of traditional scanned electronic documents and fax documents, what is faced is the missing of character strokes, the thickening and thinning of strokes due to problems such as exposure, and the low-resolution problem under fax conditions. Therefore, the current training image generation method cannot generate corresponding training data files for this situation, resulting in the low recognition rate of the current OCR model for scanned electronic documents and fax documents. Summary of the Invention

[0006] Based on this, in view of the above technical problems, it is necessary to provide an OCR training data generation method, device, computer device, and storage medium that can improve the accuracy of the OCR recognition engine in recognizing characters in scanned and fax documents.

[0007] An OCR training data generation method, the method includes:

[0008] Obtain a training corpus;

[0009] Obtain image parameter information and create a blank image according to the image parameter information;

[0010] Extract a preset string from the training corpus and write the characters in the extracted string into the blank image according to the preset generated character parameter information to generate an initial image;

[0011] Perform image enhancement processing on the initial image to obtain a training image;

[0012] Generate training data based on the training image.

[0013] In one embodiment, writing the characters in the extracted string into the blank image according to the preset generated character parameter information to generate an initial image includes:

[0014] Obtain the generated character parameter information according to the image parameter information;

[0015] Obtain the generated characters from the string in the training corpus;

[0016] Write the generated characters into the blank image according to the generated character parameter information to obtain an initial image.

[0017] In one embodiment, the image parameter information includes the image width; writing the characters into the blank image according to the generated character parameter information to obtain an initial image includes:

[0018] Starting from the character start horizontal position of the blank image, write the generated characters into the blank image according to the generated character parameter information until the character start horizontal position of the next generated character exceeds the image width, and then stop writing to obtain an initial image.

[0019] In one embodiment, performing image enhancement processing on the initial image to obtain a training image includes:

[0020] Perform random binarization on the initial image;

[0021] Randomly map the black pixels in the binarized initial image to a preset gray value range;

[0022] Perform dithering processing on the initial training image to generate a dot matrix binary image;

[0023] Perform random horizontal or vertical scaling on the dot matrix binary image to obtain a training image.

[0024] In one embodiment, performing image enhancement processing on the initial image to obtain a training image includes:

[0025] Perform random binarization on the initial image;

[0026] Obtain an image of slight morphological dilation of black pixels based on the binarized initial image;

[0027] Copy the initial image to obtain a copy image;

[0028] Perform a vertical erosion operation on the initial image;

[0029] Generate a mapping matrix of the initial image according to the image parameter information;

[0030] For the initial image and the copy image, connect the horizontal strokes in the initial image removed by the erosion operation according to the mapping matrix;

[0031] Perform random bilinear interpolation or cubic interpolation on the initial image horizontally or vertically;

[0032] Randomly binarize the interpolated initial image to obtain a training image.

[0033] In one embodiment, for the initial image and the copy image, connecting the horizontal strokes in the initial image removed by the erosion operation according to the mapping matrix includes:

[0034] If the value of the mapping matrix at a certain pixel point is the first point value, perform an AND operation on the pixel point values of the initial image and the copy image;

[0035] If the value of the mapping matrix at a certain pixel point is the second point value, delete the pixel value corresponding to the pixel point in the initial image.

[0036] In one embodiment, generating training data according to the training image includes:

[0037] Save the training image to a preset path;

[0038] Save the preset path of the training image;

[0039] Save the string text corresponding to the training image;

[0040] Integrate the training image, the preset path, and the string text to obtain training data.

[0041] An OCR training data generation device, the device includes:

[0042] A corpus acquisition module, used to acquire a training corpus;

[0043] An image establishment module, used to acquire image parameter information and establish a blank image according to the image parameter information;

[0044] A character extraction module, which is used to extract a preset string from a training corpus, and write the characters in the extracted string into a blank image according to preset generated character parameter information to generate an initial image;

[0045] An image enhancement module, which performs image enhancement processing on the initial image to obtain a training image;

[0046] A data generation module, which is used to generate training data according to the training image.

[0047] A computer device, including a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented:

[0048] Obtain a training corpus;

[0049] Obtain image parameter information, and establish a blank image according to the image parameter information;

[0050] Extract a preset string from the training corpus, and write the characters in the extracted string into the blank image according to the preset generated character parameter information to generate an initial image;

[0051] Perform image enhancement processing on the initial image to obtain a training image;

[0052] Generate training data according to the training image.

[0053] A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0054] Obtain a training corpus;

[0055] Obtain image parameter information, and establish a blank image according to the image parameter information;

[0056] Extract a preset string from the training corpus, and write the characters in the extracted string into the blank image according to the preset generated character parameter information to generate an initial image;

[0057] Perform image enhancement processing on the initial image to obtain a training image;

[0058] Generate training data according to the training image.

[0059] In the above OCR training data generation method, by obtaining a training corpus, obtaining image parameter information, creating a blank image based on the image parameter information, extracting a preset string from the training corpus, writing the characters in the extracted string into the blank image according to the preset generated character parameter information to generate an initial image, performing image enhancement processing on the initial image to obtain a training image, and generating training data for training an OCR deep learning engine based on the training image. This solution is used to optimize the recognition of fax images and scanned images. The pictures in the generated training data have the characteristics of real fax images or scanned images. When the OCR model trained using this training data performs recognition on fax images and scanned images, the recognition rate is significantly improved. Brief Description of the Drawings

[0060] Figure 1 is a schematic flowchart of the OCR training data generation method in an embodiment;

[0061] Figure 2 is a schematic flowchart of generating an initial image in an embodiment;

[0062] Figure 3 is a schematic flowchart of performing enhancement processing on the initial image in an embodiment;

[0063] Figure 4 is a schematic flowchart of performing fax enhancement processing on the initial image in an embodiment;

[0064] Figure 5 is a schematic flowchart of performing enhancement processing on the initial image in another embodiment;

[0065] Figure 6 is a schematic flowchart of performing scanned image enhancement processing on the initial image in an embodiment;

[0066] Figure 7 is an effect diagram of the initial image in an embodiment;

[0067] Figure 8 In an embodiment, for Figure 7 the effect diagram of the fax training image after performing enhancement processing on the shown image;

[0068] Figure 9 In an embodiment, for Figure 7 the effect diagram of the scanned training image after performing enhancement processing on the shown image;

[0069] Figure 10 is a structural block diagram of an OCR training data generation device in an embodiment;

[0070] Figure 11 is an internal structural diagram of a computer device in an embodiment. Detailed Description of the Embodiment

[0071] In order to make the objectives, technical solutions and advantages of this application more clear and understandable, the following further details this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0072] In one embodiment, as Figure 1 shown, an OCR training data generation method is provided. In this embodiment, taking the application of this method to a server as an example, it can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the solution includes the following steps:

[0073] Step 102, obtain a training corpus.

[0074] Specifically, obtain the training corpus through the Internet. The training corpus stores language materials that have actually appeared in the actual use of the language. The training corpus is a basic resource that carries language knowledge with an electronic computer as the carrier. After obtaining the training corpus, it is usually necessary to preprocess the training corpus.

[0075] Step 104, obtain image parameter information and create a blank image according to the image parameter information.

[0076] Specifically, the image parameter information includes the image height and the image width. A blank image refers to an image that has not yet had characters written into it. A corresponding blank image can be created according to the image height and the image width.

[0077] Step 106, extract a preset string from the training corpus and write the characters in the extracted string into the blank image according to the preset generated character parameter information to generate an initial image.

[0078] Specifically, the generated character parameter information includes the character height and the character width. The character height is the maximum character height within the image height, and the character width determines the number of characters written in each row of the blank image.

[0079] Further, obtain the generated characters from the string in the training corpus, and starting from the character starting horizontal position in the blank image, write the generated characters into the blank image according to the generated character parameter information until the character starting horizontal position of the next generated character exceeds the image width, and then stop writing to obtain the initial image.

[0080] Step 108, perform image enhancement processing on the initial image to obtain a training image.

[0081] Specifically, for the initial image, an image enhancement process is performed. A random image enhancement method is selected for the enhancement process. For example, the initial image is randomly enhanced by fax or by scanning. If the fax enhancement process is performed, a fax training image is generated. If the scanning enhancement process is performed, a scanning training image is generated. The fax training image is used as input during subsequent OCR model training to improve the OCR model's ability to recognize fax images. The scanning training image is also used as input during subsequent OCR model training to improve the OCR model's ability to recognize scanned images.

[0082] Step 110: Generate training data based on the training images.

[0083] Specifically, the training images include fax training images and scanning training images, both of which are text line images. The training images are saved to a preset path, and at the same time, the preset path of the training images, as well as the string text corresponding to the training images, are saved. The above training images, preset path, and string text are integrated to obtain the training data.

[0084] In the above OCR training data generation method, by obtaining a training corpus, obtaining image parameter information, creating a blank image based on the image parameter information, extracting a preset string from the training corpus, writing the characters in the extracted string into the blank image according to the preset generated character parameter information to generate an initial image, performing an image enhancement process on the initial image to obtain training images, and generating training data for training the OCR deep learning engine based on the training images. This solution is used to optimize the recognition of fax images and scanned images. The pictures in the generated training data have the characteristics of real fax images or scanned images. When the OCR model trained using this training data is used to recognize fax images and scanned images, the recognition rate is significantly improved.

[0085] In one embodiment, obtaining the training corpus specifically includes: selecting a training corpus that meets the requirements through the Internet, preprocessing the training corpus through a regular algorithm, deleting duplicate spaces, garbled characters, and non-existent characters in the string text of the training corpus, and performing full-width and half-width conversion on the string text.

[0086] Specifically, the language materials stored in the corpus are those that have actually appeared in the real use of the language; the corpus is a basic resource that carries language knowledge with the help of a computer; the real corpus needs to be processed (analyzed and processed) to become a useful resource. There are various types of corpora, and the main basis for determining the type is its research purpose and use, which is often reflected in the principles and methods of corpus collection. Generally, corpora can be divided into four types: (1) Heterogeneous: There is no specific corpus collection principle, and various corpora are widely collected and stored as they are; (2) Homogeneous: Only collect corpora of the same type of content; (3) Systematic: Collect corpora according to pre-determined principles and proportions, making the corpora balanced and systematic, and able to represent the language facts within a certain range; (4) Specialized: Only collect corpora for a specific purpose. In addition, according to the language of the corpus, the corpus can also be divided into monolingual, bilingual, and multilingual. According to the collection unit of the corpus, the corpus can be further divided into text-based, sentence-based, and phrase-based. Bilingual and multilingual corpora can also be divided into parallel (aligned) corpora and comparative corpora according to the organization form of the corpus. The former forms a translation relationship between the corpora and is mostly used in application fields such as machine translation and bilingual dictionary compilation. The latter collects different language texts expressing the same content together and is mostly used in language contrast research. A large number of various types of corpora have been accumulated, such as: Portuguese corpus, Chinese-English news classification corpus for text classification research, Reuters text classification training corpus, Chinese text classification corpus, multilingual parallel corpus data of the large open subtitle library OpenSubtitles (OpenSubtitles Corpus), "Bible" bilingual corpus, Short messages service (SMS) corpus, etc. In this solution, select a corpus as the training corpus according to the actual situation.

[0087] In one embodiment, writing the characters in the extracted string into a blank image according to the preset generated character parameter information to generate an initial image includes: obtaining the generated character parameter information according to the image parameter information; obtaining the generated characters from the strings in the training corpus; writing the generated characters into the blank image according to the generated character parameter information to obtain the initial image.

[0088] Further, the image parameter information includes the image width and the image height; writing characters into the blank image according to the generated character parameter information to obtain the initial image, including: starting from the character starting horizontal position of the blank image, writing the generated characters into the blank image according to the generated character parameter information until the character starting horizontal position of the next generated character exceeds the image width, then stopping writing to obtain the initial image.

[0089] Specifically, as Figure 2 shown, extracting the preset string from the training corpus, and writing the characters in the extracted string into the blank image according to the preset generated character parameter information to generate the initial image, including:

[0090] Step 202, obtaining the generated character parameter information according to the image parameter information.

[0091] In this embodiment, the font of the generated characters is preset in advance. According to the image parameter information and the font of the generated characters, the generated character parameter information can be determined. The generated character parameter information includes the character height, and the character height is the maximum character height within the image height. The specific maximum character heights of different fonts are different, so it depends on the font.

[0092] Step 204, obtaining the generated characters from the strings in the training corpus.

[0093] Specifically, obtaining the font file corresponding to the font, and obtaining the character generator according to the font file. The character generator is a file object generated after reading the font file through the code library.

[0094] In this embodiment, the generated characters can be obtained from the strings in the training corpus through the character generator.

[0095] Step 206, starting from the character starting horizontal position of the blank image, writing the generated characters into the blank image according to the generated character parameter information until the character starting horizontal position of the next generated character exceeds the image width, then stopping writing to obtain the initial image.

[0096] In this embodiment, through the character generator, starting from the character starting horizontal position of the blank image, writing the first generated character at the character starting horizontal position of the blank image to obtain the width return parameter of the first generated character. Further, according to the character starting horizontal position of the blank image and the width return parameter, obtaining the starting horizontal position of the next generated character, and writing the next generated character at the starting horizontal position of the next generated character of the blank image, thereby obtaining the width return parameter of the next generated character. And so on, until the character starting horizontal position of the current generated character exceeds the image width, indicating that the blank image has been filled with characters, then obtaining the initial image.

[0097] Specifically, with reference to the image height and width, set the starting horizontal position where the first generated character char_1 is written at w_start_1. Use the character generator to obtain the generated character from the training corpus. Use the character generator to write the first generated character char_1 at the coordinate w_start_1 of the blank image, and obtain the width return parameter w_char_1 of the first generated character char_1.

[0098] Further, according to the character starting horizontal position w_start_1 and the width return parameter w_char_1, obtain the starting horizontal position w_start_2 of the next generated character char_2 = w_start_1 + w_char_1. Use the character generator to write the generated character char_2 at w_start_2 of the blank image, obtain the width return parameter w_char_2 of the generated character char_2, obtain the starting horizontal position w_start_3 of the next generated character char_3 = w_start_2 + w_char_2, and use the character generator to write the generated character char_3 at w_start_3 of the blank image, obtain the width return parameter w_char_3 of the generated character char_3...

[0099] Finally, until the character starting horizontal position w_start_n of the nth generated character exceeds the image width, obtain the initial image with the generated string written.

[0100] In one embodiment, as Figure 3 shown, perform image enhancement processing on the initial image to obtain the training image, including:

[0101] Step 302, perform random binarization on the initial image.

[0102] Step 304, randomly map the black pixels in the binarized initial image to a preset gray value range.

[0103] Step 306, perform dithering processing on the initial training image to generate a dot matrix binary image.

[0104] Step 308, perform random horizontal or vertical scaling on the dot matrix binary image to obtain the training image.

[0105] In this embodiment, the initial image is as Figure 7 shown, perform facsimile enhancement processing to generate a facsimile training image image_1 as Figure 8 shown.

[0106] Specifically, as Figure 4 shown, the facsimile enhancement processing includes:

[0107] Step 402: Randomly binarize the initial image within the threshold range of [80, 160] to obtain image_11.

[0108] Step 404: Randomly map the black pixels of image_11 to gray values in the range of [0, 60] to obtain image_12.

[0109] Step 406: Apply the bayer image dithering algorithm or the floydSteninberg image dithering algorithm with a size of 2x2 or 4x4 to image_12 for dithering processing to generate a dot matrix binary image image_13.

[0110] Step 408: Horizontally or vertically randomly scale image_13 using the nearest neighbor pixel method to obtain the fax training image image_1.

[0111] In this embodiment, through further enhancement processing of the initial image, the training image is made closer to the real fax image. Using the enhanced image to train the OCR model can significantly improve the recognition effect of the OCR model on fax images.

[0112] In one embodiment, as Figure 5 shown, perform image enhancement processing on the initial image to obtain a training image, including:

[0113] Step 502: Randomly binarize the initial image.

[0114] Step 504: Obtain an image with slight morphological dilation of black pixels based on the binarized initial image.

[0115] Step 506: Copy the initial image to obtain a copy image.

[0116] Step 508: Perform a vertical erosion operation on the initial image.

[0117] Step 510: Generate a mapping matrix of the initial image according to the image parameter information.

[0118] Step 512: For the initial image and the copy image, connect the horizontal strokes in the initial image that were removed by the erosion operation according to the mapping matrix.

[0119] Step 514: Perform random bilinear interpolation or cubic interpolation on the initial image horizontally or vertically.

[0120] Step 516: Randomly binarize the interpolated initial image to obtain the training image.

[0121] Further, for the initial image and the copy image, the horizontal strokes in the initial image removed by the erosion operation are connected according to the mapping matrix, including: if the value of the mapping matrix at a certain pixel point is the first point value, perform an AND operation on the pixel point values of the initial image and the copy image; if the value of the mapping matrix at a certain pixel point is the second point value, then delete the pixel value corresponding to the pixel point in the initial image.

[0122] In this embodiment, the initial image is as Figure 7 shown, and scanning enhancement processing is performed to generate a scanning training image image_2 as Figure 9 shown.

[0123] Specifically, as Figure 6 shown, the scanning enhancement processing includes:

[0124] Step 602: Randomly binarize the initial image within the threshold range of [80, 160] to obtain image_21.

[0125] Step 604: Use a dilation algorithm with a kernel size of 3x3 on image_21 to obtain an image image_22 with slight morphological dilation of black pixels.

[0126] Step 606: Copy image_22, and define the copy image as image_A.

[0127] Step 608: Perform a vertical erosion operation on image_22 to obtain an image with most of the overly thin horizontal strokes eroded, which is used as image_23.

[0128] Step 610: Generate a two-dimensional random 0, 1 mapping matrix M of size [h, w] according to the image height h and image width w.

[0129] Step 612: Perform an AND operation on image_A and image_23 according to the value of the mapping matrix M (if the value of M at a pixel point is 1, perform an AND operation on the pixel point values of image_A and the initial image image; if the value of M at a pixel point is 0, then delete the pixel value corresponding to the pixel point in the initial image image). Obtain an image image_24 with randomly connected horizontal strokes eroded in step 608.

[0130] Step 614: Perform random bilinear interpolation or cubic interpolation on the horizontal or vertical direction of image_24 to add deformation-unique noise to the image, obtaining image_25.

[0131] Step 616: Randomly binarize image_25 within the threshold range of [100, 150] to obtain the scanning training image image_2.

[0132] In this embodiment, through further enhancement processing of the initial image, the training image is made closer to a real scanned image. Using the enhanced image to train the OCR model can significantly improve the recognition effect of the OCR model on the scanned image.

[0133] In one embodiment, for each initial image, any one of fax enhancement processing and scan enhancement processing can be randomly used for image enhancement processing to obtain a training image, that is, the training image may contain both fax training images and scan training images.

[0134] In this embodiment, through further enhancement processing of the initial image, the training image is made closer to real fax images and scanned images. Using the enhanced image to train the OCR model can significantly improve the recognition effect of the OCR model on fax images and scanned images.

[0135] In one embodiment, training data is generated based on the training image, including: saving the training image to a preset path, saving the preset path of the training image, saving the string text corresponding to the training image, and integrating the training image, the preset path, and the string text to obtain the training data.

[0136] Specifically, the training image includes the above-mentioned fax training image image_1 and scan training image image_2. Save the training image to the preset path, save the preset path of the training image at the same time, save the string text corresponding to the training image, and finally integrate the training image, the preset path, and the string text to obtain the training data. The OCR model to be trained uses this training data for training, which can improve the accuracy of the OCR model in recognizing fax images and scanned images.

[0137] In the above OCR training data generation method, by obtaining a training corpus, obtaining image parameter information, establishing a blank image according to the image parameter information, extracting a preset string in the training corpus, writing the characters in the extracted string into the blank image according to the preset generated character parameter information to generate an initial image, performing image enhancement processing on the initial image to obtain a training image, and generating training data for training the OCR deep learning engine based on the training image. This solution is used to optimize the recognition of fax images and scanned images. The pictures in the generated training data have the characteristics of real fax images or scanned images. When the OCR model trained with this training data recognizes fax images and scanned images, the recognition rate is significantly improved.

[0138] It should be understood that although Figures 1 to 6The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figures 1 to 6 At least a part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0139] In one embodiment, as Figure 10 shown, an OCR training data generation device is provided, including: a corpus acquisition module 1001, an image creation module 1002, a character extraction module 1003, an image enhancement module 1004, and a data generation module 1005, where:

[0140] The corpus acquisition module 1001 is used to acquire a training corpus.

[0141] The image creation module 1002 is used to acquire image parameter information and create a blank image according to the image parameter information.

[0142] The character extraction module 1003 is used to extract a preset string from the training corpus and write the characters in the extracted string into the blank image according to the preset generated character parameter information to generate an initial image.

[0143] The image enhancement module 1004 performs image enhancement processing on the initial image to obtain a training image.

[0144] The data generation module 1005 is used to generate training data according to the training image.

[0145] In one embodiment, the character extraction module 1003 includes:

[0146] The character parameter acquisition module is used to acquire generated character parameter information according to the image parameter information.

[0147] The generated character acquisition module is used to acquire generated characters from the strings in the training corpus.

[0148] The generated character writing module is used to start writing the generated characters into the blank image from the character start horizontal position of the blank image, and stop writing until the character start horizontal position of the next generated character exceeds the image width, so as to obtain an initial image.

[0149] In one embodiment, the image enhancement module 1004 includes:

[0150] A binarization module, used to perform random binarization on the initial image.

[0151] A mapping module, used to randomly map the black pixels in the binarized initial image to a preset gray value range.

[0152] A dot matrix binary image generation module, used to perform dithering on the initial training image to generate a dot matrix binary image.

[0153] A scaling module, used to randomly scale the dot matrix binary image horizontally or vertically to obtain a training image.

[0154] In another embodiment, the image enhancement module 1004 includes:

[0155] A primary binarization module, used to perform random binarization on the initial image.

[0156] A black pixel dilation module, used to obtain an image of slight morphological dilation of black pixels according to the binarized initial image.

[0157] A copy module, used to copy the initial image to obtain a copied image.

[0158] An erosion operation module, used to perform vertical erosion operation on the initial image.

[0159] A mapping module, used to generate a mapping matrix of the initial image according to image parameter information;

[0160] A stroke connection module, used to connect the horizontal strokes in the initial image that have been processed by the erosion operation for the initial image and the copied image according to the mapping matrix. If the value of the mapping matrix at a certain pixel point is the first point value, perform an AND operation on the pixel point values corresponding to the initial image and the copied image; if the value of the mapping matrix at a certain pixel point is the second point value, then delete the pixel value corresponding to that pixel point in the initial image.

[0161] An interpolation module, used to perform random bilinear interpolation or cubic interpolation on the initial image horizontally or vertically.

[0162] A secondary binarization module, used to perform random binarization on the interpolated initial image to obtain a training image.

[0163] In one embodiment, the data generation module 1005 includes:

[0164] An image saving module, used to save the training image to a preset path;

[0165] A path saving module, used to save the preset path of the training image;

[0166] A text saving module, used to save the string text corresponding to the training image;

[0167] An integrated data module for integrating training images, preset paths, and string texts to obtain training data.

[0168] For the specific limitations of the OCR training data generation device, reference can be made to the limitations of the OCR training data generation method in the above text, which will not be elaborated here. Each module in the above OCR training data generation device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0169] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 11 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store OCR training data. The network interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements an OCR training data generation method.

[0170] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 11 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an OCR training data generation method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad set on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0171] Those skilled in the art can understand that Figure 11 The structure shown in Figure 11 is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0172] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0173] Obtain a training corpus;

[0174] Obtain image parameter information and create a blank image according to the image parameter information;

[0175] Extract a preset string from the training corpus, and write the characters in the extracted string into the blank image according to the preset generated character parameter information to generate an initial image;

[0176] Perform image enhancement processing on the initial image to obtain a training image;

[0177] Generate training data according to the training image.

[0178] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0179] Obtain the generated character parameter information according to the image parameter information;

[0180] Obtain the generated characters from the strings in the training corpus;

[0181] Write the generated characters into the blank image according to the generated character parameter information to obtain an initial image.

[0182] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0183] Starting from the character starting horizontal position of the blank image, write the generated characters into the blank image according to the generated character parameter information until the character starting horizontal position of the next generated character exceeds the image width, then stop writing to obtain an initial image.

[0184] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0185] Perform random binarization on the initial image;

[0186] Randomly map the black pixels in the binarized initial image to a preset gray value range;

[0187] Perform dithering on the initial training image to generate a dot matrix binary image;

[0188] Perform random scaling on the dot matrix binary image horizontally or vertically to obtain the training image.

[0189] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0190] Perform random binarization on the initial image;

[0191] Obtain an image of morphological dilation of black pixels slightly based on the binarized initial image;

[0192] Copy the initial image to obtain a copy image;

[0193] Perform vertical erosion operation on the initial image;

[0194] Generate a mapping matrix of the initial image according to the image parameter information;

[0195] For the initial image and the copy image, connect the horizontal strokes in the initial image processed by the erosion operation according to the mapping matrix;

[0196] Perform random bilinear interpolation or cubic interpolation on the initial image horizontally or vertically;

[0197] Perform random binarization on the interpolated initial image to obtain the training image.

[0198] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0199] If the value of the mapping matrix at a certain pixel point is the first point value, perform an AND operation on the pixel point values of the initial image and the copy image corresponding to it;

[0200] If the value of the mapping matrix at a certain pixel point is the second point value, delete the pixel value corresponding to this pixel point in the initial image.

[0201] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0202] Save the training image to the preset path;

[0203] Save the preset path of the training image;

[0204] Save the string text corresponding to the training image;

[0205] Integrate the training image, the preset path, and the string text to obtain the training data.

[0206] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0207] Obtain a training corpus;

[0208] Obtain image parameter information and create a blank image according to the image parameter information;

[0209] Extract a preset string from the training corpus, and write the characters in the extracted string into the blank image according to the preset generated character parameter information to generate an initial image;

[0210] Perform image enhancement processing on the initial image to obtain a training image;

[0211] Generate training data according to the training image.

[0212] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0213] Obtain the generated character parameter information according to the image parameter information;

[0214] Obtain the generated characters from the strings in the training corpus;

[0215] Write the generated characters into the blank image according to the generated character parameter information to obtain an initial image.

[0216] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0217] Starting from the character start horizontal position of the blank image, write the generated characters into the blank image according to the generated character parameter information until the character start horizontal position of the next generated character exceeds the image width, then stop writing to obtain an initial image.

[0218] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0219] Perform random binarization on the initial image;

[0220] Randomly map the black pixels in the binarized initial image to a preset gray value range;

[0221] Perform dithering processing on the initial training image to generate a dot matrix binary image;

[0222] Perform random horizontal or vertical scaling on the dot matrix binary image to obtain a training image.

[0223] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0224] Perform random binarization on the initial image;

[0225] Obtain an image of slight morphological dilation of black pixels from the binarized initial image;

[0226] Make a copy of the initial image to obtain a copy image;

[0227] Perform a vertical erosion operation on the initial image;

[0228] Generate a mapping matrix of the initial image according to the image parameter information;

[0229] For the initial image and the copy image, connect the horizontal strokes in the initial image removed by the erosion operation according to the mapping matrix;

[0230] Perform random bilinear interpolation or cubic interpolation on the initial image horizontally or vertically;

[0231] Perform random binarization on the interpolated initial image to obtain a training image.

[0232] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0233] If the value of the mapping matrix at a certain pixel point is the first point value, perform an AND operation on the pixel point values of the corresponding pixel points of the initial image and the copy image;

[0234] If the value of the mapping matrix at a certain pixel point is the second point value, delete the pixel value corresponding to the pixel point in the initial image.

[0235] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0236] Save the training image to a preset path;

[0237] Save the preset path of the training image;

[0238] Save the string text corresponding to the training image;

[0239] Integrate the training image, the preset path, and the string text to obtain training data.

[0240] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0241] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0242] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. An OCR training data generation method, characterized in that, The method includes: Obtain a training corpus; Obtain image parameter information, and establish a blank image according to the image parameter information; Extract a preset string from the training corpus, and write the characters in the extracted string into the blank image according to the preset generated character parameter information to generate an initial image; Perform image enhancement processing on the initial image to obtain a training image; Generate training data according to the training image; The performing image enhancement processing on the initial image to obtain a training image includes: Perform random binarization on the initial image; Obtain an image of slight morphological dilation of black pixels according to the binarized initial image; Make a copy of the initial image to obtain a copied image; Perform a vertical erosion operation on the initial image; Generate a mapping matrix of the initial image according to the image parameter information; For the initial image and the copied image, connect the horizontal strokes in the initial image removed by the erosion operation according to the mapping matrix; Perform random bilinear interpolation or cubic interpolation on the horizontal or vertical direction of the initial image; Perform random binarization on the interpolated initial image to obtain a training image.

2. The OCR training data generation method according to claim 1, wherein The writing the characters in the extracted string into the blank image according to the preset generated character parameter information to generate an initial image includes: Obtain the generated character parameter information according to the image parameter information; Obtain generated characters from the strings in the training corpus; Write the generated characters into the blank image according to the generated character parameter information to obtain an initial image.

3. The OCR training data generation method according to claim 2, wherein The image parameter information includes the image width and the image height; the writing the characters into the blank image according to the generated character parameter information to obtain an initial image includes: Starting from the character starting horizontal position of the blank image, write the generated characters into the blank image according to the generated character parameter information until the character starting horizontal position of the current generated character exceeds the image width, to obtain an initial image.

4. The OCR training data generation method according to claim 1, wherein For the initial image and the copied image, connecting the horizontal strokes in the initial image removed by the erosion operation according to the mapping matrix includes: If the value of the mapping matrix at a certain pixel point is the first point value, perform an AND operation on the pixel point values corresponding to the initial image and the copied image; If the value of the mapping matrix at a certain pixel point is the second point value, delete the pixel value corresponding to the pixel point in the initial image.

5. The OCR training data generation method according to claim 1, wherein The generating training data according to the training image includes: Save the training image to a preset path; Save the preset path of the training image; Save the string text corresponding to the training image; Integrate the training image, the preset path, and the string text to obtain training data.

6. An OCR training data generation device, characterized in that, The device includes: A corpus acquisition module, configured to obtain a training corpus; An image establishment module, configured to obtain image parameter information and establish a blank image according to the image parameter information; A character extraction module, configured to extract a preset string from the training corpus and write the characters in the extracted string into the blank image according to the preset generated character parameter information to generate an initial image; An image enhancement module that performs image enhancement processing on the initial image to obtain a training image; A data generation module for generating training data according to the training image; The image enhancement module includes: A primary binarization module for randomly binarizing the initial image; A black pixel dilation module for obtaining an image of slight morphological dilation of black pixels according to the binarized initial image; A copy module for copying the initial image to obtain a copy image; An erosion operation module for performing a longitudinal erosion operation on the initial image; A mapping module for generating a mapping matrix of the initial image according to the image parameter information; A stroke connection module for connecting the horizontal strokes in the initial image that are processed away by the erosion operation on the initial image and the copy image according to the mapping matrix; An interpolation module for randomly performing bilinear interpolation or cubic interpolation on the horizontal or vertical direction of the initial image; A secondary binarization module for randomly binarizing the interpolated initial image to obtain a training image.

7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the OCR training data generation method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the OCR training data generation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Character recognition model generation method and device, equipment and medium

    CN109753968A

  • Automating creation of accurate OCR training data using specialized UI application

    US20180096200A1