An OCR training sample generation method, device and system
By combining text contour extraction and image repair techniques, high-quality OCR training samples are generated and labeled files are automatically generated, which solves the problems of poor sample authenticity and poor model generalization ability in the existing technology, and achieves efficient and accurate OCR training samples generation.
Patent Information
- Application Number
- CN202111646988.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-12-29
AI Technical Summary
The existing OCR training sample generation methods have problems such as poor sample authenticity, poor model generalization ability, difficulty in obtaining training data, expensive and time-consuming labeling, and slow inference speed.
Combining text outline extraction algorithm and image repair technology, high-quality training pictures are generated using the original image background information, and automatically generate label files corresponding to the pictures, eliminating the cumbersome labeling work.
Overcoming the problems of lack of samples, poor authenticity of generated samples, and poor generalization ability of model, the generated samples can be directly used for OCR model training, improving training efficiency and sample quality.
Smart Images

Figure CN114419632B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to an OCR training sample generation method, device and system. Background Art
[0002] OCR (Optical Character Recognition) tasks widely exist in life and business scenarios. Currently, the best method for this task is to use deep learning technology for text localization and recognition. However, in real scenarios, there is often a lack of training samples of a certain layout, and deep learning depends on a large number of training samples. Therefore, methods for automatically generating OCR training samples have emerged.
[0003] Currently, the methods for generating OCR samples are generally divided into two categories:
[0004] The first category is the template-based method. That is, a clean template image is input, and then random text is written at the specified position of the template to generate OCR samples of the template style. This method has two main defects: one is that in most cases, we cannot obtain the template image; the other is that the template image is often a clean and distortion-free ideal image, which has a large gap with real images mixed with various noise interferences, and the generated samples have poor authenticity.
[0005] The second category is the deep learning-based image editing technology. This type of method has four main defects: the first defect is the poor generalization ability of the model. Since the OCR image features and text features of different layouts vary greatly, for image layouts or text styles that have not appeared during the model training process, the model often performs poorly; the second defect is the difficulty in obtaining training data. Due to the limited generalization ability of the model, if we want to train an image editing model for a certain layout, we first need a large amount of training data similar to this layout, and the current scenario is precisely because of the lack of data that we need image editing technology; the third defect is that the samples for training the image editing model require pixel-level annotation, and this annotation work is very expensive and time-consuming; the fourth defect is that the model inference speed is relatively slow, and the operation often requires a deep learning inference framework, greatly increasing the engineering complexity. Summary of the Invention
[0006] The present invention relates to an OCR training sample generation method. Aiming at the problem of lack of samples often encountered in the OCR training process, the applicant innovatively combines technologies such as text contour extraction algorithms and image restoration, makes full use of the background information of the original image, generates high-quality training images, and at the same time generates an annotation file corresponding to the image (including text content and position information), eliminating the cumbersome and laborious annotation work, and can be directly used for OCR model training. Thus, the present invention directly processes real original OCR images based on traditional computer vision methods, overcoming the above disadvantages.
[0007] According to a first aspect of the present invention, there is provided an OCR training sample generation method, the input information being an original image and the coordinates of an erased area, the method comprising:
[0008] A text contour extraction step of extracting all text contours based on the original image, determining an erased area mask (mask) in combination with the coordinates of the erased area, and obtaining a repaired area mask;
[0009] An image repair and filling step of performing image repair and filling according to the repaired area mask and the pixel information around the repaired area to obtain a background template after the erased text;
[0010] A random text generation step of generating random text in each generation area, thereby obtaining a new sample picture and a corresponding annotation information file.
[0011] Further, the text contour extraction step specifically includes:
[0012] Converting the input original image into a single-channel grayscale image, and then adaptively binarizing it to obtain all text contour masks, with the text area value being 1 and the background area value being 0;
[0013] According to the coordinates of the erased area, obtaining an erased area mask, with the erased area value being 1 and the other area values being 0;
[0014] Multiplying the pixels at the corresponding positions of all text contour masks and the erased area mask to obtain an erased area text contour mask;
[0015] Performing morphological dilation on the erased area text contour mask, thereby obtaining a repaired area mask.
[0016] Further, the dilation kernel of the morphological dilation is 2 or 3.
[0017] Further, the image repair and filling step specifically includes:
[0018] Determining the area to be repaired in the original image according to the repaired area mask;
[0019] Polling each pixel point of the area to be repaired in the order from outside to inside, and calculating the pixel value to be filled at the repair point according to the information of the known pixels around a certain pixel point, becoming a known pixel;
[0020] Calculating the pixel value of the next pixel point inward;
[0021] Iterating step by step, the area to be repaired gradually shrinks and becomes smaller until the area to be repaired is all repaired, obtaining a repaired background template after the erased text.
[0022] Further, the sorting algorithm from outside to inside is the Fast Marching Method.
[0023] Further, the method for calculating the pixel value to be filled at the repair point according to the information of the known pixels around a certain pixel includes weighted average of the known pixel values in the neighborhood, INPAINT_NS or INPAINT_TELEA method.
[0024] Here, INPAINT_NS (repair method based on Navier - Stokes) and INPAINT_TELEA (fast matching method based on image gradient, also known as Telea method) are two common image repair algorithms.
[0025] Further, the generation area is the position area where sample text needs to be generated in the background template after erasing the text.
[0026] Further, the random text generation step specifically includes:
[0027] For a certain generation area, determine the expected length w of the random text for this generation area, set the font size as s, and estimate the number of characters n of this section of random text as n = int(w / s);
[0028] Generate redundant random text with a length of n * k characters, where k is the redundancy multiple and takes positive integer values;
[0029] According to the relationship between the length of the redundant random text and the expected length w of the random text, determine the finally generated random text and its actual length L;
[0030] Randomly determine the position of the generated text within the generation area, write the finally generated random text and determine its annotation information;
[0031] Poll each generation area, and thus obtain a new sample picture and the corresponding annotation information file.
[0032] Further, the determination of the expected length w of the random text for this generation area specifically includes:
[0033] Take the maximum value of the length and width of this generation area as the maximum length of the generated text, and the minimum value of the length and width as the minimum length of the generated text. Randomly select an integer between this minimum value and the maximum value as the expected length of the random text, denoted as w.
[0034] Further, in the step of generating redundant random text with a length of n * k characters, if the corpus type is specified, generate redundant random text with a length of n * k characters in this corpus type; if the corpus type is not specified, randomly generate redundant random text with a length of n * k characters.
[0035] Further, the method for generating the specified corpus includes:
[0036] Regular expression generation: used to generate a corpus with clear rules. The required corpus rule features are compiled into a regular expression, and then a random string is generated according to the regular expression; or
[0037] Random acquisition from a database: used to generate a corpus with unclear rules or a certain fixed type. All the content of the corpus is obtained through public information, and the data is stored in a database or a data file by entry. When generating randomly, one entry can be randomly obtained from it.
[0038] Further, the value of k is taken as 3.
[0039] Further, determining the finally generated random text and its actual length L according to the relationship between the length of the redundant random text and the expected length w of the random text specifically includes:
[0040] Starting from the first character of the redundant random text, count the actual length of each character and accumulate them in turn until the total character length just meets the condition of "exceeding w if adding one more character". Take it as the finally generated random text, denoted as "text", and record the actual length L of the finally generated random text.
[0041] Further, randomly determining the position of the generated text within the generation area, writing the finally generated random text and determining its annotation information specifically includes:
[0042] Assume that the upper left corner coordinates of the generation area are (x1, x2), and the lower right corner coordinates are (x2, y2). Then the x-axis coordinate range of the starting point of the upper left corner of the finally generated random text is [x1, (x2 - w)], and the y-axis coordinate range is [y1, (y2 - s)]. Randomly select an integer point (x, y) within this range as the starting position of the finally generated random text;
[0043] On the background template, with (x, y) as the starting position of the upper left corner of the text, write the finally generated random text, and its annotation information is: the coordinates are [(x, y), (x + L, y), (x + L, y + s), (x, y + s)]; the text content is "text".
[0044] Further, after the step of randomly determining the position of the generated text within the generation area, writing the finally generated random text and determining its annotation information, there is also a step of adjusting the font size and color of the finally generated random text.
[0045] Further, the adjustment of the font size of the finally generated random text specifically includes: determining the width h of the erased area and defaulting the size of the finally generated random text to h.
[0046] Further, the adjustment of the font color of the finally generated random text specifically includes:
[0047] For a certain generation area, select the corresponding area from the repair area mask, and perform skeleton extraction on the text contour of this area to form a text skeleton area;
[0048] Extract the corresponding area of this generation area from the original image. For the RGB three channels of this area, average all the pixel values at the text skeleton area of each channel, which is the color value of each channel of the text skeleton area;
[0049] Use this color value as the color of the generated text, so that the font color of the finally generated random text is approximately the same as the original text color.
[0050] Here, common skeleton extraction algorithms include the K3M algorithm, the Zhang-Suen algorithm, etc.
[0051] According to the second aspect of the present invention, there is provided an OCR training sample generation device, and the device operates based on the method provided in any of the foregoing aspects. The device includes:
[0052] A text contour extraction module, configured to extract all text contours based on the original image, determine the erased area mask in combination with the erased area coordinates, and obtain the repair area mask;
[0053] An image repair and filling module, configured to perform image repair and filling according to the repair area mask and the pixel information around the repair area to obtain the background template after erasing the text;
[0054] A random text generation module, configured to generate random text in each generation area, thereby obtaining a new sample picture and a corresponding annotation information file.
[0055] According to the third aspect of the present invention, there is provided an OCR training sample generation system, and the system includes: a processor and a memory for storing executable instructions; wherein, the processor is configured to execute the executable instructions to execute the OCR training sample generation method as described in any of the foregoing aspects.
[0056] According to the fourth aspect of the present invention, there is provided a computer-readable storage medium, characterized in that a computer program is stored thereon, and when the computer program is executed by a processor, it implements the OCR training sample generation method as described in any of the foregoing aspects.
[0057] Advantages of the present invention:
[0058] 1. An OCR training sample generation method is invented, which can be used to solve the problem of sample shortage when training an OCR model using deep learning;
[0059] 2. Innovatively, the dilated text contour is used as the "image area to be restored", and an image restoration algorithm is used to erase the text. This can not only ensure that the text traces are erased cleanly, but also maximize the retention of surrounding pixel information for the text area during image restoration, obtaining the best restoration effect;
[0060] 3. An optional adaptive text size module is innovatively proposed, which can adaptively restore the different text sizes at various positions in the real samples as needed, making the generated pictures closer to the real samples in terms of text size;
[0061] 4. An optional adaptive text color module is innovatively proposed, which can adaptively restore the different text colors at various positions in the real samples as needed, making the generated pictures closer to the real samples in terms of text color;
[0062] 5. Annotation files can be generated simultaneously, containing information on the text position coordinates and text content required for OCR training. This eliminates the cumbersome and laborious answer annotation work. Description of the Drawings
[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.
[0064] Figure 1 Shows an example of the original image according to an embodiment of the present invention.
[0065] Figure 2 Shows an example of the extracted text contour according to an embodiment of the present invention.
[0066] Figure 3 Shows an example of the generated background template according to an embodiment of the present invention.
[0067] Figure 4 Shows an example of the generated new sample and new text box according to an embodiment of the present invention.
[0068] Figure 5 Shows an example of the result of adding an adaptive text size module according to an embodiment of the present invention.
[0069] Figure 6 Shows an example of the result of adding an adaptive text color module according to an embodiment of the present invention.
[0070] Figure 7 Shows an example of the result of adding a corpus configuration module according to an embodiment of the present invention.
[0071] Figure 8 Shows a module flowchart according to an embodiment of the present invention.
[0072] Figure 9 Shows an algorithm flowchart of a text contour extraction module according to an embodiment of the present invention.
[0073] Figure 10 Shows an algorithm flowchart of an image restoration module according to an embodiment of the present invention.
[0074] Figure 11 Shows an algorithm flowchart of a random generation module according to an embodiment of the present invention.
[0075] Figure 12 Shows an algorithm flowchart of an adaptive text size module according to an embodiment of the present invention.
[0076] Figure 13 Shows an algorithm flowchart of an adaptive text color module according to an embodiment of the present invention.
[0077] Figure 14 Shows a schematic diagram of the text contour of a certain generation area according to an embodiment of the present invention.
[0078] Figure 15 Shows a schematic diagram of the text contour backbone according to an embodiment of the present invention.
[0079] Figure 16 Shows the original image of the generation area according to an embodiment of the present invention.
[0080] Figure 17 Shows the original Figure 1 Schematic diagram.
[0081] Figure 18 Shows the generated Figure 1 Schematic diagram.
[0082] Figure 19 Shows the original Figure 2 Schematic diagram.
[0083] Figure 20 Shows the generated Figure 2 Schematic diagram.
[0084] Figure 21 Shows the original Figure 3Schematic diagram.
[0085] Figure 22 Show the generation according to an embodiment of the present invention Figure 3 Schematic diagram.
[0086] The realization of the object of the present invention, functional features and advantages will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific embodiments
[0087] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0088] The terms "first", "second", etc. in the specification and claims of the present disclosure are used to distinguish similar objects and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein, for example.
[0089] In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0090] Multiple, including two or more.
[0091] And / or, it should be understood that for the term "and / or" used in the present disclosure, it is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone.
[0092] The present invention provides a method for generating OCR training samples, mainly related to image text contour extraction technology, image backbone extraction technology, image repair and filling technology:
[0093] Text contour extraction technology:
[0094] Text contour extraction methods include traditional computer vision-based methods and deep learning-based methods. Traditional methods mainly include various binarization methods, such as fixed-threshold binarization, OTSU binarization, adaptive binarization, etc. The advantages of such traditional methods are simple and fast algorithms and strong interpretability. The disadvantages are that the extracted content is more likely to contain noise, is susceptible to brightness, and has poor robustness.
[0095] Deep learning-based methods are mainly image segmentation algorithms. The advantages are strong anti-noise ability, not easily affected by brightness, and strong robustness. The disadvantages are that the annotation of text contour training samples is very difficult, and deep learning requires model training, resulting in slow inference speed. Therefore, text contour extraction based on deep learning methods is extremely rare in industrial scenarios.
[0096] Image backbone extraction technology:
[0097] Image backbone extraction technology is used to reduce the objects in a binary image to a 1-pixel-wide representation. Its working principle is to continuously iterate the image, removing the pixels on the object boundary without changing its connectivity until no more pixels can be deleted. The present invention uses this technology to extract the backbone of the text in the image.
[0098] Image inpainting and filling technology:
[0099] Image inpainting and filling technology can be divided into three categories: image filling, image inpainting, and texture synthesis. Image filling is generally applied when the area to be filled is large. The general idea is to set the size of the filling size for each iteration, find the area that best matches the pixels around the area to be filled, and directly copy its pixels to the area to be filled, iterating until all areas to be filled are completed. Image inpainting technology is generally applied when the area to be filled is small. The general idea is to poll each pixel point in the area to be filled from the outside to the inside, and calculate the pixel value to be filled according to the known pixel information around the pixel point through a certain algorithm until all pixels are filled. Texture synthesis is generally applied to images with obvious texture features, such as cells and brick walls, and the content to be filled is the growth of the texture of the known area. The OCR sample image does not conform to the characteristics of texture synthesis. In the experimental stage of the present invention, detailed experiments were carried out on these two methods of image filling and image inpainting, and it was found that the method of image inpainting has the best effect in this scenario.
[0100] Embodiment
[0101] The present invention relates to an OCR training sample generation method. This method can adaptively generate a large number of high-quality OCR samples to solve the problem of lack of OCR training samples.
[0102] If you want to Figure 1For the original image, a large number of samples of this layout are generated. For the convenience of display, the text boxes added later are both the "erasure areas" where we want to erase the text and the "generation areas" where we want to generate text (the erasure area and the generation area do not have to coincide).
[0103] In the first step, enter the text contour extraction module to extract the text contours within all "erasure areas" on the entire image, as Figure 2 shown.
[0104] In the second step, enter the image inpainting module. Take the white part of the text contour ( Figure 2 in it) as the damaged part of the original image, and perform image inpainting and filling based on the pixel information around the damaged part to obtain the background template after erasing the text, as Figure 3 shown.
[0105] In the third step, enter the random generation module to generate random text of the specified style within the specified generation area, and record the coordinates and text content of the generated text as annotation information, as Figure 4 shown.
[0106] The overall module flow chart is as follows Figure 8 shown. The algorithm flow chart within each module can be seen in the description part of each module.
[0107] Text contour extraction module
[0108] As Figure 9 shown, the input of the text contour extraction module is the original image and the coordinates of the "erasure area". For the input image, we first convert it into a single-channel grayscale image, and then perform adaptive binarization to obtain a mask with pixel values of 0 and 1, that is, the text contour mask, denoted as mask1, where the area with a value of 1 is the text area and the area with a value of 0 is the background area. For all the input "erasure area" coordinates, we also construct a mask, making the erasure area value 1 and other area values 0, denoted as mask2. Multiply the pixels at the corresponding positions of mask1 and mask2. For the part where the value of mask2 is 0, the multiplication result is 0; for the part where the value of mask2 is 1, the multiplication result is the original value of mask1. In this way, the mask of the "erasure area text contour" can be obtained.
[0109] Since the edges of the text contours extracted by binarization may not be complete, and the traces of the text edges cannot be completely erased after patching, we perform morphological dilation (the dilation kernel is 2 or 3) on the "erasure area text contour" mask obtained in the previous step to ensure that the text contours can completely cover the text in the original image, and thus obtain the inpainting area mask, as Figure 2 shown.
[0110] If we do not want to erase any text on the picture, we can input an empty "erasure area coordinate". At this time, all values of the constructed mask2 are 0. After multiplying with mask1, the result is also all 0. After morphological dilation, it remains all 0. Therefore, the mask value of the "restoration area" obtained by this module is all 0, which means that no area needs to be restored.
[0111] If we want to erase all the text on the picture, in addition to the method of inputting the coordinates of all texts as the erasure area, we can also directly ignore mask2, use mask1 directly as the "erasure area text mask", and then perform morphological dilation on it to obtain the restoration area mask.
[0112] Image Restoration Module
[0113] The function of the image restoration module is to erase the text that needs to be erased on the picture and generate a background template. The restoration area mask obtained from the previous module is used to tell this module the specific area that needs to be restored. Now we poll each pixel point of the area to be restored in the order from outside to inside (the sorting algorithm from outside to inside can be freely selected, such as the Fast Marching Method), and calculate the pixel value that should be filled at this restoration point according to the known pixel information around this pixel point. The reason for choosing the order from outside to inside, that is, filling the pixel points on the periphery of the area to be patched first and then the internal pixel points, is that there are more known pixels in the neighborhood of the pixel points on the contour periphery, containing richer known pixel information. In this way, the filling value calculated based on the surrounding pixel points is more accurate and smooth. When a pixel point to be restored is filled, it is considered that this pixel point becomes a known pixel. When calculating the surrounding pixel points to be filled, it can be used as known pixel information and provide a reference value. The specific method of calculating the pixel value based on the pixel information around a certain pixel point can be freely selected, such as the weighted average of the known pixel values in the neighborhood, [INPAINT_NS] that takes both speed and effect into account, [INPAINT_TELEA], etc.
[0114] Through the gradual iteration of the above methods, the area to be restored gradually shrinks and becomes smaller until the entire area is restored, obtaining the restored background template picture, as Figure 3 shown.
[0115] Random Generation Module
[0116] The random generation module is used to generate random text and annotation files at specified positions or random positions on the background template. By setting the generation quantity, this module is used to quickly generate the specified number of new samples. This module requires the user to specify the font size and color of the generated text. If the user does not specify, the set default values are used.
[0117] For each "generation area", we take the maximum value of its length and width as the maximum length of the generated text, and the minimum value of its length and width as the minimum length of the generated text. Then, we randomly select an integer between this minimum value and the maximum value as the actual length of the generated text for this time, denoted as w. This not only realizes the randomness of the text length but also ensures that the text length does not exceed the generation area. We denote the set font size as s, and estimate the number of characters in this text segment as n = int(w / s). If the corpus type is specified, we generate redundant text with a character length of n * k using this corpus type. If the corpus type is not specified, we randomly generate redundant text with a character length of n * k. Here, k is the redundancy multiple. Since the estimated number of characters above is the case where the width of each character is exactly equal to its height, and the aspect ratios of different characters are not necessarily strictly 1:1, we multiply n by k to ensure that the number of characters in the generated redundant text is sufficient. Usually, k can be taken as 3. Then, starting from the first character of the redundant text, we count the actual length of each character and accumulate them in sequence until the total character length just satisfies "adding one more character will exceed w". We take this as the finally generated text, denoted as text, and record its true length at this time as L.
[0118] Next, we implement the function of randomly positioning the text generation within the generation area. To ensure that the generated text does not exceed the generation area, let the upper-left corner coordinates of the generation area be (x1, x2), and the lower-right corner coordinates be (x2, y2). Then, the x-axis coordinate range of the starting point of the upper-left corner of the text should be [x1, (x2 - w)], and the y-axis coordinate range should be [y1, (y2 - s)]. We randomly select an integer point within this range as the starting position of the generated text to achieve the position randomization function.
[0119] Then, on the image, starting from the position (x, y), we write the "text" string with the specified color and size. The annotation information of this text line can also be obtained: the coordinates are [(x, y), (x + L, y), (x + L, y + s), (x, y + s)], and the text content is "text". By polling each "generation area" in this process, we can obtain a new image and the corresponding answer annotation file.
[0120] So far, the generation of OCR training samples has been completed, and the effect is as Figure 4 shown.
[0121] Adaptive text size module
[0122] The adaptive text size module is an optional module used to adaptively adjust the text size of each generated text line so that the text size generated in each area is approximately the same as the original text size at that position.
[0123] To retain more background information, the erasure areas we specify should be close to the text. At this time, the width of the erasure area is the size of the characters therein, denoted as h. We utilize this characteristic and set the default font size of the generated text to h, so that the size of the text generated in this area is approximately the same as the size of the original text at this position.
[0124] This module has three advantages:
[0125] 1. It can keep the size of the text generated in each area approximately the same as the size of the original text at this position, maximizing the restoration of different font features at different positions in the original image;
[0126] 2. It eliminates the cumbersome steps of manually setting a fixed font size for each generated text area;
[0127] 3. It solves the problem that it is difficult to control the size when manually setting the font.
[0128] The sample generation effect after adding this module is as Figure 5 shown.
[0129] Adaptive Text Color Module
[0130] The Adaptive Text Color Module is an optional module used to adaptively adjust the font color of each generated text line so that the color of the text generated in each area is approximately the same as the original text color.
[0131] For each generated area, we select the corresponding generated area from the "Text Outline Extraction Module", such as Figure 14 . Perform backbone extraction on the text outline of this area, such as Figure 15 . The purpose of extracting the backbone is to filter out some background pixels that may be included in the text outline boundary. The colors of these background pixels often differ greatly from the text color and will introduce errors. Then, for the RGB three channels of the original image of the generated area (such as Figure 16 ), taking the R channel as an example, we average all the pixel values at the text backbone area of the R channel, which is the R channel color value of the text backbone in this area. By obtaining the average value of each channel in this way, it is the RGB three-channel color average value of the text in the current area. Use this color value as the color of the generated text, so that the color of the text generated in each area is approximately the same as the original text color.
[0132] This module has three advantages:
[0133] 1. It can keep the text color generated in each area approximately the same as the original text color at this position, maximizing the restoration of different font color features at different positions in the original image;
[0134] 2. It eliminates the cumbersome steps of manually setting a fixed font color for each generated text area;
[0135] 3. Solve the problem that it is difficult to control the font color and color pixel value when setting them manually.
[0136] The sample generation effect after adding this module is as Figure 6 shown.
[0137] Corpus configuration module
[0138] The corpus configuration module is an optional module used to configure the corpus type of each generated text line, so that the corpus meaning of the generated text in each area is consistent with the original text. This module is embedded in the "random generation module", as Figure 11 shown.
[0139] This module uses two methods to generate the specified corpus: 1. Generate by regular expression; 2. Randomly obtain from the database.
[0140] The method of generating by regular expression is mainly used to generate corpus with clear rules, such as mobile phone numbers, ID card numbers, etc. Such corpus has fixed character lengths and encoding rules. We code the required corpus rule features into regular expressions and then randomly generate strings according to the regular expressions.
[0141] The method of randomly obtaining from the database is mainly used to generate corpus with unclear rules or certain fixed-type corpus, such as company names, bank names, country names, provinces, cities, languages, etc. For such corpus, we first obtain all the content of the corpus through public information and store the data item by item in the database or data file, and randomly obtain one of them when generating randomly.
[0142] The sample generation effect after adding this module is as Figure 7 shown. Figures 17 - 22 It is the effect diagram of generating virtual samples for OCR samples of different styles using this method.
[0143] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.
[0144] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0145] Through the description of the above embodiments, those skilled in the art can clearly understand that the above implementation methods can be realized by means of software plus a necessary general hardware platform. Of course, they can also be realized by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0146] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the claims of the present invention. All of these fall within the protection scope of the present invention.
Claims
1. An OCR training sample generation method, characterized in that The input information is the original image and the coordinates of the erased area, and the method includes: A text contour extraction step, which extracts all text contours based on the original image, determines an erased area mask in combination with the coordinates of the erased area, and obtains a repaired area mask: Convert the input original image into a single-channel grayscale image, and then adaptively binarize it to obtain all text contour masks, where the text area value is 1 and the background area value is 0; According to the coordinates of the erased area, obtain the erased area mask, where the erased area value is 1 and the other area values are 0; Multiply the pixels at the corresponding positions of all text contour masks and the erased area mask to obtain an erased area text contour mask; Perform morphological dilation on the erased area text contour mask to obtain a repaired area mask; An image repair and filling step, which performs image repair and filling according to the repaired area mask and the pixel information around the repaired area to obtain a background template after the erased text: According to the repaired area mask, determine the area to be repaired in the original image; Poll each pixel point in the area to be repaired in the order from outside to inside, and calculate the pixel value that should be filled at the repair point according to the information of the known pixels around a certain pixel point, and become a known pixel; Calculate the pixel value of the next pixel point inward; Iterate step by step, and the area to be repaired gradually shrinks until the area to be repaired is all repaired, and a background template after the erased text that has been repaired is obtained; A random text generation step, which generates random text in each generation area, thereby obtaining a new sample image and a corresponding annotation information file: For a certain generation area, determine the expected length w of the random text for this generation area, set the font size to s, and estimate the number of characters n of this section of random text = int(w / s); Generate redundant random text with a length of n*k characters, where k is the redundancy multiple and takes a positive integer value; According to the relationship between the length of the redundant random text and the expected length w of the random text, determine the finally generated random text and its actual length L; Randomly determine the position of the generated text in the generation area, write the finally generated random text and determine its annotation information; Poll each generation area, thereby obtaining a new sample image and a corresponding annotation information file.
2. The OCR training sample generation method according to claim 1, wherein The sorting algorithm from outside to inside is the fast marching method.
3. The OCR training sample generation method according to claim 1, wherein The method for calculating the pixel value that should be filled at the repair point according to the information of the known pixels around a certain pixel point includes: weighted average of the known pixel values in the neighborhood, INPAINT_NS or INPAINT_TELEA method.
4. The OCR training sample generation method according to claim 1, wherein, The determination of the expected length w of the random text for this generation area specifically includes: Take the maximum value of the length and width of this generation area as the maximum length of the generated text, and the minimum value of the length and width as the minimum length of the generated text. Randomly select an integer between this minimum value and the maximum value as the expected length of the random text, denoted as w.
5. The OCR training sample generation method according to claim 1, wherein In the step of generating redundant random text with a length of n*k characters, if the specified corpus type is specified, generate redundant random text with a length of n*k characters in the specified corpus type; if the corpus type is not specified, randomly generate redundant random text with a length of n*k characters.
6. The OCR training sample generation method according to claim 5, wherein The generation of n*k character-length redundant random text based on the specified corpus type specifically includes: Regular expression generation: used to generate corpora with clear rules. Compile the required corpus rule features into regular expressions, and then randomly generate n*k character-length redundant random text according to the regular expressions; or Random acquisition from a database: used to generate corpora with unclear rules or certain fixed types. Obtain all the content of the corpus through public information, and store the data in a database or data file by item. When randomly generating, randomly obtain a piece of n*k character-length redundant random text from it.
7. The OCR training sample generation method according to claim 1, wherein The determination of the finally generated random text and its actual length L according to the relationship between the length of the redundant random text and the expected length w of the random text specifically includes: Starting from the first character of the redundant random text, count the actual length of each character and accumulate them in turn until the total character length just satisfies "adding one more character will exceed w". Take this as the finally generated random text, denoted as "text", and record the actual length L of the finally generated random text.
8. The OCR training sample generation method according to claim 1, wherein Randomly determining the position of the generated text within the generation area, writing the finally generated random text and determining its annotation information specifically includes: Assume that the upper left corner coordinates of the generation area are (x1, x2), and the lower right corner coordinates are (x2, y2). Then the x-axis coordinate range of the starting point of the upper left corner of the finally generated random text is [x1, (x2 - w)], and the y-axis coordinate range is [y1, (y2 - s)]. Randomly select an integer point (x, y) within this range as the starting position of the finally generated random text; On the background template, with (x, y) as the starting position of the upper left corner of the text, write the finally generated random text, and its annotation information is: the coordinates are [(x, y), (x + L, y), (x + L, y + s), (x, y + s)]; the text content is "text".
9. The OCR training sample generation method according to claim 1, wherein After the step of randomly determining the position of the generated text within the generation area, writing the finally generated random text and determining its annotation information, it further includes the step of adjusting the font size and color of the finally generated random text.
10. The OCR training sample generation method according to claim 9, wherein, The adjustment of the font size of the finally generated random text specifically includes: determining the width h of the erasure area, and default setting the size of the finally generated random text to h.
11. The OCR training sample generation method according to claim 9, wherein, The adjustment of the font color of the finally generated random text specifically includes: For a certain generation area, select the corresponding area from the repair area mask, and perform backbone extraction on the text contour of this area to form a text backbone area; Extract the corresponding area of the generation area from the original image. For the RGB three channels of this area, average all the pixel values at the text backbone area of each channel, which is the color value of each channel of the text backbone area; Take this color value as the color of the generated text, so that the font color of the finally generated random text is approximately the same as the original text color.
12. An OCR training sample generation device, the device operates based on the method according to any one of claims 1 to 11, and the device includes: A text contour extraction module, which is used to extract all text contours based on the original image, determine an erasure area mask in combination with the erasure area coordinates, and obtain a repaired area mask; An image repair and filling module, which is used to perform image repair and filling according to the repaired area mask and the pixel information around the repaired area, and obtain a background template after erasing the text; A random text generation module, which is used to generate random text in each generation area, thereby obtaining a new sample picture and a corresponding annotation information file.
13. An OCR training sample generation system, the system comprising: A processor and a memory for storing executable instructions; wherein, the processor is configured to execute the executable instructions to execute the OCR training sample generation method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by the processor, the OCR training sample generation method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Text recognition data synthesis method based on image fusion
CN112949754A