Text Detection Method and Device

By estimating the kernel density and layering the image, identifying and superimposing the text area of ​​each layer, the problem of missed detection in the prior art is solved, and efficient text detection and recognition is achieved.

CN111612005BActive Publication Date: 2025-06-24XIAN WANXIANG ELECTRONICS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010266335.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-04-07
Publication Date
2025-06-24
Estimated Expiration
2040-04-07

AI Technical Summary

Technical Problem

The prior art is prone to missed detection problems in image text recognition, resulting in poor processing effects.

Method used

By obtaining the kernel density estimation map of the image to be processed, the minimum value points are determined, the image is divided into multiple layers, and the text area is recognized and superimposed on each layer to generate the final target text image.

Benefits of technology

The text detection of the image to be processed is realized, avoiding the missed detection problem, and obtaining a complete and accurate text area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111612005B_ABST
    Figure CN111612005B_ABST
Patent Text Reader

Abstract

The present disclosure provides a text detection method and apparatus, relating to the technical field of image processing, and capable of solving the problem of missed detection in existing text recognition. The specific technical solution is as follows: obtaining a kernel density estimation map of an image to be processed; determining N minimum points in the kernel density estimation map; using the N minimum points as demarcation points to layer the image to be processed, obtaining N + 1 layers; identifying text regions in each of the N + 1 layers; and superimposing the text regions of each layer to obtain a target text image. The present invention is used for text detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a text detection method and apparatus. Background Art

[0002] Currently, when performing text recognition on an image, most traditional binarization methods use a single threshold to binarize the image to identify regions of interest. Traditional binarization methods include, for example, conventional binarization or adaptive binarization algorithms such as OSTU (Otsu method or maximum inter-class variance method). However, the above binarization methods are prone to missing detection problems, resulting in poor processing effects. Summary of the Invention

[0003] Embodiments of the present disclosure provide a text detection method and apparatus, which can solve the problem of missing detection in existing text recognition. The technical solutions are as follows:

[0004] According to a first aspect of the embodiments of the present disclosure, a text detection method is provided. The method includes:

[0005] Obtain a kernel density estimation map of the image to be processed;

[0006] Determine N minimum value points in the kernel density estimation map, where N≥1;

[0007] Using the N minimum value points as demarcation points, layer the image to be processed to obtain N + 1 layers;

[0008] Identify the text regions of each of the N + 1 layers;

[0009] Overlay the text regions of each layer to obtain a target text image.

[0010] By layering the image to be processed and identifying the text regions of each layer, a complete text region can be obtained, avoiding the situation of missing detection. After obtaining the text regions of each layer, overlay the text regions of each layer to generate the final target text graphic. Therefore, compared with the prior art, the text detection method provided by the embodiments of the present disclosure can implement text detection on the image to be processed while avoiding the problem of missing detection.

[0011] In one embodiment, identifying the text regions of each of the N + 1 layers includes:

[0012] Perform binarization processing on each of the N + 1 layers to obtain N + 1 binarized layers;

[0013] Identify the target regions of interest of each of the N + 1 binarized layers;

[0014] Identify the target region of interest for each binarized layer to obtain the text region for each layer.

[0015] In one embodiment, identifying the target region of interest for each binarized layer among N + 1 binarized layers includes:

[0016] Perform dilation and erosion on each binarized layer to obtain the initial region of interest for each binarized layer;

[0017] Filter the initial regions of interest for each binarized layer to obtain the target regions of interest for each binarized layer.

[0018] In one embodiment, filtering the initial regions of interest for each binarized layer to obtain the target regions of interest for each binarized layer includes:

[0019] Calculate the area of each initial region of interest in each binarized layer, where the area of the initial region of interest is the sum of all pixel counts within the initial region of interest;

[0020] Exclude the initial regions of interest in each binarized layer with an area smaller than a preset area threshold to obtain the target regions of interest for each binarized layer.

[0021] In one embodiment, filtering the initial regions of interest for each binarized layer to obtain the target regions of interest for each binarized layer includes:

[0022] Obtain the area and corresponding number of inflection points for each initial region of interest in each binarized layer, where an inflection point is used to indicate that the pixel value of a pixel point differs from the pixel values of surrounding pixels by more than a preset threshold;

[0023] Calculate the inflection point density for each initial region of interest based on the area and corresponding number of inflection points of each initial region of interest;

[0024] Determine the uniformity of the inflection point distribution in the corresponding initial region of interest based on the inflection point density of each initial region of interest;

[0025] Exclude the initial regions of interest in each binarized layer with non-uniform inflection point density distribution to obtain the target regions of interest for each binarized layer.

[0026] In one embodiment, determining the uniformity of the inflection point distribution in the corresponding initial region of interest based on the inflection point density of each initial region of interest includes:

[0027] Divide each initial region of interest into M sub-regions on average;

[0028] Calculate the sub-inflection point density for each sub-region among the M sub-regions of each initial region of interest;

[0029] Check whether the sub - inflection - point density of each sub - region is less than the corresponding inflection - point density;

[0030] When the sub - inflection - point densities of a preset number of sub - regions are less than the corresponding inflection - point densities, it is determined that the inflection - point distribution in the corresponding initial region of interest is uneven.

[0031] In one embodiment, screening the initial regions of interest of each binary layer to obtain the target regions of interest of each binary layer, including:

[0032] Obtain the contour of each initial region of interest in each binary layer and the minimum bounding rectangle corresponding to the contour of each third region of interest;

[0033] Calculate the ratio of the contour area of each initial region of interest to the area of the corresponding minimum bounding rectangle to obtain the rectangularity of each initial region of interest;

[0034] Eliminate the initial regions of interest in the binary layer with rectangularity less than a preset threshold to obtain the target regions of interest of the binary layer.

[0035] In one embodiment, superimposing the text regions of each layer to obtain a target text image, including:

[0036] Superimpose the text regions of each layer to obtain a superimposed text image;

[0037] When an overlapping region is detected in the superimposed text image, retain the layer corresponding to the text region with the largest area to obtain the target text image, and the overlapping region is used to indicate that the same part of the superimposed text image contains repeated text regions of different layers.

[0038] In one embodiment, the method further includes:

[0039] Perform post - processing on each text region in the target text image to eliminate mis - detected regions that do not meet the preset requirements, and obtain the final target text image. The preset requirement is to expand the connected domain where the text region is located by a preset step length and connect it to the connected domains of other text regions.

[0040] According to the second aspect of the embodiments of the present disclosure, there is provided a text detection device. The text detection device includes a processor and a memory. At least one computer instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the steps performed in the text detection method described in the first aspect and any embodiment of the first aspect.

[0041] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing at least one computer instruction, which is loaded and executed by a processor to implement the steps performed in the text detection method described in the first aspect and any embodiment of the first aspect.

[0042] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.

[0044] Figure 1 is a flowchart of a text detection method provided by an embodiment of the present disclosure;

[0045] Figure 2 is a flowchart of identifying the text area of each layer provided by an embodiment of the present disclosure;

[0046] Figure 3 is a flowchart of identifying a target region of interest provided by an embodiment of the present disclosure;

[0047] Figure 4 is a kernel probability density estimation trend graph of a grayscale histogram provided by an embodiment of the present disclosure;

[0048] Figure 5 is a schematic diagram of a minimum point provided by an embodiment of the present disclosure;

[0049] Figure 6 is an image without dilation and erosion provided by an embodiment of the present disclosure;

[0050] Figure 7 is an image after dilation and erosion provided by an embodiment of the present disclosure;

[0051] Figure 8 is a schematic diagram with uniform inflection point distribution provided by an embodiment of the present disclosure;

[0052] Figure 9 is a schematic diagram with non-uniform inflection point distribution provided by an embodiment of the present disclosure;

[0053] Figure 10 is a schematic diagram of an image to be processed provided by an embodiment of the present disclosure;

[0054] Figures 11 to 13 is an image after dilation and erosion of the binary layer obtained by layering the image to be processed provided by an embodiment of the present disclosure for Figure 10 ;

[0055] Figure 14 is the text image obtained by superimposing Figures 11 to 13 as described in the embodiments of the present disclosure;

[0056] Figure 15 is a schematic diagram for detecting misdetection areas provided by the embodiments of the present disclosure;

[0057] Figure 16 is a schematic diagram of the image to be processed provided by the embodiments of the present disclosure;

[0058] Figure 17 is the text image obtained after recognizing the image to be processed provided by the embodiments of the present disclosure;

[0059] Figure 18 is a structural diagram of a text detection device provided by the embodiments of the present disclosure;

[0060] Figure 19 is a structural diagram of a text detection device provided by the embodiments of the present disclosure;

[0061] Figure 20 is a structural diagram of a text detection device provided by the embodiments of the present disclosure;

[0062] Figure 21 is a structural diagram of a text detection device provided by the embodiments of the present disclosure;

[0063] Figure 22 is a structural diagram of a text detection device provided by the embodiments of the present disclosure. Detailed implementation manners

[0064] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0065] The embodiments of the present disclosure provide a text detection method, as Figure 1 shown, the text detection method includes the following steps:

[0066] 101. Obtain the kernel density estimation map of the image to be processed.

[0067] In the embodiments of the present disclosure, obtaining the kernel density estimation map of the image to be processed includes: obtaining the grayscale histogram of the image to be processed; using the kernel density estimation method to process the grayscale histogram to obtain the kernel density estimation map of the image to be processed. Specifically, when the image to be processed is a color picture, perform grayscale processing on the image to be processed to convert the image to be processed from a color picture into a grayscale image. Of course, denoising filtering can also be performed on the grayscale image. When the grayscale image of the image to be processed is obtained, according to the grayscale value of each pixel point in the grayscale image, draw the grayscale histogram of the image to be processed, and then use the kernel probability density estimation method to draw the trend of the grayscale histogram, that is, the kernel density estimation map of the image to be processed.

[0068] 102. Determine N minimum value points in the kernel density estimation map.

[0069] In the embodiments of the present disclosure, the kernel density estimation map is a function curve that can describe the trend of the grayscale histogram. Therefore, N minimum value points are obtained from the kernel density estimation curve. The minimum value point refers to a point where the function values on both sides of the point are greater than the function value corresponding to the point. It should be noted that the minimum value point is obtained by comparing the function value corresponding to the point with the function values of the points nearby, or by comparing the function value corresponding to the point with the function values of the points in the neighborhood of the point. It is the abscissa in a certain sub-interval of the function image, and does not mean the largest or smallest within the entire domain of the function.

[0070] 103. Take the N minimum value points as the dividing interfaces, and layer the image to be processed to obtain N + 1 layers.

[0071] In the embodiments of the present disclosure, the abscissa of the function curve of the kernel density estimation is the grayscale value, and the range of this grayscale value is the minimum grayscale value to the maximum grayscale value in the grayscale image of the image to be processed. Therefore, taking the N minimum value points as the boundaries, the pixels between the minimum grayscale value and the first minimum value point are determined as the first layer, the pixels between the first minimum value point and the second minimum value point are determined as the second layer, and so on. The pixels between the Nth minimum value point and the maximum grayscale value are determined as the N + 1th layer. In this way, N + 1 layers are separated from one image.

[0072] 104. Identify the text regions in each of the N + 1 layers.

[0073] In the embodiments of the present disclosure, as shown in Figure 2 , step 104 of identifying the text regions in each of the N + 1 layers includes:

[0074] 21. Perform binarization processing on each of the N + 1 layers to obtain N + 1 binarized layers.

[0075] 22. Identify the target region of interest (ROI) for each of the N + 1 binarized layers.

[0076] 23. Identify the target region of interest for each binarized layer to obtain the text region for each layer.

[0077] The binarization process for each layer includes: for the first layer, set the pixel values of the pixels between the minimum gray value and the first extreme point in the first layer to 255, and set the remaining pixel values to 0 to obtain the first binarized layer; for the second layer, set the pixel values between the first minimum point and the second minimum point in the second layer to 255, and set the remaining pixel values to 0 to obtain the second binarized layer, and so on. For the (N + 1)th layer, set the pixel values of the pixels between the Nth minimum point and the maximum gray value to 255, and set the remaining pixel values to 0 to obtain the (N + 1)th binarized layer. In this way, N + 1 binarized layers are formed.

[0078] After obtaining the N + 1 binarized layers, identify the target region of interest for each binarized layer. The region of interest (ROI) outlines the area to be processed from the processed image in the form of a rectangle, circle, ellipse, irregular polygon, etc. Refer to Figure 3 As shown, the steps for identifying the target region of interest for each binarized layer obtained in step 202 include the following:

[0079] 31. Perform dilation and erosion on each binarized layer to obtain the region of interest for each binarized layer.

[0080] Dilation and erosion are morphological operations, which are a series of image processing operations based on shapes. By dilation and erosion, the region of interest for each binarized layer can be obtained. It should be noted that each binarized layer contains at least one region of interest.

[0081] 32. Screen the regions of interest for each binarized layer to obtain the target region of interest for each binarized layer.

[0082] In the embodiments of the present disclosure, regions of interest that are not text regions can be excluded from the regions of interest by at least one of the area of the region of interest, the uniformity of the corner distribution, and the rectangularity of the region of interest. The following is a specific description for different situations.

[0083] In the first example, screening the regions of interest for each binarized layer to obtain the target region of interest for each binarized layer includes:

[0084] Calculate the area of each initial region of interest in each binarized layer. The area of the initial region of interest is the sum of the number of all pixels within the initial region of interest;

[0085] Eliminate the initial regions of interest with an area smaller than a preset area threshold in each binarized layer to obtain the target regions of interest for each binarized layer.

[0086] In this embodiment, by screening the areas of the regions of interest, it can be considered that the regions of interest with an area smaller than the area threshold are noise misdetections and are directly eliminated. Taking the area threshold as 30 pixels as an example, calculate the area of each initial region of interest in each binarized layer, compare the area of each initial region of interest with the area threshold of 30, eliminate the initial regions of interest with an area smaller than the area threshold of 30, and retain the initial regions of interest with an area greater than the area threshold of 30 as the target regions of interest.

[0087] In the second example, screen the regions of interest in each binarized layer to obtain the target regions of interest for each binarized layer, including:

[0088] Obtain the area of each initial region of interest and the corresponding number of inflection points in each binarized layer. The inflection point is used to indicate that the absolute value of the difference between the pixel value of a pixel point and the pixel values of surrounding pixels is greater than a preset threshold;

[0089] Calculate the inflection point density of each initial region of interest based on the area of each initial region of interest and the corresponding number of inflection points;

[0090] Determine the uniformity of the inflection point distribution in the corresponding initial region of interest according to the inflection point density of each initial region of interest;

[0091] Eliminate the initial regions of interest with uneven inflection point density distribution in each binarized layer to obtain the target regions of interest for each binarized layer.

[0092] Specifically, use the fast algorithm to detect the inflection points of each initial region of interest, count the number of inflection points within each region of interest, divide the number of inflection points within each initial region of interest by the area of the corresponding initial region of interest to obtain the inflection point density of each initial region of interest; then, divide each initial region of interest into M sub-regions on average, calculate the sub-inflection point density of each sub-region, and detect whether the sub-inflection point density of each sub-region is less than the corresponding inflection point density. If there are a preset number of sub-regions within an initial region of interest whose sub-inflection point density is less than the corresponding inflection point density, it is considered that the inflection point distribution of the initial region of interest is uneven, and the initial regions of interest with uneven inflection point distribution are eliminated, while the initial regions of interest with uniform inflection point distribution are retained.

[0093] In the third example, the initial regions of interest (ROIs) of each binarized layer are filtered to obtain the target ROIs of each binarized layer, including:

[0094] Obtain the contours of each initial ROI in each binarized layer and the minimum bounding rectangles corresponding to the contours of each third ROI;

[0095] Calculate the ratio of the contour area of each initial ROI to the area of the corresponding minimum bounding rectangle to obtain the rectangularity of each initial ROI;

[0096] Eliminate the initial ROIs in the binarized layer with rectangularity less than a preset threshold to obtain the target ROIs of the binarized layer.

[0097] Specifically, extract the contour of each initial ROI, calculate the minimum bounding rectangle of the contour, and then obtain the contour area of each initial ROI and the area of the corresponding minimum bounding rectangle. Define the ratio of the contour area to the area of the minimum bounding rectangle as the rectangularity. The smaller the rectangularity, the more irregular the shape of the initial ROI, and the less likely it is considered to be a text region. The closer the rectangularity is to 1, the closer the contour is to a rectangle. Therefore, the larger the rectangularity, the more likely the initial ROI is to be text. Therefore, compare the rectangularity of each initial ROI with the preset threshold, eliminate the initial ROIs with rectangularity less than the preset threshold, and retain the initial ROIs with rectangularity greater than the preset threshold as the target ROIs of each binarized layer.

[0098] It should be noted that for the elimination of ROIs, elimination can be performed from any one of the area of the ROI, the uniformity of the corner distribution, and the rectangularity of the ROI, or any two or three of the area of the ROI, the uniformity of the corner distribution, and the rectangularity of the ROI can be combined for elimination. By combining to eliminate the ROIs that are not text regions, a complete and accurate text region can be obtained, avoiding missed detection and false detection.

[0099] 105. Overlay the text regions of each layer to obtain the target text image.

[0100] In the embodiments of the present disclosure, since the text regions are scattered in various layers, it is necessary to superimpose and integrate the text regions in all layers to obtain the superimposed text graphics. However, there may be an overlap of multiple text regions in different layers in the same part of the superimposed text image. Therefore, the layer where the text region with the largest area is selected from these multiple text regions is used as the layer with the greatest contribution, that is, only the layer with the greatest contribution is retained, which can avoid multiple duplicate text regions in the same part of the superimposed text image. At the same time, all possible regions of interest can be effectively extracted, avoiding the situation of a large number of missed detections in the prior art.

[0101] The text detection method provided by the embodiments of the present disclosure divides the image to be processed into N + 1 layers according to N minimum points in the kernel density estimation map by obtaining the kernel density estimation map of the image to be processed, layers the image to be processed, and identifies the text regions in each layer. In this way, complete text regions can be obtained, avoiding the situation of missed detections. After obtaining the text regions in each layer, the text regions in each layer are superimposed to generate the final target text graphics. Therefore, compared with the prior art, the text detection method provided by the embodiments of the present disclosure can implement text detection of the image to be processed and avoid the problem of missed detections at the same time.

[0102] Based on the above Figure 1 corresponding embodiments of the text detection method, another embodiment of the present disclosure provides a text detection method. The text detection method provided in this embodiment is used for text detection in natural scenes, and specifically includes the following steps:

[0103] Step 1, perform grayscale processing and filtering and denoising processing on the image.

[0104] Specifically, grayscale processing refers to converting a color picture into a grayscale image.

[0105] Step 2, perform adaptive hierarchical threshold processing on the image.

[0106] First, according to the grayscale value (Y value) in the pixel value (YUV) of each pixel point in the image, draw the grayscale histogram of the image. Next, use the kernel probability density estimation method to draw the trend of the grayscale histogram, as Figure 4 shown by the black curve in the figure; finally, find the minimum points in the histogram. The minimum points refer to the points where the values on both sides of the point are greater than it, as Figure 5 shown by the black circles in the figure. There are three minimum points, so the image can be divided into four layers.

[0107] Step 3, take the minimum points as the boundaries, perform layering processing on the image; perform binarization processing on each layer to generate a binarized layer.

[0108] In this step, if there are n minimum values, the image is divided into n + 1 layers; and for each layer of the image, binarization processing is performed, thus obtaining n + 1 binarized layers.

[0109] Taking the minimum points as the boundaries, the specific steps for layering the image are as follows: For the first minimum point, the pixel values between 0 and the first extreme point are set to 255, and the remaining pixel values are set to 0, forming the first binarized layer; for the second extreme point, the pixel values between the first and second extreme points are set to 255, and the remaining pixel values are set to 0, forming the second binarized layer. Similarly, a total of 4 binarized layers are formed, similar to separating 4 layers from an image.

[0110] Step 4, in each binarized layer, use dilation and erosion operations to extract the regions of interest respectively, and identify the text regions from the regions of interest.

[0111] In this step, use dilation and erosion to connect the white regions in the binarized layer to generate the regions of interest, as Figure 6 and Figure 7 shown, Figure 6 is the image without dilation and erosion, Figure 7 is the image after dilation and erosion, Figure 7 and the white part in it is the region of interest.

[0112] Since text is a continuous region with many inflection points, and common text regions are not geometric shapes with extremely irregular shapes, therefore, in order to obtain a more accurate text region, the following processing can be performed on each binarized layer:

[0113] First, screen the area of the region of interest. It can be considered that the regions of interest with an area smaller than the area threshold are noise misdetections and are directly removed. Among them, the area threshold can be 30 pixels.

[0114] Second, within the region of interest, use the fast algorithm to detect inflection points (if the difference between a certain pixel and its surrounding pixels is too large, then this pixel point can be regarded as an inflection point); calculate the inflection point density within the region of interest; determine the uniformity of the inflection point distribution in the region of interest according to the inflection point density, and remove the regions of interest with uneven inflection point distribution.

[0115] Among them, the inflection point density is the number of inflection points in the region of interest divided by the area of the region of interest. To improve the calculation efficiency, a region of interest can be evenly divided into nine parts. If more than three parts have an inflection point density much lower than the average inflection point density, then it can be considered that the inflection point distribution in this region of interest is uneven, and this region of interest may be a relatively small picture or noise, and this region of interest cannot be regarded as a text region.

[0116] For example: As Figure 8 shown, the region of interest A is divided into nine parts, and there are a total of 18 inflection points in the entire region of interest A. Figure 8 The distribution of the inflection points of the region of interest A shown is uniform; as Figure 9 shown, the region of interest B is divided into nine parts, and there are a total of 18 inflection points in the entire region of interest B. Among them, the inflection point density of 4 parts is too small. It can be considered that the region of interest B is not a text region.

[0117] Third, calculate the rectangularity of the region of interest and eliminate the regions of interest with small rectangularity.

[0118] In this step, use the sobel operator to extract the contour of the region of interest and calculate the minimum bounding rectangle of the contour. Among them, the rectangularity is defined as the ratio of the contour area to the area of the minimum bounding rectangle. The smaller the rectangularity, the more irregular the shape of the region of interest, and the less likely it is considered that the region of interest is a text region. Then, the regions of interest with a rectangularity less than the preset threshold are not text regions.

[0119] First, use the open-source library opencv to calculate the minimum bounding rectangle and contour of the region of interest, avoid the repeated inclusion of the rectangular frame, and draw the minimum bounding envelope rectangle. Among them, the contour is of any shape, and the circumscribed envelope rectangle is a rectangle outside the region of interest; then, calculate the ratio of the areas of the two. The closer it is to 1, the closer the contour is to a rectangle, the greater the rectangularity, and the more likely the region of interest is text.

[0120] In this way, after the above-mentioned first to third steps of processing, the regions of interest that are not text regions can be eliminated from the regions of interest through the area of the region of interest, the uniformity of the corner distribution, and the rectangularity of the region of interest. The remaining regions of interest are most likely text regions.

[0121] Step 5: Stack the multiple layers processed in Step 4 to obtain an image containing multiple text regions; for each part of the image, only keep the layer with the largest contribution in the text region of that part.

[0122] Since the text regions may be scattered in each layer, in this step, the text regions in all layers can be integrated. The specific steps are as follows:

[0123] First, directly merge and stack each layer; since there may be multiple text regions in each part of the superimposed image, select the layer where the text region with the largest area is located as the layer with the largest contribution; finally, only keep the layer with the largest contribution in that part. This can avoid multiple overlapping text regions in the same part of the superimposed image.

[0124] It should be noted that this can effectively extract all possible regions of interest, avoiding the situation of a large number of missed detections in conventional binarization or adaptive binarization algorithms such as OSTU.

[0125] Illustrate by way of example. Figure 10 For the image to be processed, using the text detection method provided by the present invention, multiple binarized layers can be generated. After erosion and dilation, as Figures 11 to 13 shown, the superimposed image is as Figure 14 shown. It can be seen that the present invention can extract all possible text regions (regions of interest) in the figure, so there is no missed detection.

[0126] Step 6, exclude false detection regions by post-processing the image.

[0127] In a specific implementation, there may still be false detection regions in the image generated in step 5. Post-processing techniques can be used to check for false detection regions and delete isolated points.

[0128] Specifically, false detection regions usually refer to independent connected components. For connected components with a small area, they can be dilated outward by 10 times to determine whether they are connected to other connected components. If not, the connected component can be considered an isolated point and excluded.

[0129] Illustrate by way of example, as Figure 15 shown, a connected component is falsely detected in the lower left corner of the figure, marked with a thick black rectangle. The rectangular area can be dilated ten times, and it is found that the rectangular area will still not be connected to any blue rectangle. Therefore, the rectangular area can be regarded as an isolated point and deleted. In this way, text recognition can be performed within the recognized text region of the image.

[0130] Experimental results show that, as Figure 16 and Figure 17 shown, using the text detection method of the present invention, under the win10 system, with a CPU of i7 8550U, the time taken to recognize the text region is 10 - 20 ms. If the existing machine recognition method is used, under the same conditions, the time taken is about 100 ms.

[0131] The text detection method provided by the embodiments of the present disclosure first divides the image into multiple layers, performs binarization processing on each layer to obtain multiple regions of interest; then, according to the area of the regions of interest of each layer, the uniformity of the corner distribution, and the rectangularity of the regions of interest, the regions of interest that are not text regions are excluded from the multiple regions of interest of each layer. Finally, the regions of interest left are most likely to be text regions. It can be seen that the present invention can obtain complete and accurate text regions through binarization processing of each layer, avoiding missed detections.

[0132] In addition, since the finally obtained image is generated by superimposing multiple layers, there may be multiple superimposed text regions in a part of the image. For each part of the image, the present invention only retains the text region of the layer with the greatest contribution in that part, that is, that part only contains the text region with the largest area among the multiple text regions included in the multiple layers, so as to avoid multiple overlapping text regions in the same part of the superimposed image.

[0133] Based on the above Figures 1 to 3 The text detection method described in the corresponding embodiment, the following is an embodiment of the device of the present disclosure, which can be used to execute the method embodiment of the present disclosure.

[0134] The embodiment of the present disclosure provides a text detection device, as Figure 18 shown, the text detection device 180 includes: an acquisition module 1801, a determination module 1802, a layering module 1803, an identification module 1804, and a superimposition module 1805;

[0135] The acquisition module 1801 is used to acquire the kernel density estimation map of the image to be processed;

[0136] The determination module 1802 is used to determine N minimum value points in the kernel density estimation map, N≥1;

[0137] The layering module 1803 is used to layer the image to be processed with N minimum value points as demarcation points, and obtain N+1 layers;

[0138] The identification module 1804 is used to identify the text regions of each of the N+1 layers;

[0139] The superimposition module 1805 is used to superimpose the text regions of each layer to obtain the target text image.

[0140] In one embodiment, as Figure 19 shown, the identification module 1804 includes: a binarization processing sub-module 1901 and an identification sub-module 1902;

[0141] The binarization processing sub-module 1901 is used to perform binarization processing on each of the N+1 layers to obtain N+1 binarized layers;

[0142] The identification sub-module 1902 is used to identify the target region of interest of each of the N+1 binarized layers;

[0143] The identification sub-module 1902 is used to identify the target region of interest of each binarized layer to obtain the text region of each layer.

[0144] In one embodiment, as Figure 20As shown, the recognition sub-module 1902 includes: a dilation and erosion unit 2001 and a screening unit 2002;

[0145] The dilation and erosion unit 2001 is configured to perform dilation and erosion on each binarized layer to obtain an initial region of interest for each binarized layer;

[0146] The screening unit 2002 is configured to screen the initial regions of interest for each binarized layer to obtain a target region of interest for each binarized layer.

[0147] As Figure 21 shown, the screening unit 2002 includes: a calculation subunit 2101, a rejection subunit 2102, an acquisition subunit 2103, and a determination subunit 2104;

[0148] In one embodiment, the calculation subunit 2101 is configured to calculate the area of each initial region of interest in each binarized layer, and the area of the initial region of interest is the sum of the number of all pixels within the initial region of interest;

[0149] The rejection subunit 2102 is configured to reject the initial regions of interest in each binarized layer with an area smaller than a preset area threshold to obtain a target region of interest for each binarized layer.

[0150] In one embodiment, the acquisition subunit 2103 is configured to acquire the area of each initial region of interest and the corresponding number of inflection points for each binarized layer, and an inflection point is used to indicate that the pixel value difference between a pixel point and the pixel values of surrounding pixels is greater than a preset threshold;

[0151] The calculation subunit 2101 is configured to calculate the inflection point density of each initial region of interest based on the area of each initial region of interest and the corresponding number of inflection points;

[0152] The determination subunit 2104 is configured to determine the uniformity of the inflection point distribution in the corresponding initial region of interest based on the inflection point density of each initial region of interest;

[0153] The rejection subunit 2102 is configured to reject the initial regions of interest in each binarized layer with uneven inflection point density distribution to obtain a target region of interest for each binarized layer.

[0154] In one embodiment, the determination subunit 2104 is configured to evenly divide each initial region of interest into M sub-regions; calculate the sub-inflection point density of each sub-region in the M sub-regions of each initial region of interest; detect whether the sub-inflection point density of each sub-region is less than the corresponding inflection point density; and determine that the inflection point distribution in the corresponding initial region of interest is uneven when there are a preset number of sub-regions with sub-inflection point density less than the corresponding inflection point density.

[0155] In one embodiment, an acquisition subunit 2103 is configured to acquire the contour of each initial region of interest in each binarized layer and the minimum bounding rectangle corresponding to the contour of each third region of interest.

[0156] A calculation subunit 2101 is configured to calculate the rectangularity of each initial region of interest according to the ratio of the contour area of each initial region of interest to the area of the corresponding minimum bounding rectangle.

[0157] An elimination unit 2102 is configured to eliminate the initial regions of interest in the binarized layer with rectangularity less than a preset threshold, so as to obtain the target regions of interest of the binarized layer.

[0158] In one embodiment, a superimposition module 1805 is configured to superimpose the text regions of each layer to obtain a superimposed text image; when an overlapping region is detected in the superimposed text image, keep the layer corresponding to the text region with the largest area to obtain a target text image, and the overlapping region is used to indicate that the same part of the superimposed text image contains repeated text regions of different layers.

[0159] In one embodiment, as Figure 22 shown, the text detection device 180 further includes: a post-processing module 1806;

[0160] The post-processing module 1806 is configured to perform post-processing on each text region in the target text image, and eliminate the misdetected regions that do not meet the preset requirements, so as to obtain the final target text image, and the preset requirement is to expand the connected domain where the text region is located by a preset step length and connect it to the connected domains of other text regions.

[0161] The text detection device provided by the embodiments of the present disclosure divides the image to be processed into N + 1 layers according to N minimum value points in the kernel density estimation map of the image to be processed by obtaining the kernel density estimation map of the image to be processed, hierarchically processes the image to be processed, and identifies the text regions of each layer. In this way, complete text regions can be obtained, and the situation of missed detection can be avoided. After obtaining the text regions of each layer, the text regions of each layer are superimposed to generate the final target text graphic. Therefore, compared with the prior art, the text detection method provided by the embodiments of the present disclosure can implement text detection on the image to be processed and avoid the problem of missed detection at the same time.

[0162] The embodiments of the present disclosure further provide a text detection device, which includes a receiver, a transmitter, a memory, and a processor. The transmitter and the memory are respectively connected to the processor, and at least one computer instruction is stored in the memory. The processor is configured to load and execute at least one computer instruction to implement the Figures 1 to 3 text detection method described in the corresponding embodiment above.

[0163] Based on the above Figures 1 to 3 For the text detection method described in the corresponding embodiment, the embodiments of the present disclosure also provide a computer-readable storage medium. For example, a non-transitory computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc. Computer instructions are stored on the storage medium for executing the above Figures 1 to 3 text detection method described in the corresponding embodiment, which will not be elaborated here.

[0164] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware or by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0165] After considering the specification and practicing the disclosure herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

Claims

1. A text detection method, characterized in that, The method includes: Obtaining a kernel density estimation map of the image to be processed; Determining N minimum value points in the kernel density estimation map, where N≥1; Taking the N minimum value points as demarcation points to layer the image to be processed, obtaining N+1 image layers; Identifying the text regions of each of the N+1 image layers; Overlaying the text regions of each image layer to obtain a target text image; The identifying the text regions of each of the N+1 image layers includes: performing binarization processing on each of the N+1 image layers to obtain N+1 binarized image layers; identifying the target regions of interest of each of the N+1 binarized image layers; identifying the text regions of each image layer by identifying the target regions of interest of each binarized image layer; The identifying the target regions of interest of each of the N+1 binarized image layers includes: performing dilation and erosion on each binarized image layer to obtain the initial regions of interest of each binarized image layer; screening the initial regions of interest of each binarized image layer to obtain the target regions of interest of each binarized image layer; The screening the initial regions of interest of each binarized image layer to obtain the target regions of interest of each binarized image layer includes: obtaining the area and the corresponding number of inflection points of each initial region of interest of each binarized image layer, where the inflection point is used to indicate that the pixel value difference between a pixel point and the pixel values of surrounding pixels is greater than a preset threshold; calculating the inflection point density of each initial region of interest according to the area and the corresponding number of inflection points of each initial region of interest; determining the uniformity of the inflection point distribution in the corresponding initial region of interest according to the inflection point density of each initial region of interest; removing the initial regions of interest with uneven inflection point density distribution in each binarized image layer to obtain the target regions of interest of each binarized image layer.

2. The method according to claim 1, characterized in that, The screening the initial regions of interest of each binarized image layer to obtain the target regions of interest of each binarized image layer includes: Calculating the area of each initial region of interest in each binarized image layer, where the area of the initial region of interest is the sum of the number of all pixels in the initial region of interest; Removing the initial regions of interest with an area smaller than a preset area threshold in each binarized image layer to obtain the target regions of interest of each binarized image layer.

3. The method according to claim 1, characterized in that, The determining the uniformity of the inflection point distribution in the corresponding initial region of interest according to the inflection point density of each initial region of interest includes: Dividing each initial region of interest into M sub-regions on average; Calculating the sub-inflection point density of each sub-region in the M sub-regions of each initial region of interest; Detecting whether the sub-inflection point density of each sub-region is less than the corresponding inflection point density; When there are a preset number of sub-regions with sub-inflection point density less than the corresponding inflection point density, determining that the inflection point distribution in the corresponding initial region of interest is uneven.

4. The method according to claim 1, wherein The screening the initial regions of interest of each binarized image layer to obtain the target regions of interest of each binarized image layer includes: Obtain the contour of each initial region of interest in each binarized layer and the minimum bounding rectangle corresponding to the contour of each initial region of interest; Calculate the ratio of the contour area of each initial region of interest to the area of the corresponding minimum bounding rectangle to obtain the rectangularity of each initial region of interest; Remove the initial regions of interest in the binarized layer whose rectangularity is less than a preset threshold to obtain the target regions of interest in the binarized layer.

5. The method according to claim 1, wherein The step of obtaining the target text image by superimposing the text regions of each layer includes: Superimpose the text regions of each layer to obtain a superimposed text image; When an overlapping region is detected in the superimposed text image, retain the layer corresponding to the text region with the largest area to obtain the target text image, and the overlapping region is used to indicate that the same part of the superimposed text image contains repeated text regions of different layers.

6. The method according to claim 1, wherein The method further includes: Perform post-processing on each text region in the target text image to remove misdetected regions that do not meet the preset requirements to obtain the final target text image, where the preset requirement is that the connected component where the text region is located is expanded by a preset step length and then connected to the connected components of other text regions.

7. A text detection device, characterized in that, The text detection device includes a processor and a memory, and at least one computer instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the steps performed in the text detection method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Chinese model train, Chinese image recognition method, device, equipment and medium

    CN109102037A

  • Character detection method and device, apparatus and computer readable storage medium

    CN110097046A