OCR Graphic and Text Recognition and Statistics Method for Data Collector
Through the analysis of edge lines and thick ink pixel growth area, combined with corner point gradient features, motion blur direction estimation and feature descriptor are constructed, which solves the problems of low accuracy and high computational complexity of OCR graphic and text recognition in motion blur scenes by the data collector, and realizes efficient graphic and text recognition and character statistics.
Patent Information
- Application Number
- CN202411938261.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-12-26
AI Technical Summary
Existing data collectors have problems with low recognition accuracy and high computational complexity in OCR graphics and text recognition, especially in motion fuzzy scenarios, which are difficult to meet real-time requirements.
The edge line and thick ink pixel growth area analysis were used, and the corner point gradient characteristics were combined to construct motion fuzzy direction estimation and feature descriptors, and OCR graphic recognition and character statistics were achieved through key point matching.
It improves the accuracy of OCR graphic and text recognition in motion blur scenes, reduces the length of feature descriptors, and improves the recognition speed and the accuracy of character statistics.
Smart Images

Figure CN119863798B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of graphic and text recognition technology, and specifically relates to an OCR graphic and text recognition and statistical method for data collectors. Background Art
[0002] Data collectors are usually widely used in fields such as logistics, commerce, and inventory management. Data collectors can not only scan and identify barcodes and QR codes, but also perform graphic and text recognition and character quantity statistics when encountering unbarcoded characters, dates, and texts, which can save labor, avoid manual entry errors, and achieve fast recognition and reading of graphic and text information.
[0003] The OCR recognition method based on deep learning extracts deep features of graphics and texts through a neural network model, has wide applicability to complex scenarios, and has a high recognition accuracy. However, to implement OCR recognition, a large number of graphic and text data samples need to be collected, and it has a strong dependence on the hardware of the data collector, and the technical implementation cost is high.
[0004] The SIFT (Scale-Invariant Feature Transform) algorithm extracts key points and feature descriptors of text features in graphics and texts, and realizes OCR graphic and text recognition through feature matching. However, the usage scenarios of data collectors are complex, and the collected images are prone to motion blur. Too few key points are detected by the SIFT algorithm, resulting in a high error rate of graphic and text recognition. Moreover, the SIFT algorithm generates 128-dimensional feature descriptors for key points of features, with a high computational complexity and unable to meet the real-time performance of OCR graphic and text recognition for data collectors. Summary of the Invention
[0005] In order to solve the above technical problems, this application provides an OCR graphic and text recognition and statistical method for data collectors to solve the existing problems.
[0006] The OCR graphic and text recognition and statistical method for data collectors of this application adopts the following technical solutions:
[0007] An embodiment of this application provides an OCR graphic and text recognition and statistical method for data collectors, including the following steps:
[0008] Perform graphic and text image acquisition through a data collector to obtain a graphic and text grayscale image, and perform graphic and text character segmentation on the graphic and text grayscale image to obtain each text detection frame;
[0009] Extract the edge lines and corner points in the graphic and text grayscale image, analyze the distribution of edge lines and corner points in each text detection frame, and obtain each trend direction and its trend direction significance value;
[0010] Perform threshold segmentation on the grayscale image of the text and graphics to obtain thick ink pixel points, extract thick ink strokes through the region growing algorithm, and determine the thick ink stroke directions of each growing region according to the directions of the thick ink strokes;
[0011] Based on the difference in the number of growing regions in the horizontal and vertical directions of the thick ink stroke directions, construct the confidence of the horizontal and vertical motion blur directions of the grayscale image of the text and graphics, and estimate the motion blur direction of the grayscale image of the text and graphics in combination with the number of growing regions in each thick ink stroke direction;
[0012] According to the estimated result of the motion blur direction of the grayscale image of the text and graphics, and in combination with the gradient direction, gradient amplitude of each corner point and the path information of the thick ink pixel points in the motion blur direction, construct the motion blur eigenvalue of each corner point with respect to each gradient direction;
[0013] According to the motion blur eigenvalue and the distribution of thick ink pixel points in the neighborhood of each corner point, screen the corner points to obtain the key points for feature matching of each text detection box, and in combination with the motion blur direction and the number of thick ink pixel points in the neighborhood of the key points, construct the feature descriptor of the key points in the motion blur direction, use the key point matching algorithm to automatically recognize the OCR text and graphics of the data collector, and count the number of characters through the number of text detection boxes.
[0014] Preferably, the obtaining of each trend direction further includes:
[0015] Statistically analyze the ratio of the number of edge pixel points in each text detection box to the total number of pixel points in the text detection box, and arrange the text detection boxes in ascending order according to the ratio to obtain the first preset number of text detection boxes, denoted as the detection boxes to be analyzed;
[0016] Through the corner points in the detection boxes to be analyzed, segment the edge lines in the detection boxes to be analyzed to obtain each edge segment, perform linear fitting on each edge segment, and calculate the included angle value between the fitted lines and the horizontal direction as each trend direction.
[0017] Preferably, the method for obtaining the trend direction significant value of each trend direction is: statistically analyze the number of fitted lines in the same trend direction and use it as the trend direction significant value of the same trend direction.
[0018] Preferably, the obtaining of the thick ink pixel points further includes:
[0019] Perform the first threshold segmentation on the grayscale image of the text and graphics, output the first segmentation threshold, take the pixel points with gray values lower than the first segmentation threshold as stroke pixel points, and perform the second threshold segmentation on the gray values of all stroke pixel points, output the second segmentation threshold, and take the pixel points with gray values lower than the second segmentation threshold as thick ink pixel points.
[0020] Preferably, the method for determining the thick ink stroke direction is as follows: perform skeleton extraction on each growth region, perform linear fitting on the skeleton, and determine the angle between the fitted line and the horizontal direction as the thick ink stroke direction of each growth region.
[0021] Preferably, the expression for the confidence level of the horizontal and vertical motion blur directions of the graphic grayscale image is:
[0022] In the formula, T is the confidence level of the horizontal and vertical motion blur directions of the graphic grayscale image, μ1 is the number of growth regions with a thick ink stroke direction of 90°, μ2 is the number of growth regions with a thick ink stroke direction of 180°, max() is the maximum value function, and M is the number of growth regions.
[0023] Preferably, further estimating the motion blur direction of the graphic grayscale image includes:
[0024] When the confidence level of the horizontal and vertical motion blur directions is greater than or equal to the preset confidence threshold, select the direction with the largest number of growth regions from the horizontal and vertical directions as the motion blur direction of the graphic grayscale image;
[0025] When the confidence level of the horizontal and vertical motion blur directions is less than the preset confidence threshold, calculate the motion blur direction certainty values of all trend directions except the horizontal and vertical directions, and use the trend direction corresponding to the largest motion blur direction certainty value as the motion blur direction of the graphic grayscale image;
[0026] Among them, the expression for the motion blur direction certainty value is: In the formula, R θ , n θ are respectively the motion blur direction certainty value and the trend direction significance value of the trend direction θ, is the motion blur direction weight when the thick ink stroke direction is θ, where the proportion of the number of growth regions in each thick ink stroke direction in the total number of growth regions in all thick ink stroke directions is used as the motion blur direction weight of each thick ink stroke direction.
[0027] Preferably, the expression for the motion blur eigenvalue of each corner point with respect to each gradient direction is:
[0028] In the formula, v q,x is the motion blur eigenvalue of the corner point q with respect to the gradient direction x, φ is the motion blur direction of the graphic grayscale image, Cos() is the cosine value, J is the number of gradient directions, h q,x is the gradient amplitude of the corner point q in the gradient direction x, and h q,j is the gradient amplitude of the corner point q in the gradient direction j.
[0029] Preferably, the screening of the key points includes:
[0030] Obtain a straight line with the direction of motion blur and passing through the corner point within the preset neighborhood of each corner point, and count the number of thick ink pixel points passed by the straight line passing through the corner point within the preset neighborhood of the corner point as the thick ink pixel length of the corner point;
[0031] Take the product of the maximum value of the motion blur feature value and the thick ink pixel length of each corner point as the screening robustness strength of each corner point. Arrange all the corner points in the grayscale image of the graphic according to the screening robustness strength from large to small, and select the top preset number of corner points as key points.
[0032] Preferably, the construction of the feature descriptor of the key points in the direction of motion blur includes:
[0033] Normalize the motion blur direction of the grayscale image of the graphic, the blur representation direction of each key point, the motion blur feature value, and the thick ink pixel length respectively, and use the vector composed of the normalized result values as the feature descriptor of each key point in the direction of motion blur.
[0034] This application has at least the following beneficial effects:
[0035] This application estimates the motion blur direction of the grayscale image of the graphic through multi-scale analysis based on the edge and the growth area of thick ink pixel points. The beneficial effect is that the printed text is horizontal and vertical in form. In addition to the edge lines whose trend direction is the same as the motion blur direction, there are also a relatively large number of horizontal and vertical edge lines, excluding the influence of horizontal and vertical edge lines on the estimation of the motion blur direction, and improving the accuracy of the motion blur direction estimation;
[0036] At the same time, according to the estimation result of the motion blur direction of the grayscale image of the graphic, this application screens the corner points according to the gradient feature and the stroke trend feature in the direction of motion blur in the neighborhood of the corner points, and realizes the automatic OCR graphic recognition and character number statistics of the data collector. The beneficial effect is to expand the key points of the SIFI algorithm, reduce the false detection rate, and at the same time mine the key points with significant features in the direction of motion blur, which has high applicability to the motion blur scene of the data collector, and the length of the feature descriptor is reduced, improving the speed of key point matching by the SIFI algorithm. Description of the Drawings
[0037] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0038] Figure 1 This is a flowchart of the steps of the OCR graphic and text recognition and statistics method for a data collector provided by this application. Specific embodiments
[0039] In order to further elaborate on the technical means and effects adopted by this application to achieve the intended invention purpose, the following combines the accompanying drawings and preferred embodiments to specifically describe the specific implementation manner, structure, features and effects of the OCR graphic and text recognition and statistics method for a data collector proposed according to this application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0040] Unless otherwise defined, terms such as "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a circuit structure, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the article or device including the said element. In addition, the term "and / or" used herein includes any and all combinations of one or more of the related listed items. All technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.
[0041] The following specifically describes the specific solution of the OCR graphic and text recognition and statistics method for a data collector provided by this application in conjunction with the accompanying drawings.
[0042] The OCR graphic and text recognition and statistics method for a data collector provided by an embodiment of this application, specifically, please refer to Figure 1 , including the following steps:
[0043] Step 1: Perform graphic and text image acquisition through a data collector to obtain a graphic and text grayscale image, and perform graphic and text character segmentation on the graphic and text grayscale image to obtain each text detection box.
[0044] In this embodiment, a data collector is used to obtain an image. Specifically, either the rear high-definition camera of the data collector can be used to perform image acquisition on the text to be detected, or the wireless communication method of the data collector can be used to upload the image. Perform bilateral filtering processing on the image obtained by the data collector, aiming to eliminate the interference of noise on the image and improve the accuracy of OCR graphic and text recognition. Convert the denoised image to a graphic and text grayscale image.
[0045] Data collectors are widely used in fields such as logistics, commerce, and inventory management. Usually, they perform automatic OCR graphic and text recognition on item labels. The text on item labels belongs to the category of printed characters with a white background and black characters, and has the characteristics of high standardization and neat arrangement. Therefore, in this embodiment, the graphic grayscale image is first rotationally corrected, and then the MSER (Maximally Stable Extremal Regions) algorithm is used to perform preliminary detection box calibration on the graphic grayscale image. The preliminary detection box calibration is specifically implemented using cv2.MSER in the Python language. Both the rotational correction and the MSER algorithm are well-known technologies, and the specific process will not be elaborated here.
[0046] Due to the various forms of the text, the detection boxes obtained using the MSER algorithm are irregular, and there are many overlapping detection boxes. Some text detection boxes may only enclose the radicals or strokes of Chinese characters. Therefore, in this embodiment, the minimum bounding rectangle of each text detection box in the graphic grayscale image is obtained using cv2.boundingRect in the Python language, and then the NMS (Non Maximum Suppression) algorithm is used to suppress the boxes that are not the largest size and remove the overlapping regions. In this embodiment, the overlapping area (IoU) is set to 0.3, and the text detection boxes of each text in the graphic grayscale image are output. Each text detection box contains a printed character with a white background and black characters, realizing the segmentation of graphic characters.
[0047] Step 2: Extract the edge lines and corner points in the graphic grayscale image, analyze the distribution of edge lines and corner points within each text detection box, and obtain each trend direction and its trend direction significance value.
[0048] During the process of the data collector acquiring images, due to the complex usage scenarios of the data collector, factors such as hand jitter and object movement are likely to occur during the shooting process, resulting in motion blur in the acquired images, causing the outlines of graphic characters to be unclear and the details to be lost. The motion blur direction of the graphic grayscale image reflects the relative motion direction of the object. The strokes of printed fonts are regular and have a high degree of directionality, and the applicability of traditional motion blur direction estimation methods is poor. Motion blur will cause the clear strokes of the graphic to appear blurred, and the background of the graphic characters is usually a white label paper. These blurred strokes have a relatively high contrast with the white label paper, and it is easy to form edge lines with the same trend direction as the motion blur direction.
[0049] In this embodiment, the graphic grayscale image is used as the input of the Canny edge detection algorithm, and all the edge lines of the graphic grayscale image are output. Since the more edge lines there are inside a graphic character, the more complex the stroke structure features of the character are, and the strokes in different trend directions are more likely to interfere with each other, disturbing the subsequent estimation of the motion blur direction. In this embodiment, the ratio of the number of edge line pixel points in each text detection box in the graphic grayscale image to the total number of pixel points in the text detection box is statistically calculated, and the text detection boxes are arranged in ascending order according to the ratio, and the first preset number of text detection boxes are obtained, denoted as the detection boxes to be analyzed. Here, the preset number of text detection boxes is taken as 10 in this embodiment. In actual application scenarios, the implementer selects it by himself / herself, and no special limitation is made in this embodiment.
[0050] Since the edge lines may include stroke inflection points and folding points, this embodiment uses the Harris corner detection algorithm to obtain all the corner points in the graphic grayscale image. Through the corner points in the detection boxes to be analyzed, the edge lines in the detection boxes to be analyzed are segmented to obtain each edge segment, and linear fitting is performed on each edge segment, and the included angle value between the fitted straight line and the horizontal direction is calculated as each trend direction. In this embodiment, linear fitting can be performed by the least square method. Further, for each trend direction, the number of straight lines in the same trend direction is statistically calculated and used as the trend direction significance value corresponding to the trend direction.
[0051] Step 3: Perform threshold segmentation on the graphic grayscale image to obtain thick ink pixel points, extract thick ink strokes through the region growing algorithm, and determine the thick ink stroke directions of each growth region according to the trend of the thick ink strokes.
[0052] However, in the graphic grayscale image, the printed characters have high standardization and are horizontally and vertically straight in form. In addition to the edge lines with the same trend direction and motion blur direction, there are also a relatively large number of horizontal (motion blur direction is 180°) and vertical (motion blur direction is 90°) edge lines. In the part of the strokes in the graphic grayscale image that are in the same direction as the motion blur direction, there is not much ghosting, the printed ink is thick, and the gray value is low, while the ghosted part of the text strokes is blurred and the gray value is slightly higher.
[0053] In this embodiment, the graphic grayscale image is used as the input of the first Ostu algorithm, and the first segmentation threshold is output. The pixel points with gray values higher than the first segmentation threshold are white background pixel points, and the pixel points with gray values lower than the first segmentation threshold are stroke pixel points. And the gray values of all pixel points lower than the first segmentation threshold are used as the input of the second Ostu algorithm, and the second segmentation threshold is output. The pixel points with gray values higher than the second segmentation threshold are ghost pixel points, and the pixel points with gray values lower than the second segmentation threshold are thick ink pixel points.
[0054] Furthermore, in this embodiment, the strokes with relatively thick printed ink after motion blur will be extracted and denoted as thick-ink strokes. The thick-ink pixel points are the pixel points on the thick-ink strokes, and the adjacent thick-ink pixel points together form the thick-ink strokes. The specific process of obtaining the thick-ink strokes in this embodiment is as follows: Obtain each text detection frame containing the detection frame to be analyzed. Use region growing to obtain each growing region, which is used to represent the thick-ink strokes in the text and image grayscale map. The specific region growing process in this embodiment is as follows:
[0055] S1: Randomly select a thick-ink pixel point as the initial center of region growing within each text detection frame containing the detection frame to be analyzed;
[0056] S2: Use the region growing algorithm. If there are thick-ink pixel points in the eight-neighborhood of the initial center, then merge and grow, and finally obtain the growing region;
[0057] S3: Within each detection frame to be analyzed, re-select a thick-ink pixel point outside the growing region as the initial center of region growing, repeat step S2, and continue region growing to obtain a new round of growing regions;
[0058] S4: Repeat the above steps until all thick-ink pixel points within the detection frame to be analyzed are within the growing region and stop, or when the number of growing regions is equal to M, also stop region growing. M is taken as 15 in this embodiment. It should be noted that the region growing algorithm is a prior art, and the specific growing process is not described in this embodiment.
[0059] Furthermore, obtain the direction of the thick-ink strokes. Specifically, in this embodiment, the growing region is used as the input of the Zhang-Suen skeleton extraction algorithm. Since the growing region is approximately a connected domain of thick-ink strokes, the skeleton of the growing region is output, and the straight line fitting is performed on the skeleton. The included angle between the fitted straight line and the horizontal direction is used as the direction of the thick-ink strokes of the growing region.
[0060] Step 4: According to the difference in the number of growing regions of the thick-ink stroke directions in the horizontal and vertical directions, construct the horizontal and vertical motion blur direction confidence degrees of the text and image grayscale map, and estimate the motion blur direction of the text and image grayscale map by combining the number of growing regions of each thick-ink stroke direction.
[0061] Since the stroke parts in the text and image grayscale map that are consistent with the motion blur direction belong to thick-ink pixel points only after motion blur, and there are many horizontal and vertical strokes in the text and image grayscale map. If the motion blur direction is 90° or 180°, there should also be a large number of growing regions with thick-ink stroke directions of 90° or 180°.
[0062] Therefore, based on the above analysis, construct the horizontal and vertical motion blur direction confidence degree T of the text and image grayscale map. In this embodiment, the calculation formula is:
[0063] Where, μ1 is the number of growth regions with a thick ink stroke direction of 90°, μ2 is the number of growth regions with a thick ink stroke direction of 180°, max() is the maximum value function, and M is the number of growth regions. If the motion blur direction of the text grayscale image is 90° or 180°, the text strokes in the other direction will cause ghosting. When |μ1-μ2| is larger, it means that the direction of the thick ink stroke in the growth region is only biased towards one of 90° or 180°. At the same time, when max(μ1,μ2) is larger, it means that the direction of the thick ink stroke of 90° or 180° is the most complete, and the motion blur direction of the text grayscale image is more likely to be horizontal or vertical.
[0064] This embodiment calculates the proportion of the number of growth areas in each dark ink stroke direction in the total number of growth areas in all dark ink stroke directions as the motion blur direction weight of each dark ink stroke direction. The larger the motion blur direction weight, the more likely it is that the dark ink strokes composed of dark ink pixels in the grayscale image of the image and text are consistent with the motion blur direction.
[0065] At this point, this embodiment obtains the trend direction significance value n of each trend direction of the image and text grayscale image, the horizontal and vertical motion blur direction confidence T of the image and text grayscale image, and the motion blur direction weight of each thick ink stroke direction
[0066] When the confidence T of the horizontal and vertical motion blur direction of the image and text grayscale image is greater than or equal to the confidence threshold t, the motion blur direction of the image and text grayscale image is more likely to be horizontal or vertical, where the confidence threshold is 0.7 in this embodiment. The direction containing the largest number of growth areas is selected from the horizontal and vertical directions as the motion blur direction of the image and text grayscale image.
[0067] When the confidence T of the horizontal and vertical motion blur direction of the image grayscale image is less than the confidence threshold t, the motion blur direction of the image grayscale image is less likely to be horizontal or vertical, and the interference of the horizontal and vertical edge lines in the trend direction histogram should be excluded. After excluding the horizontal and vertical directions, any trend direction is recorded as θ, and the motion blur direction confidence value R of the trend direction θ is calculated as follows θ In this embodiment, the specific calculation formula is:
[0068] In the formula, is the motion blur direction weight when the direction of the thick ink stroke is θ, n θ is the trend direction significance value of the trend direction θ. θ The larger the value, the more likely it is that the edge line trend direction of the grayscale image is consistent with the motion blur direction. The larger it is, it indicates that on the premise of excluding the horizontal and vertical directions, the thick ink strokes composed of thick ink pixels are more likely to be consistent with the motion blur direction, and the more likely θ is the motion blur direction of the text and image grayscale map.
[0069] According to the above process, after calculating and excluding the horizontal and vertical directions, calculate the motion blur direction certainty values of all other trend directions, and take the trend direction corresponding to the maximum motion blur direction certainty value as the motion blur direction of the text and image grayscale map.
[0070] So far, according to the above process of this embodiment, the estimation of the motion blur direction of the text and image grayscale map is realized.
[0071] Step 5: According to the estimation result of the motion blur direction of the text and image grayscale map, and in combination with the gradient direction, gradient amplitude of each corner point, and the path information of the thick ink pixels in the motion blur direction, construct the motion blur feature value of each corner point with respect to each gradient direction.
[0072] Motion blur causes the character contours of the text and image grayscale map obtained by the data collector to be unclear and the details to be lost, which easily leads to too few key points output by the SIFT algorithm, resulting in misjudgment of text characters as graphic characters with similar structures, and the OCR text and image recognition accuracy is low. The corner points in the text and image grayscale map have a certain robustness under the conditions of scaling and illumination changes, which can make up for the shortage of the number of key points of the SIFI algorithm, but it is easy to misdetect redundant corner points, and the corner points need to be further screened.
[0073] Considering that the key points should have strong robustness to motion blur, the text strokes consistent with the motion blur direction can retain a large amount of text stroke detail information. In this embodiment, the Sobel operator is used to calculate the gradient direction and gradient amplitude of any pixel point in the text and image grayscale map. Further, in combination with the gradient direction, gradient amplitude of each corner point, and the path information of the thick ink pixels in the motion blur direction, the motion blur feature value of any corner point with respect to the gradient direction x can be obtained. In this embodiment, the specific calculation formula is:
[0074] In the formula, v q,x is the motion blur feature value of the corner point q with respect to the gradient direction x, φ is the motion blur direction of the text and image grayscale map. Since the gradient direction is the direction in which the pixel point brightness changes fastest, in this embodiment, the complementary angle of the gradient direction is the stroke direction, Cos() is the cosine value, h q,x is the gradient amplitude of the corner point q in the gradient direction x, h q,j is the gradient amplitude of the corner point q in the gradient direction j, the lower limit of j is 1°, and J is the number of gradient directions.
[0075] It can be understood that The larger it is, the more prominent the edge trend feature of the corner point q in the gradient direction x is. At the same time, when Cos(180 - x, φ) is larger, it indicates that the gradient direction of the text stroke in the neighborhood of the corner point q and the motion blur direction are more likely to be perpendicular. Then, the text stroke direction and the motion blur direction are more consistent, the interference degree of the gradient direction x by the motion blur is smaller, the gradient direction x can better represent the motion blur direction, the edge trend feature of the gradient direction x is closer to the template text, and the corner point q is more suitable as the key point for feature matching.
[0076] Step Six: According to the motion blur feature value and the distribution of the thick ink pixel points in the neighborhood of each corner point, screen the corner points to obtain the key points for feature matching of each text detection box, and combine the motion blur direction and the number of thick ink pixel points in the neighborhood of the key points to construct a feature descriptor of the key points in the motion blur direction. Use the key point matching algorithm to automatically recognize the OCR graphics and texts of the data collector, and count the number of characters through the number of text detection boxes.
[0077] In the grayscale image of the graphics and texts, the text strokes consistent with the motion blur direction have a low degree of blurring. Therefore, in this embodiment, within the preset neighborhood of the corner point, a straight line with the direction of the motion blur direction and passing through the corner point is obtained, and the number of thick ink pixel points passed by this straight line in the local neighborhood of the corner point is counted as the thick ink pixel length of the corner point. Among them, the size of the preset neighborhood in this embodiment is 9 * 9.
[0078] Furthermore, in this embodiment, the product of the maximum value of the motion blur feature value of each corner point and the thick ink pixel length is used as the screening robustness strength of each corner point. Specifically, for the convenience of understanding and expression, in this embodiment, the gradient direction with the largest motion blur feature value of each corner point is denoted as the fuzzy representation direction of each corner point. It can be understood that the larger the motion blur feature value is, the smaller the interference degree of the fuzzy representation direction of the corner point by the motion blur is, and the more suitable it is as the key point for feature matching. At the same time, the larger the thick ink pixel length is, the higher the length of the pixel points of the text stroke with high robustness in the motion blur process is, the more it can represent the local information of the text template stroke in the motion blur direction, and the more it should be screened as the key point for feature matching.
[0079] So far, obtain the screening robustness strength of all corner points in the grayscale image of the graphics and texts, arrange the corner points in descending order according to the screening robustness strength, and select the top preset number of corner points as the key points for feature matching of the graphic characters. Among them, the preset number in this embodiment is 20. Normalize the motion blur direction of the grayscale image of the graphics and texts, the fuzzy representation direction of each key point, the motion blur feature value, and the thick ink pixel length respectively, and use the vector composed of the normalized result values as the feature descriptor of each key point in the motion blur direction.
[0080] To achieve the character quantity statistics of the graphic and text automatic recognition machine, in this embodiment, a text template library is first established, and for each text template, its corresponding key points and the feature descriptors of each key point in any motion blur direction are obtained. Further, the SIFI algorithm is used to perform key point matching on the key points in the grayscale image of the graphic and text. In the actual application scenario, the implementer can also use the FLANN fast approximate nearest neighbor search algorithm to perform key point matching. Through key point matching, the text information in the grayscale image of the graphic and text can be obtained, that is, each text detection box in the grayscale image of the graphic and text gets a text template with the highest matching degree, realizing the automatic recognition of OCR graphic and text, and counting the total number of all text detection boxes in the grayscale image of the graphic and text to perform character quantity statistics. Among them, the implementation of key point matching by the SIFI algorithm and the FLANN algorithm is a well-known technology, and the specific process will not be elaborated here.
[0081] It can be understood that the reference to "one embodiment" or "some embodiments" described in the specification of this application means that a specific feature, structure, or characteristic described in combination with the embodiment is included in one or more embodiments of this application. Thus, if "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. appear in different places in this specification, they do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0082] It should be noted that the above-mentioned sequence of the embodiments of this application is only for description and does not represent the superiority or inferiority of the embodiments. And the above describes specific embodiments of this specification. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be beneficial. At the same time, the magnitude of the serial numbers of the steps in the embodiments does not mean the sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments in this specification.
[0083] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.
Claims
1. An OCR graphic and text recognition and statistics method for a data collector, characterized in that, Including the following steps: Perform graphic and image acquisition through a data collector to obtain a graphic grayscale image, and perform graphic character segmentation on the graphic grayscale image to obtain each text detection box; Extract the edge lines and corner points in the graphic grayscale image, analyze the edge line distribution and corner point distribution in each text detection box, and obtain each trend direction and its trend direction significance value; Perform threshold segmentation on the graphic grayscale image to obtain thick ink pixel points, extract thick ink strokes through the region growing algorithm, and determine the thick ink stroke direction of each growing region according to the trend of the thick ink strokes; According to the difference in the number of growing regions in the horizontal and vertical directions of the thick ink stroke direction, construct the horizontal and vertical motion blur direction confidence of the graphic grayscale image, and estimate the motion blur direction of the graphic grayscale image in combination with the number of growing regions in each thick ink stroke direction; According to the motion blur direction estimation result of the graphic grayscale image, and in combination with the gradient direction, gradient amplitude of each corner point and the path information of the thick ink pixel points in the motion blur direction, construct the motion blur eigenvalue of each corner point with respect to each gradient direction; According to the motion blur eigenvalue and the distribution of thick ink pixel points in the neighborhood of each corner point, screen the corner points to obtain the key points for feature matching in each text detection box, and in combination with the motion blur direction and the number of thick ink pixel points in the neighborhood of the key points, construct the feature descriptor of the key points in the motion blur direction, use the key point matching algorithm to automatically recognize the OCR graphics and texts of the data collector, and perform character number statistics through the number of text detection boxes.
2. The OCR graphic and text recognition and statistics method for a data collector according to claim 1, wherein, The obtaining of each trend direction further includes: Statistically calculate the ratio of the number of edge pixels in each text detection box to the total number of pixels in the text detection box, and arrange the text detection boxes in ascending order according to the ratio, and obtain the first preset number of text detection boxes, denoted as the detection boxes to be analyzed; Through the corner points in the detection boxes to be analyzed, segment the edge lines in the detection boxes to be analyzed to obtain each edge segment, perform linear fitting on each edge segment, and calculate the included angle value between the fitted line and the horizontal direction as each trend direction.
3. The OCR graphic and text recognition and statistics method for a data collector according to claim 2, wherein, The method for obtaining the trend direction significance value of each trend direction is: statistically calculate the number of fitted lines in the same trend direction and use it as the trend direction significance value of the same trend direction.
4. The OCR graphic and text recognition and statistics method for a data collector according to claim 1, wherein, The obtaining of the thick ink pixel points further includes: Perform the first threshold segmentation on the graphic grayscale image, output the first segmentation threshold, take the pixel points with gray values lower than the first segmentation threshold as stroke pixel points, and perform the second threshold segmentation on the gray values of all stroke pixel points, output the second segmentation threshold, and take the pixel points with gray values lower than the second segmentation threshold as thick ink pixel points.
5. The OCR graphic and text recognition and statistics method for a data collector according to claim 1, characterized in that, The method for determining the thick ink stroke direction is: perform skeleton extraction on each growing region, perform linear fitting on the skeleton, and determine the included angle between the fitted line and the horizontal direction as the thick ink stroke direction of each growing region.
6. The OCR graphic and text recognition and statistics method for a data collector according to claim 1, characterized in that, The expression of the horizontal and vertical motion blur direction confidence of the graphic grayscale image is: In the formula, T is the confidence of the horizontal and vertical motion blur direction of the graphic grayscale image, μ1 is the number of growth regions with the thick ink stroke direction of 90°, μ2 is the number of growth regions with the thick ink stroke direction of 180°, max() is the maximum value function, and M is the number of growth regions.
7. The OCR graphic and text recognition and statistics method for a data collector according to claim 1, characterized in that The estimation of the motion blur direction of the graphic grayscale image further includes: When the confidence of the horizontal and vertical motion blur directions is greater than or equal to the preset confidence threshold, select the direction with the largest number of growth regions included from the horizontal and vertical directions as the motion blur direction of the grayscale image of the text and image; When the confidence of the horizontal and vertical motion blur directions is less than the preset confidence threshold, calculate the motion blur direction certainty values of all trend directions except the horizontal and vertical directions, and use the trend direction corresponding to the largest motion blur direction certainty value as the motion blur direction of the grayscale image of the text and image; Among them, the expression of the motion blur direction certainty value is as follows: In the formula, R θ , n θ are respectively the motion blur direction certainty value and the trend direction significance value of the trend direction θ. is the motion blur direction weight when the thick ink stroke direction is θ, where the proportion of the number of growth regions in each thick ink stroke direction in the total number of growth regions in all thick ink stroke directions is used as the motion blur direction weight for each thick ink stroke direction.
8. The OCR graphic and text recognition and statistics method for a data collector according to claim 1, characterized in that The expression of the motion blur eigenvalue of each corner point with respect to each gradient direction is: In the formula, v q,x is the motion blur feature value of the corner point q with respect to the gradient direction x, φ is the motion blur direction of the grayscale image of the text and image, Cos() is the cosine value, J is the number of gradient directions, h q,x is the gradient amplitude of the corner point q in the gradient direction x, h q,j is the gradient amplitude of the corner point q in the gradient direction j.
9. The OCR graphic and text recognition and statistics method for a data collector according to claim 1, wherein The screening of the key points includes: Obtain a straight line passing through the corner point with the motion blur direction within the preset neighborhood of each corner point, and count the number of dark ink pixel points passed by the straight line passing through the corner point within the preset neighborhood of the corner point as the dark ink pixel length of the corner point; Take the product of the maximum value of the motion blur eigenvalue of each corner point and the dark ink pixel length as the screening robustness strength of each corner point. Arrange all the corner points in the grayscale image of the text and image in descending order according to the screening robustness strength, and select the top preset number of corner points as the key points.
10. The OCR graphic and text recognition and statistics method for a data collector according to claim 9, characterized in that, The construction of the feature descriptor of the key points in the motion blur direction includes: Normalize the motion blur direction of the grayscale image of the text and image, the blur representation direction of each key point, the motion blur eigenvalue, and the dark ink pixel length respectively, and use the vector composed of the normalized result values as the feature descriptor of each key point in the motion blur direction.
Citation Information
Patent Citations
Video image character recognition method based on submesh characteristic adaptive weighting
CN102663382A
Text paragraph recognition method and device, equipment, medium and program product
CN116978049A