Character detection method and device, storage medium and electronic equipment
By segmenting characters on parts and multi-faceted detection methods, the problems of product confusion and quality traceability caused by character errors are solved, and the detection accuracy and control level of product quality are improved.
Patent Information
- Application Number
- CN202510081904.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
During industrial manufacturing and assembly, characters on parts are prone to omissions, errors or incompleteness, resulting in product confusion and difficulty in quality traceability, affecting the safety and reliability of the product.
A character detection method is provided, by identifying the arrangement shape of the original characters in the target image, dividing them into a single character image, performing character length verification and accuracy verification, and extracting the core shape of the characters for verification and integrity verification of the number of backbone endpoints.
Improve the detection accuracy of character errors, avoid the impact of character arrangement and spacing on detection, and can also detect errors when there are some flaws in characters, enhance the transparency of the supply chain and the traceability of product quality, and reduce returns, rework and cost losses.
Smart Images

Figure CN119992564A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and in particular to a character detection method, a character detection device, a storage medium and an electronic device. Background Art
[0002] In the process of industrial manufacturing and assembly, many parts are often used to complete the production process. Due to various reasons such as equipment failure, operating errors, material problems or environmental factors in the production process, characters on parts may be omitted, mistyped or incomplete. If not corrected in time, it will lead to product confusion, difficulty in quality traceability, and even affect the safety and reliability of the product.
[0003] Character detection is an important prerequisite for character error correction. The higher the accuracy of the detection results, the more correct the characters will be after correction. This can effectively enhance the transparency and traceability of the supply chain, improve the quality control level of products, and reduce returns, rework and cost losses caused by character errors.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0005] The purpose of the present disclosure is to provide a character detection method, a character detection device, a storage medium and an electronic device, aiming to improve the detection accuracy of character errors.
[0006] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.
[0007] According to one aspect of the present disclosure, there is provided a character detection method, comprising:
[0008] Recognize the arrangement shape of the original characters in the target image, and segment the original characters into individual characters according to the character extraction method corresponding to the arrangement shape to obtain individual character images;
[0009] Performing a character length check and / or a character accuracy check based on the character recognition results corresponding to each of the single character images and the real character corresponding to the target image; and
[0010] The core shape of the character in the single character image is extracted as the character backbone, and the backbone endpoint quantity verification and / or integrity detection are performed based on the character backbone.
[0011] Preferably, the performing backbone endpoint quantity verification and / or integrity detection according to the character backbone includes:
[0012] Extracting the endpoint number based on the character backbone, and performing backbone endpoint number verification on the endpoint number using a pre-built endpoint number mapping table; and / or
[0013] The character contour edge of the single character image is extracted, and integrity detection is performed based on the character backbone and the character contour edge.
[0014] Preferably, the extracting the number of endpoints based on the character backbone comprises:
[0015] For a backbone point in the character backbone, when the peripheral backbone points corresponding to the backbone point meet a preset condition, marking the backbone point as an endpoint;
[0016] Traversing all backbone points to count the number of endpoints corresponding to the backbone of the character;
[0017] Among them, the preset condition is that there is 1 peripheral backbone point connected to the backbone point in the connected neighborhood of the backbone point, there is 1 peripheral backbone point connected to the backbone point in the diagonal neighborhood of the backbone point, and there are no more than 2 peripheral backbone points connected in the enclosing edge of the backbone point.
[0018] Preferably, the checking the number of backbone endpoints using a pre-built endpoint number mapping table comprises:
[0019] Obtaining a character recognition result corresponding to the single character image;
[0020] Extracting the standard endpoint number matching the character recognition result from a pre-constructed endpoint number mapping table; wherein the endpoint number mapping table includes the standard endpoint number extracted based on the standard character backbone of the basic character; and performing backbone endpoint number verification on the endpoint number based on the standard endpoint number.
[0021] Preferably, the integrity detection based on the character backbone and the character outline edge includes:
[0022] For a backbone point on the character backbone, marking the intersection of the backbone point at the normal line of the character backbone and the edge of the character outline as a first intersection point and a second intersection point;
[0023] When a first distance between the backbone point and the first intersection point, or a second distance between the backbone point and the second intersection point, is less than a preset distance, marking the backbone point as a defective point;
[0024] All backbone points are traversed, and when a preset number of defective points appear continuously, the result of the integrity test is failure.
[0025] Preferably, when the arrangement shape is a rectangle, dividing the original characters into individual characters comprises:
[0026] Superimposing the grayscale values of the target image according to the arrangement direction of the original characters to obtain a dimension vector, and drawing a grayscale superposition map based on the dimension vector;
[0027] Determine the number of rows of the original characters according to the superimposed area in the grayscale superimposed image, and segment the target image into segmented images corresponding to the number of rows;
[0028] The segmented image is processed using a morphological method to segment the original character into individual characters.
[0029] Preferably, when the arrangement shape is a ring, dividing the original characters into individual characters comprises:
[0030] Performing Hough circle detection on the target image to obtain the center position of the circle, and performing contour extraction on the target image to obtain a contour map of the circumscribed rectangle containing each character;
[0031] Calculating the deviation angle of each character based on the center point coordinates of the circumscribed rectangle of each character and the center position of the circle;
[0032] After correcting each character according to the deviation angle, a single character is obtained.
[0033] According to a second aspect of the present disclosure, there is provided a character detection device, comprising:
[0034] A character segmentation module, used for identifying the arrangement shape of the original characters in the target image, and segmenting the original characters into individual characters according to the character extraction method corresponding to the arrangement shape to obtain individual character images;
[0035] A first detection module, configured to perform a character length check and / or a character accuracy check based on a character recognition result corresponding to each of the single character images and a real character corresponding to the target image; and
[0036] The second detection module is used to extract the core shape of the character in the single character image as the character backbone, and perform backbone endpoint quantity verification and / or integrity detection based on the character backbone.
[0037] According to a third aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the character detection method as described above is implemented.
[0038] According to a fourth aspect of the present disclosure, there is provided an electronic device, including:
[0039] one or more processors;
[0040] The storage device is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the character detection method as described above.
[0041] The exemplary embodiments of the present disclosure may have some or all of the following beneficial effects:
[0042] In the technical solutions provided by some embodiments of the present disclosure, the original characters in the target image are first segmented into individual characters to obtain individual character images, and then the individual character images are detected, including character length verification and character accuracy verification based on the character recognition results, as well as backbone endpoint number verification and integrity detection based on the character backbone. On the one hand, the present disclosure first segments individual characters during character detection, rather than performing overall character recognition based on the image, which can avoid the difficulty of character detection caused by character arrangement and spacing; on the other hand, the use of diversified detection content will help to complete the comprehensive detection of various character contents, thereby improving the detection accuracy of character errors; on the other hand, the detection of endpoint number and integrity is based on the character backbone, rather than relying entirely on the character recognition results, so that even if the character has some flaws but does not affect character recognition, the character backbone error can be discovered as early as possible based on the backbone endpoint number verification, and the problem of incomplete characters can be discovered based on the integrity detection, which can greatly improve the detection accuracy of character errors. This will overall optimize character detection, a prerequisite for character error correction, help enhance supply chain transparency and product quality traceability, improve product quality control, and reduce returns, rework and cost losses.
[0043] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0045] Figure 1 A schematic diagram schematically illustrates a flow chart of a character detection method in an exemplary embodiment of the present disclosure;
[0046] Figure 2 A schematic diagram of a process of segmenting characters in a rectangular arrangement shape in an exemplary embodiment of the present disclosure is schematically shown;
[0047] Figure 3 A schematic diagram of a process of segmenting characters in a circular arrangement shape in an exemplary embodiment of the present disclosure is schematically shown;
[0048] Figure 4 A schematic diagram schematically illustrating a character backbone in an exemplary embodiment of the present disclosure;
[0049] Figure 5 A schematic diagram schematically showing character backbone endpoints in an exemplary embodiment of the present disclosure;
[0050] Figure 6 A schematic diagram schematically illustrates a nine-square grid in an exemplary embodiment of the present disclosure;
[0051] Figure 7 A schematic diagram schematically illustrates an integrity detection in an exemplary embodiment of the present disclosure;
[0052] Figure 8 A schematic diagram schematically shows the composition of a character detection device in an exemplary embodiment of the present disclosure;
[0053] Fig. 9 The structure diagram of a computer system of an electronic device in an exemplary embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0054] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more comprehensive and complete and will fully convey the concept of the example embodiments to those skilled in the art.
[0055] In addition, the described features, structures or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the present disclosure.
[0056] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0057] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0058] In production and processing, due to equipment failure, operational errors, material problems or environmental factors (such as oil pollution, rust) and other reasons in the production process, the characters on the fasteners may be missing or blurred. Character errors may cause product confusion, quality traceability difficulties, and even affect product safety and reliability. It can also enhance the transparency and traceability of the supply chain, improve the quality control level of products, and reduce returns, rework and cost losses caused by character errors. Therefore, how to detect character errors in a timely manner is of great significance for character error correction.
[0059] The common existing technology is character detection based on template matching, that is, using image processing technology to extract the character features of the characters on the fasteners, and then comparing them with a series of pre-defined character templates to achieve character recognition and error correction. This technology is heavily dependent on the number and quality of templates, and in practical applications, it is impossible to exhaust all character styles. In addition, when the characters undergo transformations such as rotation and scaling, the adaptability of this method is poor, which easily leads to poor error correction results.
[0060] Another common method is character detection based on machine learning, which is to build a deep neural network model, input the labeled characters into the neural network, train and learn the characters, identify the characteristics and rules of the characters, and thus realize character recognition and error correction. However, this technology requires a large amount of labeled data as training samples, and the generalization ability of the model is affected by the quality and quantity of training data. In addition, deep learning requires high computing power and has high hardware requirements.
[0061] In addition, there is also character detection based on optical character recognition (OCR) technology, which uses OCR technology to convert characters in an image into computer-recognizable text, and then compares and corrects errors. This technology may not be effective for special characters, deformed characters, or low-contrast characters. In addition, OCR has certain requirements for the arrangement and spacing of characters. If the characters are arranged irregularly or the spacing is large / small, recognition errors may occur.
[0062] Therefore, in view of the shortcomings of the prior art, the present disclosure provides a character detection method to solve one or more problems existing in the prior art. The implementation details of the technical solution of the embodiment of the present disclosure are described in detail below.
[0063] Figure 1The following is a schematic diagram showing a flow chart of a character detection method in an exemplary embodiment of the present disclosure. Figure 1 As shown, the character detection method includes steps S101 to S105:
[0064] Step S101, identifying the arrangement shape of original characters in a target image, and segmenting the original characters into individual characters according to a character extraction method corresponding to the arrangement shape to obtain individual character images;
[0065] Step S103, performing a character length check and / or a character accuracy check based on the character recognition results corresponding to each of the single character images and the real characters corresponding to the target image; and
[0066] Step S105, extracting the core shape of the character in the single character image as the character backbone, and performing backbone endpoint quantity verification and / or integrity detection based on the character backbone.
[0067] In the technical solutions provided by some embodiments of the present disclosure, the original characters in the target image are first segmented into individual characters to obtain individual character images, and then the individual character images are detected, including character length verification and character accuracy verification based on the character recognition results, as well as backbone endpoint number verification and integrity detection based on the character backbone. On the one hand, the present disclosure first segments individual characters during character detection, rather than performing overall character recognition based on the image, which can avoid the difficulty of character detection caused by character arrangement and spacing; on the other hand, the use of diversified detection content will help to complete the comprehensive detection of various character contents, thereby improving the detection accuracy of character errors; on the other hand, the detection of endpoint number and integrity is based on the character backbone, rather than relying entirely on the character recognition results, so that even if the character has some flaws but does not affect character recognition, the character backbone error can be discovered as early as possible based on the backbone endpoint number verification, and the problem of incomplete characters can be discovered based on the integrity detection, which can greatly improve the detection accuracy of character errors. This will overall optimize character detection, a prerequisite for character error correction, help enhance supply chain transparency and product quality traceability, improve product quality control, and reduce returns, rework and cost losses.
[0068] Below, each step of the character detection method in this example implementation will be described in more detail with reference to the accompanying drawings and embodiments.
[0069] In step S101, the arrangement shape of the original characters in the target image is identified, and the original characters are segmented into single characters according to a character extraction method corresponding to the arrangement shape to obtain a single character image.
[0070] Specifically, due to the limitation of the shape of the materials of parts such as fasteners, the arrangement shape of common characters may be rectangular or circular, etc. Different character extraction methods are used to segment the characters according to different arrangement shapes.
[0071] In one embodiment of the present disclosure, when the arrangement shape is a rectangle, segmenting the original character into individual characters includes: superimposing the grayscale values of the target image according to the arrangement direction of the original character to obtain a dimensional vector, and drawing a grayscale overlay image based on the dimensional vector; determining the number of rows of the original character according to the superposition area in the grayscale overlay image, and segmenting the target image into segmented images corresponding to the number of rows; and processing the segmented image using a morphological method to segment the original character into individual characters.
[0072] Specifically, for rectangular character features, Figure 2 The following is a schematic diagram of a process of segmenting characters in a rectangular arrangement in an exemplary embodiment of the present disclosure. Figure 2 As shown by Figure 2 (a)-(d) show the schematic diagrams of the results of each step of segmenting characters under the rectangular arrangement shape. Each step is described in detail below.
[0073] Figure 2 (a) shows the original target image. The characters displayed in the target image are arranged in a rectangular shape. First, the target image is superimposed with grayscale values according to the arrangement direction of the original characters. If the characters are arranged horizontally, the grayscale values of the pixels in the same row are superimposed. If the characters are arranged vertically, the grayscale values of the pixels in the same column are superimposed, and finally a one-dimensional vector is obtained.
[0074] Figure 2 (b) shows the grayscale overlay after visualizing the dimensional vector. The overlay values of each grayscale value are plotted on the coordinate axis and the points are connected into a curve. The closed area seat grayscale overlay is obtained by filling the curve and the y-axis, as shown in Figure 2 (b) as shown.
[0075] Figure 2 (c) shows a schematic diagram after the line segmentation. The original characters are segmented according to the number of character lines according to the overlapping area in the dimensional vector graph, such as Figure 2 As shown in (b), there are two black parts as overlapping areas, so the original characters are divided into two lines, such as Figure 2 (c) shows the superposition area. The superposition area is the area formed by the accumulation of pixels in each line of two lines of characters. Figure 2 The upper black part in (b) corresponds to the first row of characters, and the lower black part corresponds to the second row of characters.
[0076] Figure 2(d) shows a single character image after segmentation. The morphological method is used to process and the coordinate information obtained from the connected region is combined to separate each character in order. Connected regions are a common method in the field of image processing. This method can return the label identification, centroid position, and boundary position information of each region. First, two characters are separated according to the centroid ordinate, and then the correct order of the characters is obtained according to the horizontal coordinate. Finally, a single character image is obtained after the single character segmentation. Figure 2 (d) as shown.
[0077] In one embodiment of the present disclosure, when the arrangement shape is a ring, segmenting the original characters into individual characters includes: performing Hough circle detection on the target image to obtain the center position of the circle, and performing contour extraction on the target image to obtain a contour map of the circumscribed rectangle containing each character; calculating the deviation angle of each character based on the center point coordinates of the circumscribed rectangle of each character and the center position of the circle; and correcting each character according to the deviation angle to obtain a single character.
[0078] Specifically, for the circular character features, Figure 3 The following is a schematic diagram of a process of segmenting characters in a circular arrangement in an exemplary embodiment of the present disclosure. Figure 3 As shown by Figure 3 (a)-(c) show schematic diagrams of the results obtained in each step of segmenting characters in a circular arrangement shape.
[0079] Figure 3 (a) shows the original target image. Figure 3 (b) shows the effect of the target image after contour extraction. The characters displayed in the target image are arranged in a circular shape. After contour extraction of the target image, the display is as follows: Figure 3 (b) as shown.
[0080] Figure 3 (c) and Figure 3 (d) shows the target image character deviation angle schematic diagram, Figure 3 (e) shows a single character image after segmentation. The target image after grayscale processing is subjected to Hough circle detection to obtain the center position of the circle, and then the image is subjected to contour extraction to obtain the contour map, as shown in FIG. Figure 3 As shown in (c), the contour image contains the circumscribed rectangle of each character. Through the center coordinates of the circumscribed rectangle of each character and the center coordinates of the circle, the deviation angle of each character can be obtained, as shown in Figure 3 (d) shows that the center of the circle is O 1 , the center coordinates of the circumscribed rectangle are O 2 , get the deviation angle ∠O of the character 2 O 1O. Then, each character is corrected according to this angle, and the segmented single character image is obtained, such as Figure 3 (e) as shown.
[0081] Based on the above method, individual characters are first segmented during character detection, rather than performing overall character recognition based on an image, which can avoid the difficulty of character detection caused by character arrangement and spacing.
[0082] In step S103, character length verification and / or character accuracy verification is performed based on the character recognition results corresponding to each of the single character images and the real characters corresponding to the target image.
[0083] Specifically, after the target image is segmented into individual characters to obtain individual character images, the character recognition results may be used to first detect the length and accuracy of the characters.
[0084] First, obtain the character recognition result. You can use OCR (Optical Character Recognition) technology to recognize the characters in a single character image. You can also train a network model that can perform character recognition based on a pre-trained neural network, and then input the image to obtain the character recognition result output by the network model.
[0085] When checking the character length, the actual character order is known. In order to better match the character order closest to the actual character order from the character recognition results, the Levenshtein (string editing) distance can be used to calculate the distance between the two strings, and the recognized character results can be corrected in sequence. The corrected character length is then checked with the actual character length to obtain a first test result. If the lengths are consistent, the first test result is passed; if the lengths are inconsistent, the first test result is failed.
[0086] When checking the accuracy of characters, the character recognition result of a single character image can be checked with the real character corresponding to the character image to determine whether the character recognition result and the real character are consistent, and then obtain a second detection result. If the characters are consistent, the second detection result is passed; if the characters are inconsistent, the second detection result is failed.
[0087] It should be noted that the content of step S103 may include only character length check, or only character accuracy check, or of course, may include both contents at the same time.
[0088] In step S105, the core shape of the character in the single character image is extracted as the character backbone, and the backbone endpoint quantity verification and / or integrity detection are performed based on the character backbone.
[0089] First, the core shape of the character in the single character image is extracted as the character backbone. Specifically, the Zhang parallel fast refinement algorithm can be used to extract the character backbone of a single character image. Among them, the Zhang fast parallel refinement algorithm is an algorithm used to extract the skeleton of a binary image in image processing, that is, to extract the central axis of the image. The Zhang fast parallel refinement algorithm is based on the 8-neighborhood system of the image. For a given pixel point P1, there are 8 adjacent pixel points around it, marked as P2 to P9 respectively. The algorithm gradually refines the image by iteratively deleting boundary points that meet specific conditions until only the skeleton of the image remains.
[0090] Figure 4 A schematic diagram schematically illustrates a character skeleton in an exemplary embodiment of the present disclosure. Figure 4 As shown, schematic diagrams of multiple basic characters and corresponding character skeletons are shown, including numbers and letters.
[0091] In one embodiment of the present disclosure, the checking the number of backbone endpoints and / or the integrity detection according to the character backbone includes:
[0092] Step 1: extracting the endpoint number based on the character backbone, and performing backbone endpoint number verification on the endpoint number using a pre-built endpoint number mapping table; and / or
[0093] Step 2: extract the character contour edge of the single character image, and perform integrity detection based on the character backbone and the character contour edge.
[0094] It should be noted that the content of step S105 may include only the backbone endpoint quantity verification in step 1, or only the integrity check in step 2, or of course, may include both contents at the same time.
[0095] In step 1, the number of backbone endpoints is checked. Specifically, the number of endpoints needs to be extracted based on the character backbone, which includes the following steps: for a backbone point in the character backbone, when the peripheral backbone point corresponding to the backbone point meets the preset conditions, the backbone point is marked as an endpoint; all backbone points are traversed to count the number of endpoints corresponding to the character backbone; wherein the preset conditions are that there is one peripheral backbone point connected to the backbone point in the adjacent neighborhood of the backbone point, one peripheral backbone point connected to the backbone point in the diagonal neighborhood of the backbone point, and no more than two peripheral backbone points are connected in the surrounding edge of the backbone point.
[0096] Figure 5 The schematic diagram schematically shows the character backbone endpoints in the exemplary embodiment of the present disclosure. Figure 5 As shown, the present disclosure provides 8 situations where endpoints appear, such as Figure 8 As shown, the middle one is For each backbone point in the character backbone, the backbone point is placed in the center of the nine-square grid, and the arrangement of the outer backbone points in the eight empty spaces around the backbone point is analyzed to determine whether the backbone point is an endpoint.
[0097] To facilitate the description of the pre-conditions, Figure 6 A schematic diagram of a nine-square grid in an exemplary embodiment of the present disclosure is schematically shown. Figure 6 As shown, number 5 is located in the center of the nine-square grid, numbers 2, 4, 6, and 8 are directly adjacent to number 5, which are connected neighbors, numbers 1, 3, 7, and 9 are indirectly adjacent to number 5, in a diagonal shape, which are diagonal neighbors, and there are four surrounding edges, namely, the line connecting positions 1, 2, and 3, the line connecting positions 1, 4, and 7, the line connecting positions 3, 6, and 9, and the line connecting positions 7, 8, and 9.
[0098] Therefore, to determine whether the backbone point is an endpoint, the following conditions must be met at the same time: there is only one peripheral backbone point in the adjacent neighborhood, there is only one peripheral backbone point in the diagonal neighborhood, and there are no more than two peripheral backbone points connected to each other in the surrounding edge. If the peripheral backbone points of the backbone point meet the above preset conditions, then the backbone point is an endpoint, and all the backbone points are traversed to count the number of endpoints of the character backbone.
[0099] Then the extracted endpoint number is verified with the standard endpoint number recorded in the endpoint number mapping table, thereby completing the backbone endpoint number verification. In one embodiment of the present disclosure, the backbone endpoint number verification of the endpoint number using the pre-constructed endpoint number mapping table includes: obtaining the character recognition result corresponding to the single character image; extracting the standard endpoint number matching the character recognition result from the pre-constructed endpoint number mapping table; wherein the endpoint number mapping table includes the standard endpoint number extracted based on the standard character backbone of the basic character; and performing backbone endpoint number verification on the endpoint number based on the standard endpoint number.
[0100] Specifically, it is necessary to pre-build an endpoint number mapping table. When building the endpoint number mapping table, the process of extracting the endpoint number of the character backbone is the same as the process of extracting the endpoint number based on the character backbone in step 1. The only difference is that the standard character backbone of the basic character is used here. The endpoints are extracted for 26 commonly used letters and 10 numbers, a total of 36 characters, and the endpoint number mapping table is shown in Table 1.
[0101] Table 1
[0102] character A B C D E F G H I J K L Number of endpoints 2 0 2 0 3 3 2 4 2 2 3 2 character M N O P Q R S T U V W X Number of endpoints 2 2 0 1 2 2 2 3 2 2 2 4 character Y Z 0 1 2 3 4 5 6 7 8 9 Number of endpoints 3 2 0 2 2 2 2 2 1 2 0 1
[0103] Therefore, after extracting the number of endpoints corresponding to a single character image, the standard number of endpoints is extracted from the endpoint number mapping table according to the character result corresponding to the single character image, and the extracted number of endpoints is compared with the standard number of endpoints to obtain the third detection result. If the numbers are consistent, the backbone endpoint number verification passes, and the third detection result is passed. If the numbers are inconsistent, the third detection result is failed.
[0104] In the second step, an integrity check is performed. Specifically, the integrity check is performed based on the character backbone and the character contour edge, including: for a backbone point on the character backbone, marking the intersection of the backbone point at the normal of the character backbone and the character contour edge with the first intersection and the second intersection; when the first distance between the backbone point and the first intersection, or the second distance between the backbone point and the second intersection is less than a preset distance, marking the backbone point as a defective point; traversing all backbone points, when a preset number of defective points appear continuously, the integrity check result is obtained as a failure.
[0105] Specifically, integrity detection is to detect whether there are partial defects in a single character. Even if there are partial defects, the character can still be recognized during fuzzy recognition. However, in this case, if the character cannot be corrected in time, the character will continue to wear and damage over time, which will increase the difficulty of subsequent character recognition. Therefore, defects need to be detected as early as possible for correction.
[0106] Figure 7 The following is a schematic diagram of an integrity detection in an exemplary embodiment of the present disclosure. Figure 7 As shown, the middle line is the character backbone extracted based on a single character image, representing the core shape of the character, and the strip area is the character contour extracted based on a single character image, so the edge of the strip area is the character contour edge.
[0107] The character backbone is obtained according to a certain granularity or coordinate axis to obtain multiple backbone points. Taking a backbone point P in the character backbone as an example, first determine the normal L of P on the character backbone, and then mark the two intersection points of L and the edge of the character contour as the first intersection point p 1 and the second intersection point p 2 Calculate Pp 1 The first distance, and Pp 2 A preset distance is also set in advance. If any of the first distance and the second distance is less than the preset distance, the backbone point P is determined to be a defective point. If a sufficient number of defective points appear continuously, the number of defective points can still be limited by the preset number, and it can be concluded that the single character is incomplete.
[0108] The preset distance can be configured as needed. For example, if the width of a character is 10 and the character backbone is located at the center of the character, the distance on one side should be 5. Then the preset distance can be designed to be 5. That is to say, when the first distance and / or the second distance is ≥5, the backbone point is not a defect point. If the first distance and / or the second distance is <5, the backbone point is a defect point. Of course, the preset distance can also be an average value obtained after fuzzy calculation in advance.
[0109] It should be noted that in addition to judging defect points based on the preset distance, defect points can also be judged based on whether they are within a preset distance interval. Taking the character width of 10 as an example, the preset distance interval can be set to an upper and lower floating interval of 5, for example [4,6]. If the first distance and / or the second distance is within [4,6], then the backbone point will not be regarded as a defect point.
[0110] The preset number may also be preconfigured. For example, if 10 defective points appear continuously, the segment is regarded as an incomplete character, that is, the fourth detection result is a failure. Otherwise, it is regarded as a pass.
[0111] Based on the above method, the number and integrity of endpoints are detected through the character backbone instead of relying entirely on the character recognition results. Therefore, even if the character has some flaws but does not affect the character recognition, the character backbone errors can be discovered as early as possible based on the backbone endpoint number check, and the problem of incomplete characters can be discovered based on the integrity check, which can greatly improve the detection accuracy of character errors.
[0112] It should be noted that the two character defect detection methods provided in the present disclosure are merely exemplary descriptions. On this basis, those skilled in the art can easily think that the character backbone can also be moved up and down by a preset distance, and then the character outline and the two moved character backbone lines can be tested for integrity. Alternatively, the character backbone can be moved up and down according to a preset distance interval to obtain two strip-shaped areas, and then the character outline and the two strip-shaped areas can be tested for integrity. All of these belong to the technical solutions protected by the present disclosure.
[0113] In one embodiment of the present disclosure, character detection needs to go through step S103 and step S105. The present disclosure does not limit the order between the two steps, which can be performed one after the other or simultaneously.
[0114] Finally, the first test result of the character length check and / or the second test result of the character accuracy check, as well as the third test result of the backbone endpoint number check and / or the fourth test result of the integrity check are obtained, and finally all the test results are combined to obtain the final test result of the target image. Specifically, if all the test results pass, the character test of the target image is considered to have passed, otherwise a test result report will be issued based on each test result, which is convenient for character error correction.
[0115] It should be noted that these four detection contents can be combined according to actual needs. For example, character detection can also be designed as vertical content, that is, the next detection content can be carried out only after the previous detection is passed, or a character detection method combining vertical and parallel detection, for example, the first detection content is character length verification, and after the detection is passed, character accuracy verification, backbone endpoint number verification and integrity detection are carried out at the same time. Those skilled in the art should understand that the above technical solutions should also be included in the protection scope of the present disclosure.
[0116] Figure 8 The following is a schematic diagram showing the composition of a character detection device in an exemplary embodiment of the present disclosure. Figure 8 As shown, the character detection device 800 may include a character segmentation module 801, a first detection module 802 and a second detection module 803. Among them:
[0117] The character segmentation module 801 is used to identify the arrangement shape of the original characters in the target image, and segment the original characters into individual characters according to the character extraction method corresponding to the arrangement shape to obtain individual character images;
[0118] A first detection module 802 is used to perform a character length check and / or a character accuracy check based on the character recognition results corresponding to each of the single character images and the real characters corresponding to the target image; and
[0119] The second detection module 803 is used to extract the core shape of the character in the single character image as the character backbone, and perform backbone endpoint quantity verification and / or integrity detection based on the character backbone.
[0120] According to an exemplary embodiment of the present disclosure, the second detection module is also used to extract the number of endpoints based on the character backbone, so as to perform backbone endpoint quantity verification on the endpoint number using a pre-built endpoint number mapping table; and / or extract the character contour edge of the single character image, and perform integrity detection based on the character backbone and the character contour edge.
[0121] According to an exemplary embodiment of the present disclosure, the second detection module is also used to target a backbone point in the character backbone, and when the peripheral backbone point corresponding to the backbone point meets a preset condition, mark the backbone point as an endpoint; traverse all the backbone points to count the number of endpoints corresponding to the character backbone; wherein the preset condition is that there is 1 peripheral backbone point connected to the backbone point in the connected neighborhood of the backbone point, there is 1 peripheral backbone point connected to the backbone point in the diagonal neighborhood of the backbone point, and there are no more than 2 peripheral backbone points connected in the enclosing edge of the backbone point.
[0122] According to an exemplary embodiment of the present disclosure, the second detection module is also used to obtain a character recognition result corresponding to the single character image; extract a standard endpoint number matching the character recognition result from a pre-constructed endpoint number mapping table; wherein the endpoint number mapping table includes a standard endpoint number extracted based on a standard character backbone of a basic character; and perform backbone endpoint number verification on the endpoint number based on the standard endpoint number.
[0123] According to an exemplary embodiment of the present disclosure, the second detection module is also used to mark the intersection of the backbone point on the character backbone and the normal of the character backbone and the edge of the character contour with the first intersection and the second intersection; when the first distance between the backbone point and the first intersection, or the second distance between the backbone point and the second intersection is less than a preset distance, mark the backbone point as a defective point; traverse all the backbone points, and when a preset number of the defective points appear continuously, obtain a detection result of failure for the integrity detection.
[0124] According to an exemplary embodiment of the present disclosure, when the arrangement shape is a rectangle, the character segmentation module is also used to superimpose the grayscale values of the target image according to the arrangement direction of the original characters to obtain a dimensional vector, and draw a grayscale overlay map based on the dimensional vector; determine the number of rows of the original character according to the superposition area in the grayscale overlay image, and segment the target image into a number of segmented images corresponding to the number of rows; and use a morphological method to process the segmented image to segment the original character into individual characters.
[0125] According to an exemplary embodiment of the present disclosure, when the arrangement shape is a ring, the character segmentation module is further used to perform Hough circle detection on the target image to obtain the center position of the circle, and to perform contour extraction on the target image to obtain a contour map of the circumscribed rectangle containing each character; calculate the deviation angle of each character based on the center point coordinates of the circumscribed rectangle of each character and the center position of the circle; and correct each character according to the deviation angle to obtain a single character.
[0126] According to an exemplary embodiment of the present disclosure,
[0127] The specific details of each module in the above-mentioned character detection device 800 have been described in detail in the corresponding character detection method, so they will not be repeated here.
[0128] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0129] In an exemplary embodiment of the present disclosure, a storage medium capable of implementing the above method is also provided. It can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a mobile phone. However, the program product of the present disclosure is not limited to this. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.
[0130] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided. Fig. 9 The structure diagram of a computer system of an electronic device in an exemplary embodiment of the present disclosure is schematically shown.
[0131] It should be noted that Fig. 9 The computer system 900 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0132] like Fig. 9 As shown, the computer system 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 902 or the program loaded from the storage part 908 to the random access memory (RAM) 903. In the RAM 903, various programs and data required for system operation are also stored. The CPU 901, the ROM 902 and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0133] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that a computer program read therefrom is installed into the storage section 908 as needed.
[0134] In particular, according to an embodiment of the present disclosure, the process described below with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part 909, and / or installed from a removable medium 911. When the computer program is executed by a central processing unit (CPU) 901, various functions defined in the system of the present disclosure are executed.
[0135] It should be noted that the computer-readable medium shown in the embodiment of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by an instruction execution system, device or device or used in combination with it. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0136] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0137] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, and the units described may also be arranged in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.
[0138] As another aspect, the present disclosure further provides a computer-readable medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiment.
[0139] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0140] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
[0141] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure.
[0142] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A character detection method, characterized in that: include: Recognize the arrangement shape of the original characters in the target image, and segment the original characters into individual characters according to the character extraction method corresponding to the arrangement shape to obtain individual character images; Performing a character length check and / or a character accuracy check based on the character recognition results corresponding to each of the single character images and the real characters corresponding to the target image; as well as The core shape of the character in the single character image is extracted as the character backbone, and the backbone endpoint quantity verification and / or integrity detection are performed based on the character backbone.
2. The character detection method according to claim 1, characterized in that: The performing backbone endpoint quantity verification and / or integrity detection according to the character backbone includes: Extracting the endpoint number based on the character backbone, and performing backbone endpoint number verification on the endpoint number using a pre-built endpoint number mapping table; and / or The character contour edge of the single character image is extracted, and integrity detection is performed based on the character backbone and the character contour edge.
3. The character detection method according to claim 2, characterized in that: The step of extracting the number of endpoints based on the character backbone comprises: For a backbone point in the character backbone, when the peripheral backbone points corresponding to the backbone point meet a preset condition, marking the backbone point as an endpoint; Traversing all backbone points to count the number of endpoints corresponding to the backbone of the character; Among them, the preset condition is that there is 1 peripheral backbone point connected to the backbone point in the connected neighborhood of the backbone point, there is 1 peripheral backbone point connected to the backbone point in the diagonal neighborhood of the backbone point, and there are no more than 2 peripheral backbone points connected in the enclosing edge of the backbone point.
4. The character detection method according to claim 2, characterized in that: The checking of the number of backbone endpoints using a pre-built endpoint number mapping table for the number of endpoints includes: Obtaining a character recognition result corresponding to the single character image; Extracting the standard endpoint number matching the character recognition result from a pre-constructed endpoint number mapping table; wherein the endpoint number mapping table includes the standard endpoint number extracted based on the standard character backbone of the basic character; and performing backbone endpoint number verification on the endpoint number based on the standard endpoint number.
5. The character detection method according to claim 2, characterized in that: The integrity detection according to the character backbone and the character outline edge includes: For a backbone point on the character backbone, marking the intersection of the backbone point at the normal line of the character backbone and the edge of the character outline as a first intersection point and a second intersection point; When a first distance between the backbone point and the first intersection point, or a second distance between the backbone point and the second intersection point, is less than a preset distance, marking the backbone point as a defective point; All backbone points are traversed, and when a preset number of defective points appear continuously, the result of the integrity test is failure.
6. The character detection method according to claim 1, characterized in that: When the arrangement shape is a rectangle, dividing the original characters into individual characters comprises: Superimposing the grayscale values of the target image according to the arrangement direction of the original characters to obtain a dimension vector, and drawing a grayscale superposition map based on the dimension vector; Determine the number of rows of the original characters according to the superimposed area in the grayscale superimposed image, and segment the target image into segmented images corresponding to the number of rows; The segmented image is processed using a morphological method to segment the original character into individual characters.
7. The character detection method according to claim 1, characterized in that: When the arrangement shape is a ring, dividing the original characters into individual characters includes: Performing Hough circle detection on the target image to obtain the center position of the circle, and performing contour extraction on the target image to obtain a contour map of the circumscribed rectangle containing each character; Calculating the deviation angle of each character based on the center point coordinates of the circumscribed rectangle of each character and the center position of the circle; After correcting each character according to the deviation angle, a single character is obtained.
8. A character detection device, characterized in that: include: A character segmentation module, used for identifying the arrangement shape of the original characters in the target image, and segmenting the original characters into individual characters according to the character extraction method corresponding to the arrangement shape to obtain individual character images; A first detection module, used for performing a character length check and / or a character accuracy check based on a character recognition result corresponding to each of the single character images and a real character corresponding to the target image; as well as The second detection module is used to extract the core shape of the character in the single character image as the character backbone, and perform backbone endpoint quantity verification and / or integrity detection based on the character backbone.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the character detection method according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enables the one or more processors to implement the character detection method as described in any one of claims 1 to 7.