Optical character recognition system and method

By combining image preprocessing, digital segmentation, neural network and K-nearest neighbor algorithm, the optical character recognition system solves the problem of inaccurate recognition in harsh environments and achieves higher recognition accuracy and stability.

CN115393856BActive Publication Date: 2025-09-23LISHUI POWER SUPPLY COMPANY OF STATE GRID ZHEJIANG ELECTRIC POWER +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210673917.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-14
Publication Date
2025-09-23
Estimated Expiration
2042-06-14

AI Technical Summary

Technical Problem

Existing optical character recognition technology cannot adapt to various harsh real-world environments. For example, when capturing gas and electricity meter images, uneven lighting causes high image noise and inaccurate recognition results.

Method used

Through the recognition system combining image acquisition, preprocessing, digital refinement and segmentation, neural network and K-nearest neighbor algorithm, the image contrast is optimized and the output results of the two recognition models are integrated to improve the recognition accuracy.

Benefits of technology

It improves the accuracy of character recognition in complex environments, reduces the impact of noise, simplifies the input of neural networks, and maintains the stability of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393856B_ABST
    Figure CN115393856B_ABST
Patent Text Reader

Abstract

The present invention provides an optical character recognition system and method. By performing processing such as thinning and segmenting on a captured image and combining two recognition models, an artificial neural network and a K-nearest neighbor algorithm, the system and method include the following steps: S1: image acquisition; S2: image preprocessing; S3: image thinning and segmentation, thinning digits with width in the image and dividing the refined digits into several line segments; S4: inputting the angle information between adjacent line segments into the neural network and the K-nearest neighbor algorithm model for digit recognition; S5: the neural network and the K-nearest neighbor algorithm model generate possible preliminary recognition results and corresponding confidence levels; S6: selecting the preliminary recognition result with the highest confidence level as the final recognition result and outputting it. This system primarily addresses the technical problem that existing optical character recognition technology cannot adapt to various harsh real-world environments and produces inaccurate recognition results, thereby improving character recognition accuracy in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular to an optical character recognition system and method. Background Art

[0002] Optical character recognition (OCR) is a method for locating and recognizing handwritten, typed, or printed text stored in images (such as JPEG or GIF images) and converting the text into a computer-recognizable form (such as ASCII or Unicode). OCR, a field of study within pattern recognition, artificial intelligence, and computer vision, converts pixel representations of letters into their equivalent character representations. OCR is primarily used in paper-intensive industries that have large volumes of paper forms and documents and face the complex real-world image environment: image degradation, high noise levels, image distortion, and low resolution.

[0003] Therefore, current OCR applications still have certain limitations. For example, many OCR programs require a trade-off between speed and accuracy. Some programs are very accurate, but as a result, they run slower and have lower recognition efficiency.

[0004] For example, Chinese patent CN201910189723.3 discloses an OCR recognition method and system, including: training an OCR deep learning model based on a preset first character image training sample set; obtaining a first similarity threshold corresponding to the OCR deep learning model; the first similarity threshold is used to distinguish different characters; obtaining characters that cannot be distinguished by the first similarity threshold from a preset character library to obtain a similar character set; training the OCR deep learning model based on the similar character set; and calling the OCR deep learning model to recognize character images. Although this invention improves the accuracy of OCR recognition to a certain extent while meeting the recognition speed requirements, its recognition accuracy cannot be guaranteed in harsh environments and when the quality of the collected images is low. Summary of the Invention

[0005] This invention primarily addresses the technical problem that existing optical character recognition technology is unable to adapt to various harsh real-world environments. For example, when capturing images of gas and electricity meters, the images are often unevenly illuminated, with some parts in shadow and others overexposed. This results in very high noise levels in the captured images, ultimately leading to inaccurate recognition results. This invention provides an optical character recognition system and method that, by performing image refinement and segmentation processing on the captured images and combining two recognition models, an artificial neural network and a K-nearest neighbor algorithm, ultimately improves character recognition accuracy in complex environments.

[0006] The above technical problems of the present invention are mainly solved by the following technical solutions: an optical character recognition system, comprising an image acquisition module, an image preprocessing module, an image digital refinement and segmentation module, a digital information processing module, a neural network recognition module, a K-nearest neighbor algorithm recognition module and a recognition result processing and output module, wherein the image acquisition module is connected to the image preprocessing module, the image preprocessing module is connected to the image digital refinement and segmentation module, the image digital refinement and segmentation module is connected to the digital information processing module, the digital information processing module is connected to the neural network recognition module and the K-nearest neighbor algorithm recognition module respectively, and the neural network recognition module and the K-nearest neighbor algorithm recognition module are both connected to the recognition result processing and output module. The image preprocessing and image digital refinement and segmentation modules in the present invention can optimize images collected in harsh environments, improve their contrast, and facilitate further specific processing of subsequent images. At the same time, the two different recognition methods of the neural network recognition module and the K-nearest neighbor algorithm recognition module are combined, and the information output by the two recognition modules is finally integrated to obtain the final recognition result, thereby improving the accuracy of character recognition in harsh environments.

[0007] Preferably, the image digital thinning and segmentation module thins the digits with width in the image and segments the thinned digits into a plurality of line segments. The digital information processing module obtains angular information between adjacent line segments based on the processing results of the image digital thinning and segmentation module, and inputs the obtained angular information into the neural network recognition module and the K-nearest neighbor algorithm recognition module, respectively. Often, due to differences in exposure, the amount of noise in an image is almost equal to the number of pixels representing the digits themselves. Therefore, a method for minimizing noise is needed. The present invention proposes a thinning method that reduces a digit to a line and then segments the line into equal or unequal segments. The main feature of the thinning process is that it maintains length.

[0008] Preferably, the neural network recognition module and the K-nearest neighbor algorithm recognition module each input their preliminary recognition results and corresponding confidence information into the recognition result processing and output module, which then determines and outputs a final recognition result based on the input information. The present invention, based on the training process of the two recognition models, synthesizes the preliminary recognition results and corresponding confidence levels output by the two recognition models, lists various possible recognition results and confidence levels for the input character, and selects the one with the highest confidence level as the final recognition result, thereby further improving the accuracy of character recognition.

[0009] The present invention also provides an optical character recognition method, comprising the following steps:

[0010] S1: Image acquisition;

[0011] S2: image preprocessing;

[0012] S3: Image thinning and segmentation, thinning the digits with width in the image and segmenting the thinned digits into several line segments;

[0013] S4: Inputting the angle information between adjacent line segments into the neural network and the K-nearest neighbor algorithm model for digital recognition;

[0014] S5: The neural network and K-nearest neighbor algorithm model obtain possible preliminary recognition results and corresponding confidence levels;

[0015] S6: Select the preliminary recognition result with the highest confidence as the final recognition result and output it.

[0016] By refining and segmenting the captured image and combining it with two recognition models, an artificial neural network and a K-nearest neighbor algorithm, the accuracy of character recognition in complex environments is ultimately improved. Using the angles between adjacent segmented segments as input to the neural network and K-nearest neighbor models offers the following advantages over algorithms that use individual pixel values: (1) the input is significantly smaller, (2) the weight of noise is reduced, and (3) the rotation of the digits does not affect the neural network architecture or its input.

[0017] Preferably, step S2 specifically includes converting the captured image into a monochrome image or grayscale image. After conversion, the contrast of an image generally having multiple colors is improved through histogram equalization. This adjustment allows the intensity to be better distributed across the histogram, allowing areas of lower local contrast to achieve higher contrast without affecting global contrast.

[0018] Preferably, step S2 also includes sharpening and adaptive thresholding. The adaptive thresholding process segments the image by setting all pixels with intensity values ​​above a threshold to foreground values ​​and all remaining pixels to background values. After this stage, the image is ready for further specific processing.

[0019] Preferably, the step S3 specifically includes:

[0020] S31: Select a boundary pixel;

[0021] S32: constructing a minimum square matrix that contains the selected pixels and does not overlap with existing adjacent matrices;

[0022] S33: Calculate the centroid pixel of the square matrix. If the center of the square matrix is ​​a 2*2 matrix pixel, take any pixel in the 2*2 matrix pixel as the centroid pixel, and replace the square matrix with the centroid pixel.

[0023] S34: Return to step S31 until there are no more available pixels;

[0024] S35: Connect the centroid pixels in sequence to obtain a number composed of several line segments, that is, the refined number is divided into several line segments.

[0025] In addition, it should be pointed out that the image refinement and image segmentation processing in step S3 can also be performed separately, and segmentation is the next step of processing after refinement. The specific process is: based on the refinement number obtained by connecting the centroid pixels in sequence in step S35, select the first point as the uppermost point; look at the adjacent points in the clockwise direction on the curve at a distance of k pixels; if the row has been passed, check further clockwise, that is, search on another row until all rows are searched; connect the searched points in sequence to obtain a refinement number composed of several line segments with a length of k pixels, and segment the refinement number into several line segments with a length of k pixels.

[0026] Preferably, each corner pixel of the square matrix in step S32 must have at most one adjacent pixel that does not belong to the square matrix, wherein the corner pixel is a pixel at the outermost corner of the square matrix.

[0027] Preferably, the number of centroid pixels = the number of line segments + 1, and the number of angles in the angle information obtained in step S4 = the number of line segments - 1. Of course, the last line segment may be adjacent to the first line segment or not. In other words, the first centroid pixel point and the last centroid pixel point may coincide, for example, when recognizing the characters "0" and "8". The angle information acquisition process is as follows: according to the refined number composed of several line segments obtained in step S3, the angle formed by the continuous line segments u and v is calculated as a vector and The angle between two two-dimensional vectors is where ||u|| and ||v|| are vectors respectively and The norm of .

[0028] Preferably, the preliminary recognition results output by the neural network and the K-nearest neighbor algorithm model include several possible numbers and corresponding confidence levels, wherein the confidence levels are obtained according to their training process, and the specific process of selecting the highest confidence level in step S6 includes: accumulating the confidence levels corresponding to the same preliminary recognition results output by the neural network and the K-nearest neighbor algorithm model, comparing the confidence levels corresponding to several different preliminary recognition results, and finally determining the highest confidence level.

[0029] The present invention has the following beneficial effects: it improves the image preprocessing process, optimizes images collected in harsh environments, and improves their contrast; uses the angles between adjacent segmented lines as inputs to the neural network and K-nearest neighbor algorithm model, and compared with the algorithm using each pixel value, has the advantages of significantly smaller inputs, reduced noise weight, and digital rotation not affecting the neural network architecture and its input; and introduces the confidence obtained during the recognition model training process to further improve the accuracy of character recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is a flow chart of an identification method according to an embodiment of the present invention.

[0031] Figure 2 2 is a schematic diagram of the refinement segmentation process according to an embodiment of the present invention. DETAILED DESCRIPTION

[0032] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.

[0033] An optical character recognition system includes an image acquisition module, an image preprocessing module, an image digital thinning and segmentation module, a digital information processing module, a neural network recognition module, a K-nearest neighbor algorithm recognition module, and a recognition result processing and output module. The image acquisition module is connected to the image preprocessing module, which is connected to the image digital thinning and segmentation module. The image digital thinning and segmentation module reduces a number with width to a line and then segments the line into equal or unequal line segments. The main feature of the thinning process is that the length is maintained. The image digital thinning and segmentation module is connected to the digital information processing module. The digital information processing module obtains angle information between adjacent line segments based on the processing results of the image digital thinning and segmentation module. The digital information processing module is respectively connected to the neural network recognition module and the K-nearest neighbor algorithm recognition module. The neural network recognition module and the K-nearest neighbor algorithm recognition module are both connected to the recognition result processing and output module. The neural network recognition module and the K-nearest neighbor algorithm recognition module respectively input their preliminary recognition results and corresponding confidence information into the recognition result processing and output module. The recognition result processing and output module determines and outputs a final recognition result based on the input information.

[0034] The image preprocessing and image digital refinement and segmentation modules in the present invention can optimize images collected in harsh environments, improve their contrast, and facilitate further specific processing of subsequent images. At the same time, they combine two different recognition methods, namely the neural network recognition module and the K-nearest neighbor algorithm recognition module. Based on the training process of the two recognition models, the preliminary recognition results and corresponding confidence levels output by the two recognition models are integrated, and various possible recognition results and confidence levels of the input characters are listed, so that the one with the highest confidence level is selected as the final recognition result output. Finally, the information output by the two recognition modules is integrated to obtain the final recognition result, which further improves the accuracy of character recognition in harsh environments.

[0035] The present invention also provides an optical character recognition method. In this embodiment, Figure 1 As shown, the following steps are included:

[0036] S1: Image acquisition, the images are acquired using a webcam installed in front of the gas or electricity meter;

[0037] S2: Image preprocessing, converting the acquired image into a monochrome image or grayscale image. After conversion, the contrast of the image, which generally has multiple colors, is improved through histogram equalization. Through this adjustment, the intensity can be better distributed on the histogram, which allows areas of lower local contrast to achieve higher contrast without affecting the global contrast. Further sharpening processing and adaptive thresholding are performed, where the specific process of adaptive thresholding is: the image is segmented by setting all pixels with intensity values ​​above the threshold to the foreground value and all remaining pixels to the background value. After this stage, the image is ready for further specific processing;

[0038] S3: Image refinement and segmentation, e.g. Figure 2 As shown, the digits with width in the image are thinned and the thinned digits are divided into several line segments, including:

[0039] S31: Select a boundary pixel, which is the leftmost one;

[0040] S32: Construct a minimum square matrix that contains the selected pixels and does not overlap with existing adjacent matrices, where each corner pixel of the square matrix must have at most one adjacent pixel that does not belong to the square matrix, wherein the corner pixel is a pixel at the outermost corner of the square matrix;

[0041] S33: Calculate the centroid pixel of the square matrix. If the center of the square matrix is ​​a 2*2 matrix pixel, take any pixel in the 2*2 matrix pixel as the centroid pixel, and replace the square matrix with the centroid pixel.

[0042] S34: Return to step S31 until there are no more available pixels;

[0043] S35: sequentially connecting the centroid pixels to obtain a number consisting of a plurality of line segments, that is, dividing the refined number into a plurality of line segments;

[0044] S4: Inputting the angle information between adjacent line segments into the neural network and the K-nearest neighbor algorithm model for digital recognition;

[0045] S5: The neural network and K-nearest neighbor algorithm model obtain possible preliminary recognition results and corresponding confidence levels;

[0046] S6: Select the preliminary recognition result with the highest confidence as the final recognition result and output it. The preliminary recognition results output by the neural network and the K-nearest neighbor algorithm model include several possible numbers and corresponding confidence levels, where the confidence levels are obtained based on the confusion matrix constructed during their training process. The specific process of selecting the highest confidence level in step S6 includes: accumulating the confidence levels corresponding to the same preliminary recognition results output by the neural network and the K-nearest neighbor algorithm model, comparing the confidence levels corresponding to several different preliminary recognition results, and finally determining the highest confidence level.

[0047] By refining and segmenting the captured image and combining it with two recognition models, an artificial neural network and a K-nearest neighbor algorithm, the accuracy of character recognition in complex environments is ultimately improved. Using the angles between adjacent segmented segments as input to the neural network and K-nearest neighbor models offers the following advantages over algorithms that use individual pixel values: (1) the input is significantly smaller, (2) the weight of noise is reduced, and (3) the rotation of the digits does not affect the neural network architecture or its input.

[0048] In addition, the number of centroid pixels in step S3 = the number of line segments + 1, and the number of angles in the angle information obtained in step S4 = the number of line segments - 1. Of course, the last line segment may be adjacent to the first line segment, or it may not be adjacent. In other words, the first centroid pixel point and the last centroid pixel point may coincide, for example, when recognizing the characters "0" and "8". The angle information acquisition process is as follows: according to the refined number composed of several line segments obtained in step S3, the angle formed by the continuous line segments u and v is calculated as a vector and The angle between two two-dimensional vectors is where ||u|| and ||v|| are vectors respectively and The norm of .

[0049] The recognition system and recognition method of the present invention improve the image preprocessing process, optimize images collected in harsh environments, and improve their contrast; the angles between adjacent segmented lines are used as inputs to the neural network and K-nearest neighbor algorithm model. Compared with the algorithm using each pixel value, the system has the advantages of significantly smaller inputs, reduced noise weight, and the rotation of numbers not affecting the neural network architecture and its input; the confidence level obtained during the recognition model training process is introduced to further improve the accuracy of character recognition.

[0050] Example 2:

[0051] The difference between this embodiment and embodiment 1 is that the image thinning and image segmentation processing in step S3 are performed separately, and segmentation is the next step after thinning. The specific process is: based on the thinning number obtained by sequentially connecting the centroid pixels in step S35, select the first point as the uppermost point; look at the adjacent points in the clockwise direction on the curve at a distance of k pixels; if the row has been passed, further check clockwise, that is, search on another row until all rows are searched; sequentially connect the searched points to obtain a thinning number composed of several line segments with a length of k pixels, that is, the thinning number is segmented into several line segments with a length of k pixels. The angle formed by consecutive adjacent line segments is sequentially obtained as angle information, such as taking adjacent line segments u and v, and the angle formed by them is calculated as a vector and The angle between two two-dimensional vectors is where ||u|| and ||v|| are vectors respectively and The norm of .

[0052] The above-described embodiments are merely preferred implementations of the present invention and are not intended to limit the present invention in any form. Other variations and modifications are possible without exceeding the technical solutions described in the claims.

Claims

1. An optical character recognition method, characterized in that: The following steps are involved: S1: Image acquisition; S2: image preprocessing; S3: Image thinning and segmentation, thinning the digits with width in the image and segmenting the thinned digits into several line segments, selecting a boundary pixel, constructing the smallest square matrix that contains the selected pixels and does not overlap with the existing adjacent matrices, calculating the centroid pixel of the square matrix, and if the center of the square matrix is ​​a 2*2 matrix pixel, taking any pixel in the 2*2 matrix pixel as the centroid pixel and replacing the square matrix with the centroid pixel, reselecting boundary pixels until there are no more available pixels, and sequentially connecting the centroid pixels to obtain a digit consisting of several line segments; S4: Inputting the angle information between adjacent line segments into the neural network and the K-nearest neighbor algorithm model for digital recognition; S5: The neural network and K-nearest neighbor algorithm model obtain possible preliminary recognition results and corresponding confidence levels; S6: Select the preliminary recognition result with the highest confidence as the final recognition result and output it.

2. An optical character recognition method according to claim 1, characterized in that: The step S2 specifically includes: converting the acquired image into a monochrome image or a grayscale image.

3. The optical character recognition method according to claim 2, wherein: The S2 also includes sharpening processing and adaptive threshold processing, wherein the specific process of the adaptive threshold processing is: segmenting the image by setting all pixels with intensity values ​​higher than the threshold as foreground values ​​and setting all remaining pixels as background values.

4. The optical character recognition method according to claim 1, wherein: Each corner pixel of the square matrix in S3 must have at most one adjacent pixel that does not belong to the square matrix.

5. The optical character recognition method according to claim 1, wherein: The number of centroid pixels=the number of line segments+1, and the number of angles in the angle information obtained in step S4=the number of line segments-1.

6. The optical character recognition method according to claim 1, wherein: Also includes: The preliminary recognition results output by the neural network and the K-nearest neighbor algorithm model include several possible numbers and corresponding confidence levels, where the confidence levels are obtained based on their training process. The specific process of selecting the highest confidence level in step S6 includes: accumulating the confidence levels corresponding to the same preliminary recognition results output by the neural network and the K-nearest neighbor algorithm model, comparing the confidence levels corresponding to several different preliminary recognition results, and finally determining the highest confidence level.

7. An optical character recognition system, executing an optical character recognition method according to any one of claims 1 to 6, characterized in that: It includes an image acquisition module, an image preprocessing module, an image digital refinement and segmentation module, a digital information processing module, a neural network recognition module, a K-nearest neighbor algorithm recognition module and a recognition result processing and output module. The image acquisition module is connected to the image preprocessing module, the image preprocessing module is connected to the image digital refinement and segmentation module, the image digital refinement and segmentation module is connected to the digital information processing module, the digital information processing module is connected to the neural network recognition module and the K-nearest neighbor algorithm recognition module respectively, and the neural network recognition module and the K-nearest neighbor algorithm recognition module are both connected to the recognition result processing and output module.

8. An optical character recognition system according to claim 7, characterized in that: The image digital thinning and segmentation module thins the digits with width in the image and divides the thinned digits into a number of line segments. The digital information processing module obtains the angle information between adjacent line segments based on the processing results of the image digital thinning and segmentation module, and inputs the obtained angle information into the neural network recognition module and the K nearest neighbor algorithm recognition module respectively.

9. An optical character recognition system according to claim 7, characterized in that: The neural network recognition module and the K-nearest neighbor algorithm recognition module respectively input their preliminary recognition results and corresponding confidence information into the recognition result processing and output module, and the recognition result processing and output module determines the final recognition result based on the input information and outputs it.

Citation Information

Patent Citations

  • An OCR recognition method and terminal

    CN109871847B

  • Combustion gas index automatic identification method based on images

    CN106169080A

  • Character recognizing device

    JP1992177485A

  • Method for recognizing handwritten character

    JP1997027010A