Image recognition rate improving method and system based on Tesseract-OCR (Optical Character Recognition)
By performing image preprocessing and model training on Tesseract-OCR optical character recognition technology, and using long and short-term memory network LSTM, the problem of low image recognition accuracy in the prior art is solved, and a more efficient image recognition effect is achieved.
Patent Information
- Application Number
- CN202510039367.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-13
AI Technical Summary
The existing Tesseract-OCR optical character recognition technology uses general recurrent neural network RNN to memorize numerical values, and cannot effectively transmit and express information in long-term series, resulting in a decrease in the accuracy of image recognition.
The image recognition rate improvement method based on Tesseract-OCR is adopted to pre-process the input image (including grayscale processing, binary processing, denoising processing, contrast enhancement and tilt correction), and the box file is generated using the command line tool of Tesseract-OCR, and the model training is performed using the long and short-term memory network LSTM.
By optimizing the image recognition process and using the LSTM network, the information in a long time series is effectively transmitted and expressed, and the accuracy and efficiency of image recognition are improved.
Smart Images

Figure CN119992571A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and system for improving image recognition rate based on Tesseract-OCR. Background Art
[0002] With the widespread application of text recognition technology in many fields, Tesseract-OCR optical character recognition technology came into being. This technology was developed by HP Labs and later maintained by Google as an open source OCR engine. OCR is the abbreviation of Optical Character Recognition, which means optical character recognition. It is a technology that recognizes printed or handwritten text through computer software.
[0003] Currently, the existing Tesseract-OCR optical character recognition technology uses a general recurrent neural network RNN to memorize numerical values.
[0004] However, the existing technology uses general recurrent neural networks (RNNs) to memorize numerical values, which cannot effectively transmit and express information in long time series, so that useful information from a long time ago will be ignored, which will lead to a decrease in the accuracy of image recognition. Summary of the invention
[0005] The embodiments of the present invention provide a method and system for improving image recognition rate based on Tesseract-OCR, which can improve the accuracy of image recognition.
[0006] In a first aspect, an embodiment of the present invention provides a method for improving image recognition rate based on Tesseract-OCR, the method comprising:
[0007] Obtain at least one first input image in a training data set;
[0008] Performing an image preprocessing operation on the at least one first input image to generate a preprocessed image of the first input image, wherein the image preprocessing operation includes: grayscale processing, binary processing, denoising processing, contrast enhancement and tilt correction;
[0009] Using a command line tool of Tesseract-OCR, generating a box file corresponding to the preprocessed image, wherein the box file includes at least one character position and character size information, and the Tesseract-OCR uses a long short-term memory network LSTM;
[0010] A character set file is generated based on the box file, and a model training is performed on the Tesseract-OCR based on the character set file and the training data set.
[0011] Preferably,
[0012] Before performing an image preprocessing operation on the at least one first input image to generate a preprocessed image of the first input image, and before using the Tesseract-OCR command line tool to generate a box file corresponding to the preprocessed image, the method further includes:
[0013] Convert the pre-processed image into a tif image format using a pre-installed ImageMagick auxiliary tool;
[0014] The Tesseract-OCR command line tool is used to generate a box file corresponding to the preprocessed image, including:
[0015] The Tesseract-OCR command line tool is used to generate a box file corresponding to the tif image format.
[0016] Preferably,
[0017] After generating the box file corresponding to the tif format using the command line tool of the Tesseract-OCR, generating a character set file based on the box file, and before training the Tesseract-OCR model based on the character set file and the training data set, further comprising:
[0018] Specify at least one page segmentation mode using the --psm parameter of the Tesseract-OCR, and determine a target page segmentation mode based on actual needs, wherein different page segmentation modes correspond to different first input images;
[0019] The box file is corrected using the pre-installed jTessBoxEditor auxiliary tool, and the corrected target box file is stored.
[0020] Preferably,
[0021] The generating of the character set file based on the box file comprises:
[0022] Based on the target box file, the character set file is extracted from the training data set, and the output file name, output path and character set extraction range of the character set file are determined.
[0023] Preferably,
[0024] After generating a character set file based on the box file, and performing model training on the Tesseract-OCR based on the character set file and the training data set, further comprising:
[0025] Performing a model test on the Tesseract-OCR using a test data set, wherein the test data set and the training data set contain similar font, size, and background features;
[0026] Using the Tesseract-OCR command line tool to read the second input image in the test data set, and output the current test result;
[0027] The Tesseract-OCR model is optimized based on the current test result.
[0028] In a second aspect, an embodiment of the present invention provides an image recognition rate improvement system based on Tesseract-OCR, the system comprising:
[0029] An acquisition module, used to acquire at least one first input image in a training data set;
[0030] A first processing module, configured to perform an image preprocessing operation on the at least one first input image to generate a preprocessed image of the first input image, wherein the image preprocessing operation includes: grayscale processing, binary processing, denoising processing, contrast enhancement and tilt correction;
[0031] A second processing module is used to generate a box file corresponding to the preprocessed image using a command line tool of Tesseract-OCR, wherein the box file includes at least one character position and character size information, and the Tesseract-OCR uses a long short-term memory network LSTM;
[0032] A training module is used to generate a character set file based on the box file, and perform model training on the Tesseract-OCR based on the character set file and the training data set.
[0033] Preferably,
[0034] After the first processing module and before the second processing module, it further includes: a format conversion module;
[0035] The format conversion module is used to convert the pre-processed image into a tif image format using a pre-installed ImageMagick auxiliary tool;
[0036] The second processing module is also used to generate a box file corresponding to the tif image format using the command line tool of the Tesseract-OCR.
[0037] Preferably,
[0038] After the second processing module and before the training module, it further includes: a correction module;
[0039] The correction module is used to perform:
[0040] Specify at least one page segmentation mode using the --psm parameter of the Tesseract-OCR, and determine a target page segmentation mode based on actual needs, wherein different page segmentation modes correspond to different first input images;
[0041] The box file is corrected using the pre-installed jTessBoxEditor auxiliary tool, and the corrected target box file is stored.
[0042] Preferably,
[0043] The training module is also used to extract the character set file from the training data set based on the target box file, and determine the output file name, output path and character set extraction range of the character set file.
[0044] Preferably,
[0045] After the training module, it further includes: a testing module;
[0046] The test module is used to perform:
[0047] Performing a model test on the Tesseract-OCR using a test data set, wherein the test data set and the training data set contain similar font, size, and background features;
[0048] Using the Tesseract-OCR command line tool to read the second input image in the test data set, and output the current test result;
[0049] The Tesseract-OCR model is optimized based on the current test result.
[0050] In a third aspect, an embodiment of the present invention provides a system for improving image recognition rate based on Tesseract-OCR, including: at least one memory and at least one processor;
[0051] The at least one memory is used to store a machine-readable program;
[0052] The at least one processor is used to call the machine-readable program to execute any method described in the first aspect.
[0053] In a fourth aspect, an embodiment of the present invention provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor executes any one of the methods described in the first aspect.
[0054] The embodiment of the present invention provides a method and system for improving the image recognition rate based on Tesseract-OCR. Tesseract-OCR is an open source OCR engine, which is an optical character recognition technology that can recognize printed or handwritten text through computer software. It uses a deep learning method to perform text recognition, and improves the accuracy of image recognition by processing pictures and training models of data files. Therefore, the implementation of this method is based on Tesseract-OCR to perform character recognition and extraction on the first input image, and improves the accuracy of image recognition by optimizing the recognition process, specifically performing preprocessing operations on the first input image to be recognized in the training data set and training the model on the training data set. Image preprocessing is one of the key steps to improve the accuracy of OCR. By performing a series of preprocessing operations on the first input image, the image quality can be optimized and the recognition ability of the OCR engine to the text area can be improved. The training data set file is a variety of different types of images collected from the Internet or locally, with clear text and sufficient diversity, which can cover all the characters and fonts you want to recognize. Using the training data set to train the model can further improve the accuracy of image recognition. Through the implementation of the above process, the existing Tesseract-OCR image recognition process can be optimized. At the same time, the use of long short-term memory network LSTM can effectively transmit and express information in long time series, so that useful information from a long time ago is not ignored, further improving the accuracy of image recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0056] Figure 1 It is a flow chart of a method for improving image recognition rate based on Tesseract-OCR provided by one embodiment of the present invention;
[0057] Figure 2 is a flowchart of another method for improving image recognition rate based on Tesseract-OCR provided by an embodiment of the present invention;
[0058] Figure 3is a schematic diagram of a system for improving image recognition rate based on Tesseract-OCR provided by an embodiment of the present invention;
[0059] Figure 4 is a schematic diagram of another system for improving image recognition rate based on Tesseract-OCR provided by an embodiment of the present invention;
[0060] Figure 5 is a schematic diagram of another system for improving image recognition rate based on Tesseract-OCR provided by an embodiment of the present invention;
[0061] Figure 6 1 is a schematic diagram of another system for improving image recognition rate based on Tesseract-OCR provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0063] like Figure 1 As shown, an embodiment of the present invention provides a method for improving image recognition rate based on Tesseract-OCR, and the method may include the following steps:
[0064] Step 101: Obtain at least one first input image in a training data set;
[0065] Step 102: performing an image preprocessing operation on at least one first input image to generate a preprocessed image of the first input image, wherein the image preprocessing operation includes: grayscale processing, binary processing, denoising processing, contrast enhancement and tilt correction;
[0066] Step 103: Generate a box file corresponding to the preprocessed image using the command line tool of Tesseract-OCR, where Tesseract-OCR uses a long short-term memory network LSTM, and the box file includes at least one character position and character size information;
[0067] Step 104: Generate a character set file based on the box file, and train the Tesseract-OCR model based on the character set file and the training data set.
[0068] In an embodiment of the present invention, a method for improving image recognition rate based on Tesseract-OCR is provided. Tesseract-OCR is an open source OCR engine, which is an optical character recognition technology that can recognize printed or handwritten text through computer software. It uses a deep learning method for text recognition, and improves the accuracy of image recognition by processing pictures and training models of data files. Therefore, the implementation of this method is based on Tesseract-OCR to perform character recognition and extraction on the first input image, and improves the accuracy of image recognition by optimizing the recognition process, specifically performing preprocessing operations on the first input image to be recognized in the training data set and training the model on the training data set. Image preprocessing is one of the key steps to improve the accuracy of OCR. By performing a series of preprocessing operations on the first input image, the image quality can be optimized and the recognition ability of the OCR engine to the text area can be improved. The training data set file is a variety of different types of images collected from the Internet or locally, with clear text and sufficient diversity, which can cover all the characters and fonts you want to recognize. Using the training data set to train the model can further improve the accuracy of image recognition. Through the implementation of the above process, the existing Tesseract-OCR image recognition process can be optimized. At the same time, the use of long short-term memory network LSTM can effectively transmit and express information in long time series, so that useful information from a long time ago is not ignored, further improving the accuracy of image recognition.
[0069] Specifically, the image preprocessing method used in this embodiment includes the following aspects:
[0070] 1. Grayscale. Grayscale is the process of converting a color image into a grayscale image. In OCR, grayscale can remove color information from an image, reduce the amount of computation, and retain enough image information for the OCR engine to recognize. For example, you can use OpenCV and other libraries for grayscale processing;
[0071] 2. Binarization. Binarization is the process of converting a grayscale image into a binary image, that is, an image containing only black and white. In OCR, binarization can further simplify image information, highlight text areas, and improve the OCR engine's ability to recognize text. Commonly used binarization methods include global thresholding and local thresholding.
[0072] 3. De-noising. The noise in the image may affect the OCR engine's recognition of the text area. Therefore, the image needs to be denoised in the preprocessing stage. Common denoising methods include median filtering and Gaussian filtering. The above filtering methods can effectively remove noise points in the image and improve image quality;
[0073] 4. Contrast enhancement. Contrast enhancement can improve the contrast between the text and the background in the image, making the text more visible. Common contrast enhancement methods include histogram equalization, Laplace sharpening, etc. The above contrast enhancement methods can adjust the brightness and contrast of the image to make the text area more prominent;
[0074] 5. Tilt correction. In actual applications, the input image may be tilted. Tilt correction can adjust the direction of the image so that the text is arranged horizontally. Commonly used tilt correction methods include Hough transform and projection method. The above tilt correction methods can detect and correct the tilt angle in the image and improve the OCR engine's ability to recognize text.
[0075] In order to improve the efficiency of image recognition, in one embodiment of the present invention, in the above embodiment, after performing an image preprocessing operation on at least one first input image to generate a preprocessed image of the first input image in step 102, before generating a box file corresponding to the preprocessed image using the Tesseract-OCR command line tool in step 103, further includes:
[0076] Convert the pre-processed image into a tif image format using a pre-installed ImageMagick auxiliary tool;
[0077] The Tesseract-OCR command line tool is used to generate a box file corresponding to the preprocessed image, including:
[0078] The Tesseract-OCR command line tool is used to generate a box file corresponding to the tif image format.
[0079] In an embodiment of the present invention, since the tif image format is a commonly used image format with the characteristics of lossless compression and high-quality output, it is very suitable for OCR training. Therefore, in order to improve the efficiency of image recognition, the first input image in the training data set can be converted into a tif image format using the pre-installed ImageMagick auxiliary tool, and then the Tesseract-OCR command line tool needs to be used to generate a box file corresponding to the tif image file. The box file contains the character position and character size information of each character in the tif file, which is important data for training Tesseract-OCR.
[0080] In order to correct the box file, in one embodiment of the present invention, in the above embodiment, after using the command line tool of the Tesseract-OCR in the step to generate the box file corresponding to the tif format, the character set file is generated based on the box file, and before the Tesseract-OCR model is trained based on the character set file and the training data set, it further includes:
[0081] Specify at least one page segmentation mode using the --psm parameter of the Tesseract-OCR, and determine a target page segmentation mode based on actual needs, wherein different page segmentation modes correspond to different first input images;
[0082] The box file is corrected using the pre-installed jTessBoxEditor auxiliary tool, and the corrected target box file is stored.
[0083] In an embodiment of the present invention, the --psm parameter of Tesseract-OCR can be used to specify the page segmentation mode when generating a box file. Different page segmentation modes are suitable for different types of images, and the appropriate mode can be selected according to the specific situation (for example, different scenes can be specified for parameters 0-11, and parameter 7 is specified to convert the image into a single text line). Since there may be errors in the automatic character recognition of Tesseract-OCR, the generated box file may contain some erroneous character position information, so it is necessary to use the pre-installed auxiliary tool jTessBoxEditor to correct the box file. After opening the tif image file and the corresponding box file in jTessBoxEditor, check the position and size information of each character one by one, and make necessary adjustments. After the correction is completed, save the modified target box file.
[0084] In order to train the model, in one embodiment of the present invention, the step 104 in the above embodiment of generating a character set file based on the box file includes:
[0085] Based on the target box file, extract the character set file from the training data set, and determine the output file name, output path and character set extraction range of the character set file;.
[0086] In an embodiment of the present invention, the character set file contains all the character information required for training Tesseract-OCR, so the character set file can be extracted from the training data set using the command line tool of Tesseract-OCR. When extracting the character set file, it is necessary to specify an output file name and output path, and specify the character set range to be extracted (such as the ASCII code range). After the extraction is completed, a text file containing all the character information will be obtained to train the model.
[0087] In order to verify the performance of the trained OCR model, in one embodiment of the present invention, in the above embodiment, after generating a character set file based on the box file in step 104 and training the Tesseract-OCR model based on the character set file and the training data set, the method further includes:
[0088] Performing a model test on the Tesseract-OCR using a test data set, wherein the test data set and the training data set contain similar font, size, and background features;
[0089] Using the Tesseract-OCR command line tool to read the second input image in the test data set, and output the current test result;
[0090] The Tesseract-OCR model is optimized based on the current test result.
[0091] In an embodiment of the present invention, in order to verify the performance of the trained OCR model, a test data set can be used to test it. The test data set should have similar font, size, background and other features as the training data set. During the test, the command line tool or API of Tesseract-OCR can be used to read the first input image and output the current test result. The OCR model can be optimized and adjusted according to the current test result. By continuously optimizing and adjusting the model, its recognition rate and robustness can be further improved.
[0092] In the embodiment of the present invention, due to the use of optical character recognition technology based on Tesseract-OCR, the optimization of the preprocessing process of the first input image, and the model training of the training data set file, it is only necessary to train the data set file in advance according to the needs and preprocess the image, which can greatly improve the accuracy of image recognition. Because the long short-term memory network LSTM is used, the long-term dependency problem commonly existing in general recursive neural networks is solved, and the information in the long time series can be effectively transmitted and expressed, and the useful information from a long time ago will not be ignored, which solves the problems of manual input of picture text in traditional image recognition and low automatic recognition rate.
[0093] In the embodiment of the present invention, text recognition technology has a wide range of applications in many fields, such as graphic verification code recognition, document processing, automated office, text input on mobile devices, etc. As an open source OCR engine, Tesseract-OCR has received widespread attention and application for its efficient and accurate text recognition capabilities. Tesseract-OCR uses a deep learning method for text recognition and can recognize multiple languages, including English, Chinese, German, French, etc. It can efficiently and accurately recognize characters in pictures, helping users extract text information from pictures, and is widely used in crawlers, PDF recognition, and text input on mobile devices.
[0094] like Figure 2 As shown, in order to more clearly illustrate the technical solutions and advantages of the present invention, a method for improving image recognition rate based on Tesseract-OCR is described in detail below in an embodiment of the present invention, which may specifically include the following steps:
[0095] Step 201: Obtain at least one first input image in a training data set;
[0096] Step 202: performing an image preprocessing operation on at least one first input image to generate a preprocessed image of the first input image, wherein the image preprocessing operation includes: grayscale processing, binary processing, denoising processing, contrast enhancement and tilt correction;
[0097] Step 203: using the pre-installed ImageMagick auxiliary tool to convert the pre-processed image into a tif image format;
[0098] Specifically, the original Tesseract only works well in a controlled environment. If the image has too much background noise or is out of focus, Tesseract will not work properly. To overcome this problem, Tesseract has used deep learning models to recognize characters and even handwriting since 4.0. Tesseract 4.0 uses a long short-term memory network (LSTM) to improve the accuracy of its OCR engine. First, you need to download and install the Tesseract-OCR software and other necessary auxiliary tools. You can download the latest version of the installation package from the official website for installation. For different operating systems, you need to select the corresponding installation package for download and installation. In addition, some additional auxiliary tools (such as jTessBoxEditor, ImageMagick) are required to process images and box files more conveniently.
[0099] Step 204: Generate a box file corresponding to the tif image format using the command line tool of Tesseract-OCR, wherein the box file includes at least one character position and character size information, and Tesseract-OCR uses a long short-term memory network LSTM;
[0100] Specifically, OCR is the abbreviation of Optical Character Recognition, which refers to the process of an electronic device (such as a scanner or digital camera) checking the characters printed on paper, determining their shapes by detecting dark and light patterns, and then translating the shapes into computer text using character recognition methods; that is, for printed characters, using optical methods to convert the text in paper documents into black and white dot matrix image files, and using recognition software to convert the text in the image into text format for further editing and processing by word processing software. LSTM (Long Short-Term Memory Network) is a time recurrent neural network that is specially designed to solve the long-term dependency problem of general RNN (Recurrent Neural Network). All RNNs have a chain form of repeated neural network modules. LSTM is a type of neural network containing LSTM blocks or other neural networks. In literature or other materials, LSTM blocks may be described as intelligent network units because they can remember values of indefinite lengths of time. There is a gate in the block that can determine whether the input is important enough to be remembered and whether it can be output.
[0101] Step 205: using the --psm parameter of Tesseract-OCR to specify at least one page segmentation mode, and determining a target page segmentation mode based on actual needs, wherein different page segmentation modes correspond to different first input images;
[0102] Step 206: Correct the box file using the pre-installed jTessBoxEditor auxiliary tool, and store the corrected target box file;
[0103] Step 207: Based on the target box file, extract the character set file from the training data set, determine the output file name, output path and character set extraction range of the character set file, and perform model training on Tesseract-OCR based on the character set file and the training data set;
[0104] Specifically, you can use the command line tool of Tesseract-OCR to specify the training parameters and configuration files and start the training process. During the training process, Tesseract-OCR will use the training data set to optimize the model parameters to improve its image recognition rate. The training time depends on the size and complexity of the training data set and the performance of the computer. After the training is completed, an OCR model that can be used for practical applications will be obtained.
[0105] Step 208: Testing the Tesseract-OCR model using a test data set, wherein the test data set and the training data set contain similar font, size, and background features;
[0106] Step 209: Use the Tesseract-OCR command line tool to read the second input image in the test data set and output the current test result;
[0107] Step 210: Optimize the Tesseract-OCR model based on the current test results.
[0108] like Figure 3 As shown, an embodiment of the present invention provides an image recognition rate improvement system based on Tesseract-OCR, the system comprising:
[0109] An acquisition module 301 is used to acquire at least one first input image in a training data set;
[0110] A first processing module 302, configured to perform an image preprocessing operation on the at least one first input image to generate a preprocessed image of the first input image, wherein the image preprocessing operation includes: grayscale processing, binary processing, denoising, contrast enhancement and tilt correction;
[0111] The second processing module 303 is used to generate a box file corresponding to the pre-processed image using a command line tool of Tesseract-OCR, wherein the box file includes at least one character position and character size information, and the Tesseract-OCR uses a long short-term memory network LSTM;
[0112] The training module 304 is used to generate a character set file based on the box file, and perform model training on the Tesseract-OCR based on the character set file and the training data set.
[0113] based on Figure 3 A system for improving image recognition rate based on Tesseract-OCR is shown in Figure 4As shown, after the first processing module 302 and before the second processing module 303, it further includes: a format conversion module 305;
[0114] The format conversion module 305 is used to convert the pre-processed image into a tif image format using a pre-installed ImageMagick auxiliary tool;
[0115] The second processing module is also used to generate a box file corresponding to the tif image format using the command line tool of the Tesseract-OCR.
[0116] based on Figure 4 A system for improving image recognition rate based on Tesseract-OCR is shown in Figure 5 As shown, after the second processing module 303 and before the training module 304, it further includes: a correction module 306;
[0117] The correction module 306 is used to perform:
[0118] Specify at least one page segmentation mode using the --psm parameter of the Tesseract-OCR, and determine a target page segmentation mode based on actual needs, wherein different page segmentation modes correspond to different first input images;
[0119] The box file is corrected using the pre-installed jTessBoxEditor auxiliary tool, and the corrected target box file is stored.
[0120] like Figure 5 As shown, the training module 304 is also used to extract the character set file from the training data set based on the target box file, and determine the output file name, output path and character set extraction range of the character set file.
[0121] based on Figure 5 A system for improving image recognition rate based on Tesseract-OCR is shown in Figure 6 As shown, after the training module 304, it further includes: a testing module 307;
[0122] The test module 307 is used to perform:
[0123] Performing a model test on the Tesseract-OCR using a test data set, wherein the test data set and the training data set contain similar font, size, and background features;
[0124] Using the Tesseract-OCR command line tool to read the second input image in the test data set, and output the current test result;
[0125] The Tesseract-OCR model is optimized based on the current test result.
[0126] It is understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on a system for improving image recognition rate based on Tesseract-OCR. In other embodiments of the present invention, a system for improving image recognition rate based on Tesseract-OCR may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0127] The information interaction, execution process and other contents between the units in the above-mentioned device are based on the same concept as the embodiment of the method of the present invention. For specific contents, please refer to the description in the embodiment of the method of the present invention, and no further description is given here.
[0128] The embodiment of the present invention also provides a system for improving image recognition rate based on Tesseract-OCR, comprising: at least one memory and at least one processor;
[0129] at least one memory for storing a machine-readable program;
[0130] At least one processor is used to call a machine-readable program to execute a method for improving image recognition rate based on Tesseract-OCR in any embodiment of the present invention.
[0131] An embodiment of the present invention further provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor executes a method for improving image recognition rate based on Tesseract-OCR in any embodiment of the present invention.
[0132] Specifically, a system or device equipped with a storage medium can be provided, on which software program code that implements the functions of any of the above-mentioned embodiments is stored, and a computer (or CPU or MPU) of the system or device can be enabled to read and execute the program code stored in the storage medium.
[0133] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.
[0134] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer by a communication network.
[0135] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.
[0136] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.
[0137] Each embodiment of the present invention has at least the following beneficial effects:
[0138] 1. In an embodiment of the present invention, a method for improving image recognition rate based on Tesseract-OCR is provided. Tesseract-OCR is an open source OCR engine, which is an optical character recognition technology that can recognize printed or handwritten text through computer software. It uses a deep learning method for text recognition, and improves the accuracy of image recognition by processing pictures and training models of data files. Therefore, the implementation of this method is based on Tesseract-OCR to perform character recognition and extraction on the first input image, and improves the accuracy of image recognition by optimizing the recognition process, specifically performing preprocessing operations on the first input image to be recognized in the training data set and training the model on the training data set. Image preprocessing is one of the key steps to improve the accuracy of OCR. By performing a series of preprocessing operations on the first input image, the image quality can be optimized and the recognition ability of the OCR engine to the text area can be improved. The training data set file is a variety of different types of images collected from the Internet or locally, with clear text and sufficient diversity, which can cover all the characters and fonts you want to recognize. Using the training data set to train the model can further improve the accuracy of image recognition. Through the implementation of the above process, the existing Tesseract-OCR image recognition process can be optimized. At the same time, the use of long short-term memory network LSTM can effectively transmit and express information in long time series, so that useful information from a long time ago is not ignored, further improving the accuracy of image recognition;
[0139] 2. In the embodiment of the present invention, since the tif image format is a commonly used image format with the characteristics of lossless compression and high-quality output, it is very suitable for OCR training. Therefore, in order to improve the efficiency of image recognition, the first input image in the training data set can be converted into the tif image format using the pre-installed ImageMagick auxiliary tool, and then the Tesseract-OCR command line tool needs to be used to generate a box file corresponding to the tif image file. The box file contains the character position and character size information of each character in the tif file, which is important data for training Tesseract-OCR;
[0140] 3. In an embodiment of the present invention, the --psm parameter of Tesseract-OCR can be used to specify the page segmentation mode when generating a box file. Different page segmentation modes are suitable for different types of images, and the appropriate mode can be selected according to the specific situation (for example, different scenes can be specified for parameters 0-11, and parameter 7 is specified to convert the image into a single text line). Since there may be errors in the automatic character recognition of Tesseract-OCR, the generated box file may contain some erroneous character position information, so it is necessary to use the pre-installed auxiliary tool jTessBoxEditor to correct the box file. After opening the tif image file and the corresponding box file in jTessBoxEditor, check the position and size information of each character one by one, and make necessary adjustments. After the correction is completed, save the modified target box file.
[0141] It should be noted that not all steps and modules in the above-mentioned processes and system structure diagrams are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The system structure described in the above-mentioned embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or some components in multiple independent devices may be implemented together.
[0142] In the above embodiments, the hardware unit can be realized by mechanical means or electrical means. For example, a hardware unit can include permanent dedicated circuits or logic (such as special processors, FPGA or ASIC) to complete the corresponding operation. The hardware unit can also include programmable logic or circuits (such as general-purpose processors or other programmable processors), which can be temporarily set by software to complete the corresponding operation. Concrete implementation (mechanical means or dedicated permanent circuits or temporarily set circuits) can be determined based on cost and time considerations.
[0143] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. The image recognition rate improvement method based on Tesseract-OCR is characterized by: The method includes: Obtain at least one first input image in a training data set; Performing an image preprocessing operation on the at least one first input image to generate a preprocessed image of the first input image, wherein the image preprocessing operation includes: grayscale processing, binary processing, denoising processing, contrast enhancement and tilt correction; Using a command line tool of Tesseract-OCR, generating a box file corresponding to the preprocessed image, wherein the box file includes at least one character position and character size information, and the Tesseract-OCR uses a long short-term memory network LSTM; A character set file is generated based on the box file, and a model training is performed on the Tesseract-OCR based on the character set file and the training data set.
2. The method according to claim 1, characterized in that After performing an image preprocessing operation on the at least one first input image to generate a preprocessed image of the first input image, and before using the Tesseract-OCR command line tool to generate a box file corresponding to the preprocessed image, the method further includes: Convert the pre-processed image into a tif image format using a pre-installed ImageMagick auxiliary tool; The Tesseract-OCR command line tool is used to generate a box file corresponding to the preprocessed image, including: The Tesseract-OCR command line tool is used to generate a box file corresponding to the tif image format.
3. The method according to claim 2, characterized in that After generating the box file corresponding to the tif format using the command line tool of the Tesseract-OCR, generating a character set file based on the box file, and before training the Tesseract-OCR model based on the character set file and the training data set, further comprising: Specify at least one page segmentation mode using the --psm parameter of the Tesseract-OCR, and determine a target page segmentation mode based on actual needs, wherein different page segmentation modes correspond to different first input images; The box file is corrected using the pre-installed jTessBoxEditor auxiliary tool, and the corrected target box file is stored.
4. The method according to claim 3, characterized in that The generating of the character set file based on the box file comprises: Based on the target box file, extract the character set file from the training data set, and determine the output file name, output path and character set extraction range of the character set file; and / or, After generating a character set file based on the box file, and performing model training on the Tesseract-OCR based on the character set file and the training data set, further comprising: Performing a model test on the Tesseract-OCR using a test data set, wherein the test data set and the training data set contain similar font, size, and background features; Using the Tesseract-OCR command line tool to read the second input image in the test data set, and output the current test result; The Tesseract-OCR model is optimized based on the current test result.
5. The image recognition rate improvement system based on Tesseract-OCR is characterized by: The system includes: An acquisition module, used to acquire at least one first input image in a training data set; A first processing module, configured to perform an image preprocessing operation on the at least one first input image to generate a preprocessed image of the first input image, wherein the image preprocessing operation includes: grayscale processing, binary processing, denoising processing, contrast enhancement and tilt correction; A second processing module is used to generate a box file corresponding to the preprocessed image using a command line tool of Tesseract-OCR, wherein the box file includes at least one character position and character size information, and the Tesseract-OCR uses a long short-term memory network LSTM; A training module is used to generate a character set file based on the box file, and perform model training on the Tesseract-OCR based on the character set file and the training data set.
6. The system according to claim 5, characterized in that After the first processing module and before the second processing module, it further includes: a format conversion module; The format conversion module is used to convert the pre-processed image into a tif image format using a pre-installed ImageMagick auxiliary tool; The second processing module is also used to generate a box file corresponding to the tif image format using the command line tool of the Tesseract-OCR.
7. The system according to claim 6, characterized in that After the second processing module and before the training module, it further includes: a correction module; The correction module is used to perform: Specify at least one page segmentation mode using the --psm parameter of the Tesseract-OCR, and determine a target page segmentation mode based on actual needs, wherein different page segmentation modes correspond to different first input images; The box file is corrected using the pre-installed jTessBoxEditor auxiliary tool, and the corrected target box file is stored.
8. The system according to claim 7, characterized in that The training module is further used to extract the character set file from the training data set based on the target box file, and determine the output file name, output path and character set extraction range of the character set file; and / or, After the training module, it further includes: a testing module; The test module is used to perform: Performing a model test on the Tesseract-OCR using a test data set, wherein the test data set and the training data set contain similar font, size, and background features; Using the Tesseract-OCR command line tool to read the second input image in the test data set, and output the current test result; The Tesseract-OCR model is optimized based on the current test result.
9. The image recognition rate improvement system based on Tesseract-OCR is characterized by: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 4.
10. A computer readable medium, characterized in that The computer readable medium stores computer instructions, which, when executed by a processor, cause the processor to perform any one of the methods of claims 1 to 4.
Citation Information
Cited By
OCR (Optical Character Recognition) image recognition method and system based on large model self-learning
CN121459369A
Ocr image recognition method and system based on large model self-learning
CN121459369B