High-speed character recognition method based on PaddleOCRModel

By adopting a high-speed character recognition method based on PaddleOCRModel in character recognition, including data preprocessing and model inference, the problems of slow recognition speed and low accuracy in traditional methods are solved, and efficient and accurate character recognition is achieved.

CN120236288AInactive Publication Date: 2025-07-01SHENZHEN SIPO TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510713313.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional character recognition methods have problems such as slow recognition speed and low recognition accuracy, which is difficult to meet the needs of practical applications.

Method used

High-speed character recognition method based on PaddleOCRModel is adopted, including data loading, data preprocessing, model loading, model inference and post-processing steps. Data preprocessing includes resizing image, grayscale processing, normalization and standardization processing, model inference uses pre-trained PaddleOCR model, and combines post-processing steps to extract and splice recognition results.

Benefits of technology

It significantly improves the speed of character recognition, meets the real-time requirements in actual applications, and improves the accuracy of recognition, and is suitable for character recognition in various fonts, font sizes, and printing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236288A_ABST
    Figure CN120236288A_ABST
Patent Text Reader

Abstract

The invention discloses a high-speed character recognition method based on PaddleOCRModel, and the method comprises the steps: loading a to-be-recognized character picture into a code operation environment from a disk in an image data form through a data loading module, and enabling the character picture to comprise various characters with different fonts, font sizes and printing modes; carrying out preprocessing on the loaded image data so as to meet the input requirement of the PaddleOCR model; in the reasoning process, a pre-trained PaddleOCR model is loaded into an equipment memory; inputting the preprocessed image data into the loaded PaddleOCR model for reasoning to obtain a reasoning result, and performing post-processing on the reasoning result; according to the method, the data preprocessing and model reasoning processes are optimized, the character recognition speed can be remarkably increased, the requirement for real-time performance in practical application is met, reasoning is conducted through the pre-trained PaddleOCR model, the post-processing step is combined, the characters of various fonts, font sizes and printing modes can be accurately recognized, and the recognition accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of character recognition, and particularly relates to a high-speed character recognition method based on the PaddleOCRModel. Background Art

[0002] Character recognition is one of the important research directions in the field of computer vision, and its goal is to extract the character information in the image and convert it into text information that can be processed by a computer. Traditional character recognition methods often have problems such as slow recognition speed and low recognition accuracy, and it is difficult to meet the requirements of practical applications. With the continuous development of deep learning technology, character recognition methods based on deep learning have gradually emerged. PaddleOCR is an OCR tool based on deep learning launched by Baidu, which has advantages such as high efficiency and accuracy. However, in practical applications, how to further improve the recognition speed and accuracy of PaddleOCR is still a hot topic and a difficult point in current research. Summary of the Invention

[0003] In view of this, the main purpose of the present invention is to provide a high-speed character recognition method based on the PaddleOCRModel.

[0004] To achieve the above object, the technical solution of the present invention is realized as follows: The embodiment of the present invention provides a high-speed character recognition method based on the PaddleOCRModel, including the following steps: S1: Data loading module: Through the data loading module, the character image to be recognized is loaded from the disk into the code running environment in the form of image data, and the character image includes characters with various different glyphs, font sizes, and printing methods; S2: Data preprocessing module: Preprocess the loaded image data to meet the input requirements of the PaddleOCR model; S3: Model loading module: During the inference process, load the pre-trained PaddleOCR model into the device memory; S4: Model inference and post-processing module: Input the preprocessed image data into the loaded PaddleOCR model for inference, obtain the inference result, and perform post-processing on the inference result.

[0005] In the above solution, according to step S2: Preprocess the loaded image data, and the preprocessing method includes: S21: Adjust the image size, and use the bilinear interpolation method to expand the image to the target size to ensure that the size of the processed image is consistent with the input size required by the model; S22: Grayscale processing, and convert the color image to a grayscale image using the weighted average method; S23: Normalization and standardization processing, which confines the preprocessed data within the range of [0, 1] to eliminate the adverse effects caused by singular sample data.

[0006] In the above solution, according to step S22: Through the calculation formula of the grayscale processing: (gray(i,j)=0.229×R(i,j)+0.587×G(i,j)+0.114×B(i,j)); Among them, R(i,j), G(i,j), and B(i,j) respectively represent the values of the red, green, and blue color channels at the position (i,j) in the image, and convert the three-channel color picture into a single-channel grayscale image.

[0007] In the above solution, according to step S23: The normalization and standardization processing confines the target image data of the preprocessing within a certain range [0, 1] to eliminate the adverse effects caused by singular sample data, and its calculation formula is: normalized_pixel_value=(pixel_value - mean) / std); Among them, mean is the average value obtained by dividing the sum of all pixel values of the image by the number of all pixel values, and std is the standard deviation of the image.

[0008] In the above solution, according to step S4: The post-processing steps include: S41: Extract the confidence of each character from the inference result. The shape of the inference result is (batch, 40, 6625), where batch is the number of images inferred simultaneously, 40 is the maximum number of detected characters per row, and 6625 is the total number of Chinese, English, and various characters; S42: Traverse each unit (the confidence of 6625 characters) in the memory, find the character corresponding to the maximum confidence value, and splice to obtain the recognition result of each row.

[0009] In the above solution, the bilinear interpolation method calculates the target pixel value by performing interpolation once in the sub-horizontal and vertical directions respectively to ensure the smoothness and continuity of the image when magnifying and reducing.

[0010] In the above solution, this method is applicable to character recognition of various glyphs, font sizes, and printing methods, including but not limited to printed characters, handwritten characters, etc.

[0011] Compared with the prior art, the beneficial effects of the present invention: By optimizing the data preprocessing and model inference processes, the present invention can significantly improve the speed of character recognition, meet the real-time requirements in practical applications. By using a pre-trained PaddleOCR model for inference and combining post-processing steps, the present invention can accurately recognize characters of various glyphs, font sizes, and printing methods, improving the recognition accuracy. Description of the Drawings

[0012] The drawings described herein are used to disclose a further understanding of the present invention and form a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is a schematic structural diagram of a high-speed character recognition method based on the PaddleOCRModel described in the present invention; Figure 2 is a diagram 1 of the bilinear interpolation method in Resize in the present invention; Figure 3 is a diagram 2 of the bilinear interpolation method in Resize in the present invention. Detailed Embodiments

[0013] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0014] In the drawings of this embodiment, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that terms such as "first", "second", "third", etc. are only used to facilitate the description of the same components and do not indicate or imply the number of the components referred to, and should not be construed as a limitation of this patent. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0015] It should be noted that in this article, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, article, or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, article, or device comprising such element.

[0016] The present invention provides a high-speed character recognition method based on the PaddleOCRModel, as Figures 1 to 3 shown, including the following steps: S1: Data Loading Module: Through the data loading module, load the character images to be recognized into the code running environment in the form of image data from the disk. The character images include characters with various different glyphs, font sizes, and printing methods. S2: Data Preprocessing Module: Preprocess the loaded image data to meet the input requirements of the PaddleOCR model. S3: Model Loading Module: During the inference process, load the pre-trained PaddleOCR model into the device memory. S4: Model Inference and Post-processing Module: Input the preprocessed image data into the loaded PaddleOCR model for inference, obtain the inference result, and perform post-processing on the inference result.

[0017] In the above solution, according to step S2: Preprocess the loaded image data. The preprocessing method includes: S21: Resize the image. Use the bilinear interpolation method to expand the image to the target size to ensure that the size of the processed image is consistent with the input size required by the model. S22: Grayscale processing. Convert the color image to a grayscale image using the weighted average method. S23: Normalization and Standardization Processing. Limit the preprocessed data within the range of [0, 1] to eliminate the adverse effects caused by singular sample data.

[0018] In the above solution, according to step S22: The calculation formula for the grayscale processing is: (gray(i,j) = 0.229×R(i,j) + 0.587×G(i,j) + 0.114×B(i,j)); Among them, R(i,j), G(i,j), and B(i,j) respectively represent the values of the red, green, and blue color channels at the position (i,j) in the image, and convert the three-channel color image into a single-channel grayscale image.

[0019] In the above solution, according to step S23: For the normalization and standardization processing, limit the target image data of the preprocessing within a certain range of [0, 1] to eliminate the adverse effects caused by singular sample data. The calculation formula is: normalized_pixel_value = (pixel_value - mean) / std); Among them, mean is the average value obtained by dividing the sum of all pixel values of the image by the number of all pixel values, and std is the standard deviation of the image.

[0020] In the above solution, according to step S4: The post-processing steps include: S41: Extract the confidence of each character from the inference result. The shape of the inference result is (batch, 40, 6625), where batch is the number of images inferred simultaneously, 40 is the maximum number of detected characters per line, and 6625 is the total number of Chinese, English, and various characters. S42: Traverse each unit (the confidence of 6625 characters) in memory, find the character corresponding to the maximum confidence value, and splice to obtain the recognition result of each line.

[0021] In the above solution, the bilinear interpolation method calculates the target pixel value by performing interpolation once in the sub-horizontal and vertical directions respectively, ensuring the smoothness and continuity of the image when magnifying and reducing.

[0022] In the above solution, this method is applicable to character recognition of various glyphs, font sizes, and printing methods, including but not limited to printed characters, handwritten characters, etc.

[0023] Embodiment

[0024] Data loading module: First, prepare a series of character pictures containing various glyphs, font sizes, and printing methods. These pictures are saved on the disk in the form of image data. Through the data loading module of the present invention, these character pictures can be batch-loaded into the code running environment to provide input for subsequent data preprocessing and model inference.

[0025] Data preprocessing module: a. Resize the image: Select a character picture with a size of 640x480 as an example. According to the input requirements of the PaddleOCR model, the picture size needs to be adjusted to 320x240. The bilinear interpolation method is used to adjust the image size. Specifically, for each pixel (x’, y’) in the target image, map it to the original image to obtain the coordinates of two nearest neighbor pixels (x, y) and (x + 1, y + 1), and calculate the red, green, and blue values of the target pixel respectively. For example, for a pixel (160, 120) in the target image, its corresponding pixel coordinates in the original image can be calculated, and its red, green, and blue values can be calculated according to the bilinear interpolation formula.

[0026] b. Grayscale processing: Next, perform grayscale processing on the resized image. According to the weighted average method formula: (gray(i,j)=0.229×R(i,j)+0.587×G(i,j)+0.114×B(i,j)), traverse each pixel in the image, calculate its grayscale value, and convert the original color image into a grayscale image.

[0027] c. Normalization and standardization: Normalize and standardize the grayscale image. Calculate the mean and standard deviation (std) of all pixel values in the image, and then normalize each pixel value according to the formula: normalized_pixel_value = (pixel_value - mean) / std, so that it is limited within the range of [0, 1].

[0028] Model loading module: During the inference process, load the pre-trained PaddleOCR model into the device memory. This model has undergone a large amount of training and optimization, and has high recognition accuracy and generalization ability.

[0029] Model inference and post-processing module: Input the preprocessed image data into the loaded PaddleOCR model for inference. The inference result is a three-dimensional array with a shape of (batch, 40, 6625), where batch is the number of images inferred simultaneously (in this example, it is 1), 40 is the maximum number of detected characters per line, and 6625 is the total number of Chinese, English, and various characters. Each character corresponds to a confidence value from 0 to 1.

[0030] Traverse each unit in the inference result (i.e., the confidence of 6625 characters), find the character corresponding to the maximum confidence value, and use it as the recognition result at that position. Then, concatenate these characters to form the final line recognition result.

[0031] Experiment 1

[0032] To verify the effectiveness of the proposed solution of the present invention, this embodiment provides specific experimental data and implementation details.

[0033] First, the training dataset involved in the present invention includes a large number of real-scene images and synthetic images. Specifically, for the text recognition task, a total of 17.9M (million) training images are used, including 6.9M real-scene images and 16M synthetic images. These real-scene images are sourced from multiple public datasets such as LSVT, RCTW-17, MTWI 2018, and CCPD 2019, as well as Baidu Image Search. The synthetic images mainly focus on complex scenarios such as different backgrounds, translation, rotation, perspective transformation, line interference, noise, and vertical text, and their corpus is sourced from real-scene images. In addition, 18.7K (thousand) validation images are used, and all validation images are from real scenes.

[0034] During the experiment, to maintain the consistency of the hardware platform, tests were conducted based on the Windows 10 operating system, Intel Core i7 10700 CPU, and NVIDIA GeForce RTX 3060 12G graphics card. It was mainly based on three experimental scenarios: natural scenes, forms, and metal surfaces. For the objectivity of the experiment, the same number of test images (10,000) were used for testing and comparison in each scenario, and the consistency of the inference image size, inference device, and inference model was maintained.

[0035] The experimental results show that in the natural scene, the text recognition accuracy of the present invention reached 91.27%, and the inference speed was 12 ms per image; in the form scene, the accuracy was increased to 98.91%, and the inference speed was 14.7 ms per image; in the metal surface scene, the accuracy remained at 97.73%, while the inference speed was increased to 10.3 ms per image. These results indicate that the present invention has excellent text recognition performance and inference speed in different scenarios.

[0036] Experiment 2

[0037] To further illustrate the practicality and effectiveness of the present invention, this embodiment provides another set of experimental data and specific implementation methods.

[0038] Similarly, a total of 17.9M training images were used for model training, including real scene images and synthetic images. The validation dataset consisted of 18.7K real scene images. During the experiment, the consistency of the hardware platform was also maintained, and tests were conducted based on three experimental scenarios.

[0039] The difference is that this embodiment pays more attention to comparison and analysis in the experimental design. The effects of different amounts of training data on the model performance and the effects of different synthetic image ratios on the model generalization ability were respectively tested. The experimental results show that when the synthetic image ratio is appropriately increased, the generalization ability of the model is significantly improved, and at the same time, the inference speed is optimized while maintaining a high accuracy.

[0040] In addition, the model was further optimized and adjusted to improve its recognition performance and stability in specific scenarios. For example, in the metal surface scene, by adjusting the model parameters and enhancing data preprocessing, etc., the recognition accuracy and inference speed of the model were further improved.

[0041] In summary, the experimental results of this embodiment show that the present invention has excellent text recognition performance and inference speed in different scenarios, and has good generalization ability and adaptability. Through appropriate adjustment and optimization, the recognition performance and stability of the model in specific scenarios can be further improved.

Claims

1. A high-speed character recognition method based on the PaddleOCRModel, characterized in that, It includes the following steps: S1: Data loading module: Through the data loading module, the character pictures to be recognized are loaded from the disk into the code running environment in the form of image data. The character pictures include characters with various different glyphs, font sizes, and printing methods; S2: Data preprocessing module: Preprocess the loaded image data to meet the input requirements of the PaddleOCR model; S3: Model loading module: During the inference process, load the pre-trained PaddleOCR model into the device memory; S4: Model inference and post-processing module: Input the preprocessed image data into the loaded PaddleOCR model for inference to obtain the inference result, and post-process the inference result.

2. A high-speed character recognition method based on the PaddleOCRModel according to claim 1, characterized in that, According to step S2: Preprocess the loaded image data, and its preprocessing method includes: S21: Resize the image. Use the bilinear interpolation method to expand the image to the target size to ensure that the size of the processed image is consistent with the input size required by the model; S22: Grayscale processing. Use the weighted average method to convert the color image into a grayscale image; S23: Normalization and standardization processing. Limit the preprocessed data within the range of [0, 1] to eliminate the adverse effects caused by singular sample data.

3. The high-speed character recognition method based on the PaddleOCRModel according to claim 2, wherein According to step S22: Through the calculation formula of the grayscale processing: (gray(i, j) = 0.229×R(i, j) + 0.587×G(i, j) + 0.114×B(i, j)); Among them, R(i, j), G(i, j), and B(i, j) respectively represent the values of the red, green, and blue color channels at the position (i, j) in the image, and convert the three-channel color picture into a single-channel grayscale image.

4. A high-speed character recognition method based on the PaddleOCRModel according to claim 3, characterized in that, According to step S23: For the normalization and standardization processing, limit the target image data of the preprocessing within a certain range [0, 1] to eliminate the adverse effects caused by singular sample data. Its calculation formula is: normalized_pixel_value = (pixel_value - mean) / std); where mean is the average value obtained by dividing the sum of all pixel values of the image by the number of all pixel values, and std is the standard deviation of the image.

5. A high-speed character recognition method based on the PaddleOCRModel according to claim 4, characterized in that, According to step S4: The post-processing steps include: S41: Extract the confidence of each character from the inference result. The shape of the inference result is (batch, 40, 6625), where batch is the number of images inferred simultaneously, 40 is the maximum number of detected characters per line, and 6625 is the total number of Chinese, English, and various characters; S42: Traverse each unit (the confidence of 6625 characters) in the memory, find the character corresponding to the maximum confidence value, and splice to obtain the recognition result of each line.

6. The high-speed character recognition method based on the PaddleOCRModel according to claim 5, characterized in that, The bilinear interpolation method calculates the target pixel value by performing interpolation once in the sub-horizontal and vertical directions respectively to ensure the smoothness and continuity of the image when it is enlarged and reduced.

7. A high-speed character recognition method based on the PaddleOCRModel according to any one of claims 1-6, characterized in that, This method is applicable to the character recognition of various glyphs, font sizes, and printing methods, including but not limited to printed characters, handwritten characters, etc.

Citation Information

Patent Citations

  • Character recognition method and device, equipment and storage medium

    CN115690797A

  • Engineering drawing character detection and recognition method and system based on improved YOLOv5s

    CN116597466A

  • PadleDesection-based steel plate jet printing character image recognition system and recognition method

    CN117690140A