Intelligent Document Processing Method and Recognition System Based on OCR and AI Technologies

Through an intelligent document processing method based on OCR and AI technology, strong light scanning is used to remove handwriting imprints on paper, solving the problem that handwriting imprints affect character recognition and improving the accuracy of recognition.

CN119107658BActive Publication Date: 2025-05-30SHANGHAI ZHIHUI QIKE INFORMATION TECHNOLOGY CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411087052.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2025-05-30
Estimated Expiration
2044-08-08

AI Technical Summary

Technical Problem

During the text character recognition process, the handwriting imprint of the previous piece of paper on the paper will affect the scanned image, resulting in inaccurate character recognition.

Method used

Through an intelligent document processing method based on OCR and AI technology, the image to be recognized is obtained and scanned. During the character recognition process, the scanning method is judged based on the similarity between the characters and the preset characters. If the similarity is low, a strong light scan is performed to remove the handwriting imprint.

Benefits of technology

It realizes the removal of handwriting imprints through strong light scanning, improving the accuracy and accuracy of character recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119107658B_ABST
    Figure CN119107658B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of image recognition, particularly character recognition technology, and specifically relates to an intelligent document processing method and recognition system based on OCR and AI technologies, including: obtaining an image to be recognized; scanning the image to be recognized to identify the characters in the image to be recognized; wherein, during the process of recognizing the characters in the image to be recognized, the scanning method is judged according to the similarity degree between the characters and the preset characters; through the recognition of the characters, the key information required by the enterprise can be automatically recognized and extracted, the OCR recognition result is optimized through a deep learning algorithm to improve the accuracy and integrity of information extraction, and the handwriting is scanned by means of strong light scanning to facilitate the recognition of the handwriting. After the handwriting is recognized, the handwriting can be removed from the image to be recognized so as to accurately recognize the characters.
Need to check novelty before this filing date? Find Prior Art

Claims

1. An intelligent document processing method based on OCR and AI technology, characterized in that: include: Obtain an image to be recognized; Scanning the image to be recognized and recognizing characters in the image to be recognized; in In the process of identifying characters in the image to be identified, a scanning mode is determined according to the similarity between the characters and the preset characters; The method of scanning the image to be recognized and recognizing characters in the image to be recognized includes: Scanning the image to be recognized in a normal scanning manner, recognizing the characters therein, arranging the recognized characters from small to large according to the number of strokes, and comparing the characters in sequence with the characters in a preset character library; Determine the number of characters to be recognized, determine the number of samples according to the number of characters, and the number of samples is less than the number of characters; The method for determining the scanning mode according to the similarity between the character and the preset character in the process of recognizing the character in the image to be recognized includes: After normal scanning, the similarity of each character in the sample is obtained based on the comparison results of each character with the characters in the preset character library, and the similarities are sorted from small to large; The minimum similarity is selected, and whether strong light scanning is required is determined based on the minimum similarity.

2. The intelligent document processing method based on OCR and AI technology as claimed in claim 1, characterized in that: The method for determining whether strong light scanning is required based on the minimum similarity includes: If the minimum similarity is greater than the preset standard similarity, it is determined that strong light scanning is not necessary; At this time, the characters are rotated and scaled according to the similarity of each character, and the number of times the image to be recognized needs to be scanned is determined according to the minimum similarity; Get the scan result after scanning the determined number of times and rotating and scaling the characters.

3. The intelligent document processing method based on OCR and AI technology as claimed in claim 2, characterized in that: The method for determining whether strong light scanning is required based on the minimum similarity also includes: If the minimum similarity is less than the preset standard similarity, it is determined that strong light scanning is required; The characters obtained by the strong light scanning are compared with the characters obtained by the normal scanning to obtain trace samples, and the trace samples are compared.

4. The intelligent document processing method based on OCR and AI technology as claimed in claim 3, characterized in that: The method for comparing trace samples comprises: Sort the characters in the trace sample by the number of strokes from small to large, compare each character with the characters in the preset character library in order, obtain the similarity of each character in the trace sample, and determine the number of samples corresponding to the trace sample; The traces are judged based on their similarity.

5. The intelligent document processing method based on OCR and AI technology as claimed in claim 4, characterized in that: The method for judging traces according to similarity includes: Arrange the similarities of the characters in the trace sample from small to large, and determine the number of strong light scans required based on the smallest similarity; Obtain trace results after performing a corresponding number of strong light scans; Process the current image to be identified according to the trace results.

6. The intelligent document processing method based on OCR and AI technology as claimed in claim 5, characterized in that: The method for processing the current image to be identified according to the trace result includes: Determine the corresponding image to be identified to which the trace in the current image to be identified belongs according to the trace result, so as to complete the trace, and after the trace is completed, remove the trace from the current image to be identified, and obtain the scanning result of the current image to be identified; and According to the corresponding image to be identified to which the trace in the current image to be identified belongs, the sorting position of the current image to be identified is determined.

7. The intelligent document processing method based on OCR and AI technology as claimed in claim 6, characterized in that: The intelligent document processing method based on OCR and AI technology also includes: After obtaining the scanning result, the scanning result is input into a text image recognition model containing a mixed convolution kernel to obtain a convolution feature map corresponding to the scanning result.

8. The intelligent document processing method based on OCR and AI technology as claimed in claim 7, characterized in that: The convolution feature map is input into the recurrent neural network of the text image recognition model for feature extraction to obtain sequence features; Input the sequence features into the fully connected layer of the text image recognition model to obtain the character probability distribution result; The preset loss function is used to calculate the error loss of the character probability distribution result to obtain the text recognition result of the scan result.

9. A recognition system using the intelligent document processing method based on OCR and AI technology as claimed in claim 1, characterized in that: include: An acquisition module, configured to acquire an image to be recognized; A scanning module, which is configured to scan the image to be recognized and recognize characters in the image to be recognized; The judgment module is configured to judge the scanning mode according to the similarity between the characters and the preset characters in the process of recognizing the characters in the image to be recognized.

Citation Information

Patent Citations

  • Online exercise method and device with character recognition optimization function

    CN113283304A