Handwritten original record recognition method based on OCR (Optical Character Recognition) technology
By introducing error correction points of the wandering error correction port and stroke comparison database in OCR recognition, the problem of difficulty in identifying handwritten original records in unsatisfactory environments is solved, and higher recognition accuracy and error correction effects are achieved.
Patent Information
- Application Number
- CN202510262996.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, the image quality of handwritten original records is affected by factors such as the recording environment, paper quality and writing tools, which makes it difficult to identify, especially records completed in environments such as the field or workshop are more difficult to accurately identify.
By setting error correction points that include the roaming error correction port and the stroke comparison database, OCR recognition, semantic understanding and text error correction are performed, single characters are initially extracted using the roaming error correction port, and locally identify characters through the recognition of tentacles, merge them into identifyable characters, and similar features are compared to correct errors.
It improves the accuracy of the recognition of handwritten original records, effectively solves the difficulties caused by the fuzzy and unclear record content, and ensures that records completed in various environments can be accurately identified and corrected.
Smart Images

Figure CN120220167A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of OCR recognition, and particularly to a method for recognizing handwritten original records based on OCR technology. Background Art
[0002] The method for recognizing handwritten original records based on OCR (Optical Character Recognition) technology is a process of automatically recognizing and extracting measurement data and related information handwritten on paper documents by using computer technology.
[0003] Through devices such as scanners and digital cameras, paper documents, pictures, etc. containing text are converted into digital images that can be processed by a computer. For example, scanning paper documents in daily office work or taking pictures of pictures containing text with a mobile phone.
[0004] In the prior art, handwritten original records may be affected by various factors such as the recording environment, paper quality, writing tools, etc., resulting in poor image quality of the records. Situations such as paper wrinkles, ink penetration, and stain occlusion will make the recorded content unclear, bringing great difficulties to recognition. In actual measurement work, records may be completed in environments such as the wild and workshops, and these environments cannot guarantee the ideal recording conditions, thus affecting subsequent recognition. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for recognizing handwritten original records based on OCR technology to solve the deficiencies in the background art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for recognizing handwritten original records based on OCR technology, comprising the following steps:
[0007] Set the record recognition range, and collect an image of the handwritten original record information based on the record recognition range;
[0008] Preprocess the image of the handwritten original record information to obtain initial data information;
[0009] Set a character carrier, input the initial data information into the character carrier, and set error correction points based on the character carrier. The error correction points include a wandering error correction port and a stroke comparison database;
[0010] Perform OCR recognition based on the error correction points, and perform semantic understanding and text error correction to obtain the target text.
[0011] In a preferred embodiment, the step of setting the record recognition range and collecting an image of the handwritten original record information based on the record recognition range includes:
[0012] Set the record recognition range, and obtain an image of the handwritten original record information within the record recognition range;
[0013] The image of the handwritten original record information is composed of pixel points.
[0014] In a preferred embodiment, the step of preprocessing the image of the handwritten original record information to obtain initial data information includes:
[0015] Perform grayscale processing on the obtained image of the handwritten original record information to obtain a grayscale image;
[0016] Perform noise reduction processing on the grayscale image to obtain a smoothed image based on the noise reduction processing;
[0017] Perform skew correction on the smoothed image to obtain an initial data pattern based on the skew correction;
[0018] Extract single characters in the initial data pattern to obtain initial data information.
[0019] In a preferred embodiment, the steps of setting a character carrier, inputting the initial data information into the character carrier, and setting error correction points based on the character carrier, where the error correction points include a wandering error correction port and a stroke comparison database, include:
[0020] Set a character carrier, where the character carrier includes multiple byte units; input the initial data information into the character carrier, and input one single character into each byte unit;
[0021] Set a wandering error correction port based on the character carrier, and the wandering error correction port wanders and recognizes between multiple byte units;
[0022] The multiple single characters include text marks and symbol marks; generate a stroke comparison database based on multiple character models and identifier models by collecting multiple handwritten character samples as character models and collecting multiple handwritten symbol samples as identifier models;
[0023] Equip the wandering error correction port with a stroke comparison database;
[0024] Take the wandering error correction port and the stroke comparison database as error correction points.
[0025] In a preferred embodiment, the step of setting a wandering error correction port based on the character carrier, where the wandering error correction port wanders and recognizes between multiple byte units, includes:
[0026] The wandering error correction port wanders between multiple byte units, and the wandering error correction port performs preliminary feature extraction on the single character in each byte unit to generate preliminary character features;
[0027] Multiple wandering error correction ports are respectively provided with multiple recognition antennae, and two marking points are set based on a single byte unit, which are respectively located on both sides of a byte unit, and the area between the two marking points is set as the recognition range of the recognition antennae; the recognition range of the recognition antennae corresponds to the single byte unit; the local features of the monomer characters in the byte unit are recognized to generate detailed character features;
[0028] Based on the recognition tentacles, the detailed character features and the preliminary character features are merged into identifiable characters.
[0029] In a preferred embodiment, the step of performing OCR recognition based on the error correction points, performing semantic understanding and text error correction, and obtaining the target text includes:
[0030] Enter the initial data information into the character carrier and start OCR recognition based on the error correction point;
[0031] The identifiable characters collected by the recognition antenna are matched in the stroke comparison database. If there is a mismatch, the recognition antenna marks the byte unit and marks the individual characters in the unit as a state to be corrected;
[0032] The recognition antenna will feed back the pending error correction state to the wandering error correction port; the wandering error correction port will compare the single character in the pending error correction state with the stroke comparison database for similar features, obtain a character model or identifier model with similar features, and replace the recognized single character with a character model or identifier model with similar features; the wandering error correction port releases the pending error correction state of the single character and completes the recognition of the byte unit;
[0033] After completing the recognition and marking of a byte unit, the wandering error correction port continues to wander to the next byte unit until all byte units are traversed to obtain the target text.
[0034] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0035] The present invention sets an error correction point including a wandering error correction port and a stroke comparison database. During the recognition process, the wandering error correction port can not only perform preliminary feature extraction on a single character, but also perform local feature recognition on the character through recognition tentacles, and merge the detailed character features with the preliminary character features into an identifiable character, so as to make the feature description of the character more complete. At the same time, during the OCR recognition process, if a character mismatch occurs, the wandering error correction port can be used to compare again with the stroke comparison database, and replaced with a character model or identifier model with similar features, thereby greatly improving the accuracy of recognition and effectively solving the great difficulty caused by unclear recorded content. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments described in the present invention. For those of ordinary skill in the art, other accompanying drawings can also be obtained based on these drawings.
[0037] Figure 1 It is a flowchart of the method of the present invention. Specific embodiments
[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0039] Embodiment 1, please refer to Figure 1 As shown, a method for recognizing handwritten original records based on OCR technology in this embodiment includes the following steps:
[0040] S1. Set the record recognition range, and collect the image of the handwritten original record information based on the record recognition range;
[0041] S2. Preprocess the image of the handwritten original record information to obtain initial data information;
[0042] S3. Set a character carrier, input the initial data information into the character carrier, and set error correction points based on the character carrier. The error correction points include a wandering error correction port and a stroke comparison database;
[0043] S4. Perform OCR recognition based on the error correction points, and perform semantic understanding and text error correction to obtain the target text.
[0044] As described in the above steps S1 - S4, handwritten original records may be affected by various factors such as the recording environment, paper quality, writing tools, etc., resulting in poor image quality of the records. Situations such as paper wrinkles, ink penetration, and stain occlusion will make the recorded content unclear, causing great difficulties for recognition. In actual metering work, the records may be completed in environments such as the wild or workshops, and these environments cannot guarantee the ideal recording conditions, thus affecting subsequent recognition. In this application, by setting error - correction points including a wandering error - correction port and a stroke comparison database, during the recognition process, the wandering error - correction port can not only perform preliminary feature extraction on single characters, but also perform local feature recognition on characters through recognition antennas, combining the detailed character features with the preliminary character features into identifiable characters, making the feature description of characters more complete. At the same time, during the OCR recognition process, if there is a situation where the character matching is incorrect, it can be compared again through the wandering error - correction port and the stroke comparison database, and replaced with a character model or identifier model with similar features, thus greatly improving the recognition accuracy and effectively solving the great difficulties brought by unclear recorded content for recognition.
[0045] In one embodiment, the step S1 of setting a record recognition range and collecting an image of handwritten original record information based on the record recognition range includes:
[0046] S11. Set a record recognition range and obtain an image of handwritten original record information within the record recognition range;
[0047] S12. The image of the handwritten original record information is composed of pixel points.
[0048] As described in the above steps S11 - S12, first, by setting a specific record recognition range, the boundary of the information collection area is clarified, ensuring that the collected image only contains content related to the handwritten original record and avoiding interference from irrelevant information. Subsequently, an image of handwritten original record information is obtained within the set record recognition range, and these images are composed of pixel points. Each pixel point carries information such as color and brightness, laying a foundation for subsequent analysis, processing, and recognition of the image of handwritten original record information. The combination and characteristics of pixel points will determine how to accurately extract and interpret various data and content in the handwritten original record.
[0049] In one embodiment, the step S2 of pre - processing the image of the handwritten original record information to obtain initial data information includes:
[0050] S21. Perform grayscale processing on the obtained image of the handwritten original record information to obtain a grayscale image;
[0051] S22. Perform noise reduction processing on the grayscale image to obtain a smooth image based on the noise reduction processing;
[0052] S23. Perform skew correction on the smoothed image to obtain an initial data graph based on the skew correction;
[0053] S24. Extract single characters in the initial data graph to obtain initial data information. As described in the above steps S21 - S24, first, perform grayscale processing on the image of the obtained handwritten original record information to convert the color image into a grayscale image, so as to simplify the subsequent processing flow and reduce the complexity brought by color information. Then, perform noise reduction processing on the grayscale image to eliminate possible noise interference in the image, thereby obtaining a smoothed image, which helps to improve the image quality and make the subsequent processing more accurate. After that, perform skew correction on the smoothed image to ensure that the image is at the correct angle and avoid information reading deviation caused by skew. Finally, extract single characters in the initial data graph. After a series of image preprocessing steps, initial data information is obtained, providing a relatively clear, accurate, and easy - to - operate data basis for subsequent further data processing and analysis, facilitating the subsequent interpretation, analysis, and use of this initial data information.
[0054] In one embodiment, the step S3 of setting a character carrier, inputting the initial data information into the character carrier, and setting error - correction points based on the character carrier, where the error - correction points include a wandering error - correction port and a stroke comparison database, includes:
[0055] S31. Set a character carrier, where the character carrier includes multiple byte units; input the initial data information into the character carrier, and input one single character into each byte unit;
[0056] S32. Set a wandering error - correction port based on the character carrier, and the wandering error - correction port wanders and identifies between multiple byte units;
[0057] S33. The multiple single characters include text marks and symbol marks; generate a stroke comparison database based on multiple character models and identifier models by collecting multiple handwritten character samples as character models and collecting multiple handwritten symbol samples as identifier models;
[0058] S34. Equip the wandering error - correction port with the stroke comparison database;
[0059] S35. Use the wandering error - correction port and the stroke comparison database as error - correction points.
[0060] As described in the above steps S31 - S35, first, a character carrier is set up. The character carrier contains multiple byte units, and its purpose is to input initial data information into it, and each byte unit can carry a single character, providing an ordered storage structure for subsequent data processing. Then, based on the character carrier, a roaming error correction port is set up. This roaming error correction port can roam and identify among multiple byte units, providing a flexible inspection mechanism for subsequent error correction operations. The multiple single characters include text marks and symbol marks. By collecting multiple handwritten character samples and handwritten symbol samples as the character model and identifier model respectively, a stroke comparison database is generated on this basis. This database will provide an important reference basis for error correction. Then, the stroke comparison database is equipped for the roaming error correction port. Finally, the roaming error correction port and the stroke comparison database are used as error correction points. Through this series of operations, the initial data information input into the character carrier can be carefully error - corrected. By the collaborative work of the stroke comparison database and the roaming error correction port, the strokes of single characters are compared and inspected, and possible errors can be timely detected and corrected to ensure the accuracy and reliability of the data.
[0061] In one embodiment, the step S32 of setting up a roaming error correction port based on the character carrier and the roaming error correction port roaming and identifying among multiple byte units includes:
[0062] S321. The roaming error correction port roams among multiple byte units, and the roaming error correction port performs preliminary feature extraction on the single character in each byte unit to generate preliminary character features;
[0063] S322. Each of the multiple roaming error correction ports is provided with multiple recognition tentacles, and two marking points are set based on a single byte unit, located on both sides of a byte unit respectively. The area between the two marking points is set as the recognition range of the recognition tentacles; the recognition range of the recognition tentacles corresponds to a single byte unit; local features of the single character in the byte unit are recognized to generate detailed character features;
[0064] S323. Based on the recognition tentacles, the detailed character features and the preliminary character features are merged into an identifiable character.
[0065] As described in the above steps S321-S323, first, the wandering error correction port will wander between multiple byte units, and perform preliminary feature extraction on the individual characters in each byte unit to generate preliminary character features, which provides an overall feature information basis for subsequent error correction and recognition. Next, multiple recognition antennae are set for the multiple wandering error correction ports, and two marking points are set on both sides of a single byte unit. The area between them is used as the recognition range of the recognition antennae, and the recognition range corresponds to the single byte unit. On this basis, the local features of the individual characters in the byte unit are recognized, and then the detailed character features are generated. Through this local feature recognition, the feature information of the individual characters can be grasped more meticulously. Finally, based on the recognition tentacles, the detailed character features and the preliminary character features are merged into identifiable characters. Through this feature integration, not only the overall characteristics of the individual characters are taken into account, but also the local detailed features are taken into account, which provides richer and more comprehensive character feature information for more accurate character identification and subsequent error correction processing, and helps to improve the accuracy and reliability of error correction; the extracted features include but are not limited to the number of strokes of the character, the starting position of the strokes, the ending position of the strokes, the curvature of the strokes and other basic information; the collected local feature information includes the continuity of the local strokes of the character, the density of the local strokes, the smoothness of the character edges, etc.; the detailed character features supplement and refine the preliminary character features, and together with the preliminary character features, they constitute a complete feature description of the individual characters.
[0066] In one embodiment, the step S4 of performing OCR recognition based on the error correction points, and performing semantic understanding and text error correction to obtain the target text includes:
[0067] S41, inputting the initial data information into the character carrier, and starting OCR recognition based on the error correction point;
[0068] S42, matching the identifiable characters collected by the recognition antenna in the stroke comparison database, if there is a mismatch, the recognition antenna marks the byte unit, and marks the individual characters in the unit as a state to be corrected;
[0069] S43, the recognition antenna will feed back the pending error correction state to the wandering error correction port; the wandering error correction port will compare the single character in the pending error correction state with the stroke comparison database for similar features, obtain a character model or identifier model with similar features, and replace the recognized single character with a character model or identifier model with similar features; the wandering error correction port releases the pending error correction state of the single character and completes the recognition of the byte unit;
[0070] S44, after completing the recognition and marking of a byte unit, the wandering error correction port continues to wander to the next byte unit until the traversal of all byte units is completed to obtain the target text.
[0071] As described in the above steps S41-S44, the initial data information is first entered into the character carrier, and then the OCR recognition operation is started based on the error correction point, laying the foundation for subsequent character processing and error correction. Next, the recognition antenna will match the collected identifiable characters in the stroke comparison database. If there is a mismatch, the recognition antenna will mark the corresponding byte unit, and at the same time mark the monomer character in the unit as a state to be corrected, so as to find characters that may have errors. After that, the recognition antenna will feed back the state to be corrected to the wandering error correction port, and the wandering error correction port will compare the monomer character in the state to be corrected with the stroke comparison database again for similar features to find a character model or identifier model with similar features, replace the recognized monomer character with a character model or identifier model with similar features, and release the state to be corrected of the monomer character, thereby completing the recognition of the byte unit. In this way, the correction of the erroneous character can be achieved. Finally, after completing the recognition and marking of a byte unit, the wandering error correction port will continue to wander to the next byte unit and repeat the above operations until the traversal of all byte units is completed, and finally the target text after recognition and error correction is obtained. Through this series of orderly operations, the consistency and accuracy of the entire recognition and error correction process are ensured, and the quality and reliability of the conversion from initial data information to target text are improved.
[0072] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for recognizing handwritten original records based on OCR technology, characterized in that: The following steps are involved: Setting a record recognition range, and collecting an image of the handwritten original record information based on the record recognition range; Preprocessing the image of the original handwritten record information to obtain initial data information; Setting a character carrier, entering initial data information into the character carrier, and setting error correction points based on the character carrier, wherein the error correction points include a wandering error correction port and a stroke comparison database; OCR recognition is performed based on the error correction points, and semantic understanding and text error correction are performed to obtain the target text.
2. The method for recognizing handwritten original records based on OCR technology according to claim 1, characterized in that: The step of setting a record recognition range and collecting an image of the handwritten original record information based on the record recognition range includes: Setting a record recognition range, and acquiring an image of the handwritten original record information within the record recognition range; The image of the handwritten original record information is composed of pixels.
3. The method for recognizing handwritten original records based on OCR technology according to claim 1, characterized in that: The step of preprocessing the image of the handwritten original record information to obtain initial data information includes: Grayscale processing is performed on the image of the obtained original handwritten record information to obtain a grayscale image; The grayscale image is subjected to denoising, and a smooth image is obtained based on the denoising process; Performing tilt correction on the smoothed image, and obtaining an initial data graph based on the tilt correction; The monomer characters in the initial data graph are extracted to obtain the initial data information.
4. The method for recognizing handwritten original records based on OCR technology according to claim 1, characterized in that: The step of setting a character carrier, inputting initial data information into the character carrier, and setting an error correction point based on the character carrier, wherein the error correction point includes a wandering error correction port and a stroke comparison database, comprises: Setting a character carrier, wherein the character carrier includes a plurality of byte units; entering initial data information into the character carrier, wherein a single character is entered into each byte unit; A wandering error correction port is set based on the character carrier, and the wandering error correction port wanders and identifies between multiple byte units; The plurality of monomer characters include text marks and symbol marks; by collecting a plurality of handwritten character samples as character models, by collecting a plurality of handwritten symbol samples as identifier models, a stroke comparison database is generated based on the plurality of character models and identifier models; Equip the wandering error correction port with a stroke comparison database; The wandering error correction port and the stroke comparison database are used as error correction points.
5. The method for recognizing handwritten original records based on OCR technology according to claim 4, characterized in that: The step of setting a wandering error correction port based on a character carrier, and the wandering error correction port wanderingly identifying between multiple byte units, comprises: The wandering error correction port wanders between multiple byte units, and the wandering error correction port performs preliminary feature extraction on a single character in each byte unit to generate preliminary character features; Multiple wandering error correction ports are respectively provided with multiple recognition antennae, and two marking points are set based on a single byte unit, which are respectively located on both sides of a byte unit, and the area between the two marking points is set as the recognition range of the recognition antennae; the recognition range of the recognition antennae corresponds to the single byte unit; the local features of the monomer characters in the byte unit are recognized to generate detailed character features; Based on the recognition tentacles, the detailed character features and the preliminary character features are merged into identifiable characters.
6. The method for recognizing handwritten original records based on OCR technology according to claim 1, characterized in that: The step of performing OCR recognition based on the error correction points, performing semantic understanding and text error correction, and obtaining the target text includes: Enter the initial data information into the character carrier and start OCR recognition based on the error correction point; The identifiable characters collected by the recognition antenna are matched in the stroke comparison database. If there is a mismatch, the recognition antenna marks the byte unit and marks the individual characters in the unit as a state to be corrected; The recognition antenna will feed back the pending error correction state to the wandering error correction port; the wandering error correction port will compare the single character in the pending error correction state with the stroke comparison database for similar features, obtain a character model or identifier model with similar features, and replace the recognized single character with a character model or identifier model with similar features; the wandering error correction port releases the pending error correction state of the single character and completes the recognition of the byte unit; After completing the recognition and marking of a byte unit, the wandering error correction port continues to wander to the next byte unit until all byte units are traversed to obtain the target text.