Handwritten Chinese text recognition method
By building a segmentation module based on ResNet-18 and FCN and semantic error correction of the BERT model, the problems of high computational complexity and high annotation cost in handwritten Chinese character recognition methods were solved, and efficient and accurate text recognition effects were achieved.
Patent Information
- Application Number
- CN202511351788.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-09-22
Smart Images

Figure CN120853192A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image data processing technology, and in particular to a method for recognizing handwritten Chinese text. Background Technology
[0002] Handwritten Chinese character recognition involves processing text images to extract the characters contained within them. Initially, handwritten Chinese character recognition focused on recognizing individual characters, categorized into traditional methods and deep learning-based methods. Traditional single-character recognition methods primarily involve three steps: preprocessing, feature extraction, and character classification. Deep learning-based single-character recognition mainly utilizes convolutional neural networks and has achieved high accuracy rates. Subsequently, handwritten Chinese character recognition based on text lines began to develop. The number of characters in text line images is variable, and issues such as slanted handwriting and excessively large or small spacing between characters increase the difficulty of text line-based handwritten Chinese character recognition. Currently, there are two main types of methods for text line-based handwritten Chinese character recognition.
[0003] One type is the character segmentation-based recognition method, which first divides the text line image into several small blocks, then performs recognition processing on each small block, and finally connects the recognized characters. First, the text line image is over-segmented into a series of original segments using a connection component-based method, and then the original segments are recognized to obtain a candidate character set. The path is evaluated from the Bayesian decision view by combining multiple contexts (including character classification scores, geometric and linguistic contexts) and the recognition result is output. This method improves accuracy and efficiency through an improved beam search algorithm, and the character set accuracy reaches 90.75% on the CASIA-HWDB dataset. However, this method (1) relies on sliding window or stroke merging algorithms (such as HMM, stroke bounding box merging), which has high computational complexity and poor robustness to connected characters, (2) requires character-by-character annotation of training data, which has high annotation cost, low annotation efficiency, and is easily affected by segmentation errors in recognition accuracy.
[0004] Another category is segmentation-free text recognition methods, such as segmentation-free scene text recognition methods based on CTC (Center-Driven Transcription) (CTC recognition algorithms) and segmentation-free scene text recognition methods based on attention mechanisms (attention mechanism recognition algorithms). These methods do not require character segmentation annotations but are based on segmentation strategies. They first use a bidirectional long short-term memory network to extract features, then use a dynamic programming algorithm to find the optimal character segmentation points, and finally use a convolutional neural network to recognize each character. While these methods avoid annotation dependence, the CTC recognition algorithm ignores the spatial relationships between characters, making it difficult to simulate the cognitive logic of human character-by-character reading. The computational cost of the attention mechanism recognition algorithm increases quadratically with the sequence length, making it difficult to meet real-time requirements and limiting inference speed.
[0005] Definitions: ResNet-18: ResNet (Residual Network) is a deep neural network architecture that addresses the vanishing gradient and representation bottleneck problems during deep network training by introducing residual connections. ResNet-18 is a lightweight model in the ResNet series, containing 18 convolutional layers and 1 fully connected layer.
[0006] BERT Model: BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained natural language processing model. The core of the BERT model is the Transformer encoder, which can be unsupervised pre-trained on large-scale corpora and then fine-tuned on various NLP tasks. BERT is a bidirectional deep learning model that can consider all words in the context simultaneously, thus better understanding the meaning of sentences. Pre-training tasks for the BERT model include MLM (Masked Language Model) and NSP (Next Sentence Prediction). MLM, also known as masked language modeling, involves BERT randomly masking approximately 15% of the words and using a bidirectional Transformer architecture to predict these masked words. Its purpose is to force the encoder to utilize both left- and right-side information, learning deep bidirectional representations. NSP, also known as next sentence prediction, is used to determine whether sentence B is the true next sentence after sentence A. Its goal is to give [CLS] vectors sentence-level semantics and inter-sentence relationship modeling capabilities, directly serving downstream tasks such as question answering and reasoning.
[0007] FCN (Fully Convolutional Network) is a type of neural network in which all layers consist of convolutional layers (and associated nonlinear, normalization, upsampling, and other operations). Compared to other methods that use CNN convolutional neural networks, its advantage is that it has no requirements on the size of the input data and does not need to consider the size and pixel size of the input image. Summary of the Invention
[0008] The purpose of this invention is to provide a handwritten Chinese text recognition method that solves the above-mentioned problems and can greatly reduce annotation costs and improve recognition efficiency and accuracy.
[0009] To achieve the above objectives, the technical solution adopted by the present invention is as follows: a handwritten Chinese text recognition method, comprising the following steps; S1, obtain the handwritten Chinese dataset D1, where the samples are handwritten text images and the text in the samples is randomly labeled; S2, construct a segmentation module, including S21~S24; S21, Obtain an object detection network M1, including a pre-trained ResNet-18 and FCN, used to extract image features of handwritten text via ResNet-18 and predict the localization result of each character in the handwritten text via FCN, wherein the localization result includes a prediction box and the corresponding confidence score; S22, preset iteration number T, preset first threshold θ1, second threshold θ2, 0 < θ1 < θ2 < 1, preset set G; S23, train M1, where the t-th training includes Sa1~Sa3, 1≤t≤T; Sa1, input sample X into M1 to obtain the localization result of each character, and calculate the localization loss L. loc To minimize L loc Adjust the network parameters of M1; Sa2 generates the first training region R on sample X for the t-th time. 1,t Second region R 2,t Third region R 3,t Specifically, the range of prediction boxes with confidence levels less than θ1 is used to construct R. 1,t The range of prediction boxes with a confidence level greater than θ2 constitutes R. 3,t The remaining prediction box ranges constitute R. 2,t ; Sa3, if t=1, complete this training; if t>1, update R sequentially. 2,t Inner prediction box and confidence level, for R 2,t Inner prediction box B t If the updated confidence level is greater than θ2, then B t Updated to R 3,t ; S24, when t=T, the iteration ends, and R is pruned. 3,t The inner prediction box is used to obtain the single character image, which is stored in the set G, and the trained object detection network M1 is used as the segmentation model. S3, constructs a handwritten Chinese text recognition network, including a segmentation module, a recognition module, and a regularization module connected in sequence; The recognition module uses FCN to input single-character images within the set G, output predicted character categories, and arranges them as initial text according to the coordinates of the positioning boxes; The regularization module is a BERT model pre-trained on MLM and NSP tasks, used to correct errors in the initial text and generate the final text. S4. Use D1 to train the handwritten Chinese text recognition network. In each training session, calculate the localization loss of the segmentation module, the classification loss of the recognition module, and the semantic loss of the regularization module. The three constitute the total loss. Adjust the handwritten Chinese text recognition network to minimize the total loss to obtain the handwritten Chinese text recognition model. S5: Obtain the handwritten Chinese text to be recognized, and obtain the corresponding final text through the handwritten Chinese text recognition model.
[0010] Preferably, the handwritten Chinese dataset includes the offline handwritten Chinese dataset CASIA-HWDB, the online handwritten Chinese dataset CASIA-OLHWDB, the ICDAR2013 competition test set, the complex scene handwritten Chinese dataset SCUT-HCCDoc, and the scene Chinese dataset ReCTS.
[0011] Preferably, θ1 is 0.4~0.5 and θ2 is 0.8~0.9.
[0012] As a preferred option, in Sa3, for R 2,t Inner prediction box B t The update method is as follows: If B t The predicted bounding box B generated in the (t-1)th iteration t-1 There is overlap, calculate B. t confidence level p t and B t-1 confidence level p t-1 weighted average ,use Update p t ; Compare and p t-1 ,like >p t-1 Then B t Otherwise, use B. t-1 Update B t .
[0013] Compared with the prior art, the advantages of the present invention are as follows: (1) This invention constructs a novel segmentation module that requires two training sessions. The first session is the pre-training of the object detection network M1, which does not require extensive data training. It only requires a small amount of labeled data for pre-training and already possesses a certain ability to segment characters. The second session trains the object detection network M1 using dataset D1 according to the training method described in step S23 of this invention, continuously updating the predicted bounding boxes and confidence scores of the characters to generate the final localization result. This localization result can also generate localization results for unlabeled characters, enabling the generation of recognition results without the need for manual annotation of the ground truth bounding boxes. This reduces the cost of actual model deployment, avoids the need for extensive data annotation, reduces costs, and improves work efficiency.
[0014] (2) After generating the location results of the characters in the sample, the predicted character category is identified by the recognition module and then sent to the regular expression module in combination with the semantic context to correct the single characters with low recognition probability, so as to ensure that the semantic context is smooth and fluent, and the final text has high accuracy. Attached Figure Description
[0015] Figure 1 This is a flowchart of the present invention; Figure 2 Network structure diagram for handwritten Chinese text recognition; Figure 3 This is a diagram illustrating the recognition process of a handwritten text image in Example 2 via step S23. Figure 4 The final text obtained from image 1 using this invention; Figure 5 The final text obtained from image 2 using this invention; Figure 6 The final text obtained from image 3 using this invention; Figure 7 The final text obtained from image 4 using this invention. Detailed Implementation
[0016] The present invention will be further described below with reference to the embodiments and accompanying drawings.
[0017] Example 1: See Figure 1 and Figure 2 A method for recognizing handwritten Chinese text, comprising the following steps; S1, obtain the handwritten Chinese dataset D1, where the samples are handwritten text images and the text in the samples is randomly labeled; S2, construct a segmentation module, including S21~S24; S21, Obtain an object detection network M1, including a pre-trained ResNet-18 and FCN, used to extract image features of handwritten text via ResNet-18 and predict the localization result of each character in the handwritten text via FCN, wherein the localization result includes a prediction box and the corresponding confidence score; S22, preset iteration number T, preset first threshold θ1, second threshold θ2, 0 < θ1 < θ2 < 1, preset set G; S23, train M1, where the t-th training includes Sa1~Sa3, 1≤t≤T; Sa1, input sample X into M1 to obtain the localization result of each character, and calculate the localization loss L. loc To minimize L loc Adjust the network parameters of M1; Sa2 generates the first training region R on sample X for the t-th time. 1,t Second region R 2,t Third region R 3,t Specifically, the range of prediction boxes with confidence levels less than θ1 is used to construct R. 1,t The range of prediction boxes with a confidence level greater than θ2 constitutes R. 3,t The remaining prediction box ranges constitute R. 2,t ; Sa3, if t=1, complete this training; if t>1, update R sequentially. 2,t Inner prediction box and confidence level, for R 2,t Inner prediction box B t If the updated confidence level is greater than θ2, then B t Updated to R 3,t ; S24, when t=T, the iteration ends, and R is pruned. 3,t The inner prediction box is used to obtain the single character image, which is stored in the set G, and the trained object detection network M1 is used as the segmentation model. S3, constructs a handwritten Chinese text recognition network, including a segmentation module, a recognition module, and a regularization module connected in sequence; The recognition module uses FCN to input single-character images within the set G, output predicted character categories, and arranges them as initial text according to the coordinates of the positioning boxes; The regularization module is a BERT model pre-trained on MLM and NSP tasks, used to correct errors in the initial text and generate the final text. S4. Use D1 to train the handwritten Chinese text recognition network. In each training session, calculate the localization loss of the segmentation module, the classification loss of the recognition module, and the semantic loss of the regularization module. The three constitute the total loss. Adjust the handwritten Chinese text recognition network to minimize the total loss to obtain the handwritten Chinese text recognition model. S5. Obtain the handwritten Chinese text to be recognized, and obtain the corresponding final text through the handwritten Chinese text recognition model.
[0018] The handwritten Chinese dataset includes the offline handwritten Chinese dataset CASIA-HWDB, the online handwritten Chinese dataset CASIA-OLHWDB, the ICDAR2013 competition test set, the complex scene handwritten Chinese dataset SCUT-HCCDoc, and the scene Chinese dataset ReCTS.
[0019] The θ1 is 0.4 to 0.5, and the θ2 is 0.8 to 0.9.
[0020] In the Sa3, for R 2,t inside a prediction box B t , its update method is: If B t overlaps with the prediction box B t-1 generated in the (t-1)th time, calculate the confidence p t of B t and the confidence p t-1 of B t-1 weighted average , use to update p t ; Compare and p t-1 , if > p t-1 , then B t remains unchanged, otherwise use B t-1 to update B t .
[0021] Example 2: Refer to Figures 1 to 7 , regarding the segmentation module, we take a handwritten text picture as an example. Its original picture is as Figure 3 shown in (a). Since the segmentation module first needs to use the pre-trained object detection network M1, during its pre-training, the labeled data in the dataset includes "及", "时", "勉", "岁", "不", ".", so the object detection network M1 can accurately locate these characters, output their prediction boxes and confidences. The prediction boxes are as Figure 3 shown by the red boxes in (b), and the confidences of these prediction boxes are generally relatively high, higher than the second threshold θ2. However, in addition to being able to generate these prediction boxes more accurately, the object detection network M1 also generates many other prediction boxes as Figure 3 shown in (c). The confidence of each prediction box is different. We label the boxes with a confidence less than θ1 as yellow, the boxes with a confidence greater than θ2 as red, and the boxes with a confidence between θ1 and θ2 as blue. Then train M1 according to S23, specifically including: The first training: Sa1, input the Figure 3 image in (a) into M1 to obtain the localization result of each character, and calculate the localization loss L loc to minimize L loc adjust the network parameters of M1; Sa2, generate the first region R, the second region R, and the third region R for the first training on the sample X. R is the yellow boxed region in (c), R is the blue region in (c), and the third region R is the red boxed region in (c); 1,1 、Second region R 2,1 、Third region R 3,1 , R 1,1 as Figure 3 shown by the yellow boxed area in (c), R 2,1 as Figure 3 shown by the blue area in (c), the third region R 3,1 as Figure 3 shown by the red boxed area in (c); The second training: Sa1, input the Figure 3 image in (a) into M1 to obtain the localization result of each character, and calculate the localization loss L loc to minimize L loc adjust the network parameters of M1; Sa2, generate the first region R, the second region R, and the third region R for the second training on the sample X as described in (d); 1,2 、Second region R 2,2 、Third region R 3,2 , as Figure 3 described in (d); Sa3, update the predicted box and confidence within R in sequence, that is, the predicted box and confidence of each blue box in (d). We randomly select the predicted box corresponding to the character "month". It corresponds to a predicted box and confidence in (c), and also corresponds to a predicted box and confidence in (d). Observe that the two predicted boxes overlap, then calculate the weighted average of the two predicted boxes, use it to update p2; then compare it with p1. If it is greater than p1, retain the predicted box in (d), otherwise use the predicted box in (c) to update the predicted box of the character "month" in (d); then compare it with θ2. If it is greater than θ2, update the predicted box of the character "month" to R, that is, change it to the red box as shown in (e), otherwise do not process; 2,2 、也就是 Figure 3 shown by the blue boxed area in (c), R Figure 3 shown by the blue boxed area in (c), R Figure 3 shown by the blue boxed area in (d), R , use to update p2; then compare with p1. If > p1, retain the predicted box in (d), otherwise use the predicted box in (c) to update the predicted box of the character "month" in (d); then compare Figure 3 shown by the blue boxed area in (d), R Figure 3 shown by the blue boxed area in (c), R Figure 3 shown by the blue boxed area in (d), R with θ2. If > θ2, update the predicted box of the character "month" to R 3,t , that is, change it to the red box as shown in (e), otherwise do not process; Figure 3 shown by the red boxed area in (e), R The third to the Tth training is the same as the second. When t = T, the iteration ends, and clip R3,t The inner prediction box obtains a single - character image, which is stored in the set G. At this time, character classification and recognition have not been performed. The single - character image is a prediction box containing these characters, such as Figure 3 the red box shown in (e).
[0022] The single - character images are sequentially recognized by the recognition module. The recognition results are "及", "时", "当", "勉", "厉", "力", "岁", "月", "不", "待", "人", "。". Then, they are arranged into the initial text according to the coordinates of the positioning box, and the initial text is obtained as "及时当勉厉力岁月不待人。". Finally, it is sent to the regularization module for error correction. The regularization module has been pre - trained, can understand the context semantics, can correct "厉力" to "励", and add punctuation marks to generate the final text "及时当勉励,岁月不待人。".
[0023] Based on the above method, we obtain 4 handwritten text images, which are respectively labeled as Picture 1 - Picture 4. The final text obtained by using the method of the present invention is as Figures 4-7 shown. Figures 4-7 Among them, the first row is the original handwritten text picture, the second row is the schematic diagram of the prediction box generated by the segmentation module for the original picture, and the third row is the final text output by the regularization module.
[0024] Embodiment 3: To illustrate the effect of the present invention, a comparative experiment is carried out by using the present invention and two existing technologies, and Table 1 is obtained:[[]] Dataset: CASIA - HWDB and ICDAR2013 form a mixed dataset; Comparison methods: the present invention, CTC recognition algorithm, attention - mechanism recognition algorithm; Evaluation metrics: character error rate CER, inference speed, annotation cost; Table 1. Comparison table of experimental results of different methods on the mixed dataset Evaluation indicators This invention CTC identification algorithm Attention control recognition algorithm Character Error Rate (CER) 3.2% 5.8% 4.9% Inference speed (ms) 2.9 16.0 21.5 Labeling cost Text line level Text line level Text line level As can be seen from Table 1, the character error rate CER of the present invention is reduced by 27.6% relative to CTC, the inference speed is increased by 5 times, and the annotation cost is reduced by 90%.
[0025] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for recognizing handwritten Chinese text, characterized in that, Includes the following steps; S1, obtain the handwritten Chinese dataset D1, where the samples are handwritten text images, and the text in the samples is randomly labeled; S2, construct a segmentation module, including S21~S24; S21, Obtain an object detection network M1, including a pre-trained ResNet-18 and FCN, used to extract image features of handwritten text via ResNet-18 and predict the localization result of each character in the handwritten text via FCN, wherein the localization result includes a prediction box and the corresponding confidence score; S22, preset iteration number T, preset first threshold θ1, second threshold θ2, 0 < θ1 < θ2 < 1, preset set G; S23, train M1, where the t-th training includes Sa1~Sa3, 1≤t≤T; Sa1, input sample X into M1 to obtain the localization result of each character, and calculate the localization loss L. loc To minimize L loc Adjust the network parameters of M1; Sa2 generates the first training region R on sample X for the t-th time. 1,t Second region R 2,t Third region R 3,t Specifically, the range of prediction boxes with confidence levels less than θ1 is used to construct R. 1,t The range of prediction boxes with a confidence level greater than θ2 constitutes R. 3,t The remaining prediction box ranges constitute R. 2,t ; Sa3, if t=1, complete this training; if t>1, update R sequentially. 2,t Inner prediction box and confidence level, for R 2,t Inner prediction box B t If the updated confidence level is greater than θ2, then B t Updated to R 3,t ; S24, when t=T, the iteration ends, and R is pruned. 3,t The inner prediction box is used to obtain the single character image, which is stored in the set G, and the trained object detection network M1 is used as the segmentation model. S3, constructs a handwritten Chinese text recognition network, including a segmentation module, a recognition module, and a regularization module connected in sequence; The recognition module uses FCN to input single-character images within the set G, output predicted character categories, and arranges them as initial text according to the coordinates of the positioning boxes; The regularization module is a BERT model pre-trained on MLM and NSP tasks, used to correct errors in the initial text and generate the final text. S4. Use D1 to train the handwritten Chinese text recognition network. In each training session, calculate the localization loss of the segmentation module, the classification loss of the recognition module, and the semantic loss of the regularization module. The three constitute the total loss. Adjust the handwritten Chinese text recognition network to minimize the total loss to obtain the handwritten Chinese text recognition model. S5: Obtain the handwritten Chinese text to be recognized, and obtain the corresponding final text through the handwritten Chinese text recognition model.
2. The handwritten Chinese text recognition method according to claim 1, characterized in that, The handwritten Chinese datasets include the offline handwritten Chinese dataset CASIA-HWDB, the online handwritten Chinese dataset CASIA-OLHWDB, the ICDAR2013 competition test set, the complex scene handwritten Chinese dataset SCUT-HCCDoc, and the scene Chinese dataset ReCTS.
3. The handwritten Chinese text recognition method according to claim 1, characterized in that, θ1 is 0.4~0.5, and θ2 is 0.8~0.
9.
4. The handwritten Chinese text recognition method according to claim 1, characterized in that, In Sa3, for R 2,t Inner prediction box B t The update method is as follows: If B t The predicted bounding box B generated in the (t-1)th iteration t-1 There is overlap, calculate B. t confidence level p t and B t-1 confidence level p t-1 weighted average ,use Update p t ; Compare and p t-1 ,like >p t-1 Then B t Otherwise, use B. t-1 Update B t .
Citation Information
Patent Citations
Method and device for recognizing characters in image, medium and electronic equipment
CN112801085A
Document-level relation extraction method and system based on reasoning path and related equipment
CN119047561A
Computer vision-based surgical workflow recognition system using natural language processing techniques
US20230017202A1