A method for handwritten Chinese character recognition and correction
Through the EfficientDet dual network and feature pyramid model combined with hypergraph neural network, the accuracy and correction of scribbled fonts in handwritten Chinese characters are solved, and the rapid and accurate recognition and correction of scribbled Chinese characters are achieved.
Patent Information
- Application Number
- CN202411878005.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-12-19
AI Technical Summary
The prior art has problems such as low accuracy and insufficient model generalization ability in the recognition of handwritten Chinese characters, especially poor recognition and correction effects for scribbled fonts.
The EfficientDet dual network and feature pyramid model are combined with the hypergraph neural network, and the Chinese character border is obtained through Gaussian filtering processing and template matching, a training data set is established, and the deep convolution generation adversarial network is used to generate scribbled samples, and the scribbled font correction is performed in combination with the hypergraph neural network.
The recognition speed and correction accuracy of scribbled Chinese characters have been improved, and the rapid and accurate recognition and correction of scribbled Chinese characters have been achieved.
Smart Images

Figure CN119323794B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of handwritten Chinese characters, and in particular to a method for recognizing and correcting handwritten Chinese characters. Background Art
[0002] As an important information carrier, Chinese characters play an important role in the fields of large-scale handwritten material input, document recognition, and teaching assistance. Handwritten Chinese character recognition technology is an important research direction in the field of computer vision and pattern recognition. It involves converting handwritten Chinese character images into editable and recognizable text information. With the rapid development of deep learning technology, the field of handwritten Chinese character recognition has ushered in new opportunities.
[0003] Chinese patent application CN109800763A discloses a method for handwritten Chinese character recognition based on deep learning. The method improves the performance of handwritten Chinese character recognition by using P2DMN normalization, NCFE feature extraction, ADBN coarse classification and MQDF fine classification, combined with deep learning technology. However, the above method is not very accurate in recognizing untrained handwritten Chinese characters, the generalization ability of the model is not strong, and a large data set is required for training to prevent the model from overfitting. It is impossible to provide a large amount of training data, and it also fails to recognize sloppy handwritten Chinese characters.
[0004] In the prior art, the recognition of handwritten Chinese characters relies on a single neural network model, which easily leads to low accuracy of handwritten Chinese character recognition and cannot provide the large data set required for training. In addition, when handwritten Chinese characters encounter sloppy fonts, they need to be quickly recognized and accurately corrected through a neural network. Summary of the invention
[0005] The purpose of the present invention is to provide a handwritten Chinese character recognition and correction method to solve the above-mentioned problems existing in the prior art.
[0006] The specific application is as follows:
[0007] A method for recognizing and correcting handwritten Chinese characters comprises the following steps:
[0008] S1: Obtain handwritten Chinese character images and perform noise reduction on them;
[0009] S2: performing image segmentation on the handwritten Chinese character image after noise reduction processing to extract the image to be recognized with Chinese characters;
[0010] S3: Obtain the bounding box of a single Chinese character in the image to be recognized based on template matching;
[0011] S4: establishing a training data set and a first recognition model for target handwritten Chinese characters, and training the first recognition model for target handwritten Chinese characters based on the training data set;
[0012] S5: Detect the image to be recognized based on the trained first recognition model for target handwritten Chinese characters, and obtain the first recognition result of the handwritten Chinese characters in the image to be recognized. The first recognition result includes scribbled fonts and non-scribbled fonts;
[0013] S6: Correct the scribbled fonts in the first recognition result based on the hypergraph neural network;
[0014] In step S4, the specific implementation process of establishing the training data set and the first recognition model for target handwritten Chinese characters is as follows:
[0015] S41: Obtain scribbled and non-scribbled Chinese character image data through handwritten Chinese characters. Part of it is used as the test set, and part of it is used as the training set. Two EfficientDet network models are trained through the training set. The feature acquisition network of the EfficientDet network model is ResNeXt. The two network models obtained after training are respectively denoted as ResNeXt-A and ResNeXt-B. The Chinese character image data is defined by manual labeling, including two labels: scribbled and non-scribbled;
[0016] S42: Divide the EfficientDet network model in S41 into a feature extraction network ResNeXt, a feature pyramid, and a scribbled classification and regression network. Use the feature extraction network ResNeXt to extract the features of scribbled and non-scribbled images. Connect the feature pyramid to the feature extraction network ResNeXt, perform convolution on each layer of the feature pyramid, and perform classification and regression through the scribbled classification and regression network;
[0017] S43: Use the test set to perform network iterative testing on the feature extraction network ResNeXt, and output the recognition result. When the recognition accuracy reaches the set threshold, stop the training of this model. The recognition result is used to judge whether the image is scribbled.
[0018] Further, in step S1, the noise reduction process is Gaussian filtering, specifically:
[0019] ,
[0020] where g(x, y) is the handwritten Chinese character image, (x, y) is the handwritten Chinese character image after Gaussian filtering, u(i, j) is the weight, and D(x, y) is the size range of (N + 1) × (N + 1) centered on (x, y);
[0021] Among them, the weight is specifically:
[0022]
[0023] In the formula, (i, j) is the spatial channel proximity weight value, (i, j) is the spatial channel similarity weight value, is the input random coefficient.
[0024] Furthermore, in step S2, the image segmentation of the handwritten Chinese character image after noise reduction is specifically as follows: The handwritten Chinese character image after noise reduction is segmented based on finding the minimum bounding rectangle of the target, and specifically includes the following steps:
[0025] S21: Perform binarization processing on the handwritten Chinese character image after noise reduction to obtain a binarized image;
[0026] S22: Based on the contour detection algorithm, search for and draw the contours of the binarized image, then fit the drawn target contours to find the minimum bounding square of the contours;
[0027] S23: Based on the center point of the minimum bounding square being at 2 / 4 - 3 / 4 of the handwritten Chinese character image, search for the Chinese character at the center and intercept the segmented image with the central Chinese character;
[0028] S24: Correct the segmented image to obtain the image to be recognized.
[0029] Furthermore, in step S24, the correction of the segmented image to obtain the image to be recognized is specifically as follows:
[0030] S251: Make the four corner points of the minimum bounding square correspond to the four corners of the image, and then based on these four pairs of corresponding corner points, obtain the perspective transformation matrix using the fitgeotrans function in MATLAB;
[0031] S252: Based on the perspective transformation matrix, call the imwarp function in MATLAB to transform the segmented image into the corrected view to obtain the image to be recognized;
[0032] S253: Output the image to be recognized.
[0033] Furthermore, S3 specifically includes the following steps:
[0034] S31: Construct a first image template and a second image template, where the first image template is the standard image of the Chinese character, and the second image template is the standard inner frame image of the Chinese character;
[0035] S32: Slide the first image template and the second image template on the image to be recognized respectively, and find the position where the maximum best matching value is obtained, that is, obtain the inner and outer frames of the Chinese character in the image to be recognized.
[0036] Further, the specific construction steps of the feature extraction network ResNeXt are as follows:
[0037] S411: Judge the images in the training set data according to the two labels of scribbled and non-scribbled. If the label of the image is a scribbled image, process it with ResNeXt-A. If the label of the image is a non-scribbled image, process it with ResNeXt-B. If it is judged that the currently input training set data image is scribbled, construct a deep convolutional generative adversarial network for scribbled images;
[0038] S412: Use the deep convolutional generative adversarial network to generate scribbled image training set samples for the scribbled images;
[0039] The processing flow for scribbled and non-scribbled features is as follows:
[0040] The feature acquisition network adopts the network framework of the feature extraction network ResNeXt, extracts the Chinese character features of scribbled and non-scribbled respectively, first performs superposition on the feature maps of the scribbled and non-scribbled convolutional layers in the second layer based on channels, and then performs dimensionality reduction processing through a 2×2 convolution. The same operation is performed on the third, fourth, fifth, and sixth convolutional layers;
[0041] The specific process of using the deep convolutional generative adversarial network to generate scribbled image training set samples for the scribbled images is as follows: According to the original scribbled image given in the current training, generate adversarial scribbled image samples. The specifications of the adversarial scribbled image samples are: diverse font size features, scribbled stroke features, and overall scribbled font features. Input 1 original given scribbled image, add a perturbation factor through the deep convolutional generative adversarial network to generate 200 scribbled sample images. The scribbled stroke features at least include the feature of non-fixed stroke length and the feature of deviation of the relative position of strokes. The overall scribbled font features at least include the feature of deviation of the overall font width-to-height ratio and the feature of deviation of the overall font center of gravity position.
[0042] Further, the specific process of connecting the feature pyramid to the feature extraction network ResNeXt is as follows:
[0043] Sample the feature map of the third convolutional layer, that is, the scribbled and non-scribbled fused feature map, and then superimpose it on the scribbled and non-scribbled fused feature maps of the second convolutional layer to obtain the first layer of the pyramid. Continue to perform this step on the fourth, fifth, and sixth convolutional layers. Superimpose the feature maps of every two adjacent layers on channels to obtain one layer of the pyramid. Finally, a total of four layers of feature pyramids are obtained. The specific processing process of the four layers of the pyramid is as follows:
[0044] The first layer of the pyramid is used to process the first image resolution, and the brightness of the image belongs to the first-level brightness;
[0045] The second layer of the pyramid is used to process the second image resolution, and the brightness of the image belongs to the second-level brightness;
[0046] The third layer of the pyramid is used to process the third image resolution, and the brightness of the image belongs to the third-level brightness;
[0047] The fourth layer of the pyramid is used to process the fourth image resolution, and the brightness of the image belongs to the fourth-level brightness;
[0048] The resolution levels of the said image are specifically divided as follows:
[0049] The first image resolution: w < 160, h < 160;
[0050] The second image resolution: 160 < w ≤ 320, 160 < h ≤ 320;
[0051] The third image resolution: 320 < w ≤ 1080, 320 < h ≤ 1080;
[0052] The fourth image resolution: 1080 < w, 1080 < h, where w represents the width of the image and h represents the height of the image;
[0053] The brightness of the said image adopts the method of calculating the average brightness of the image by using a histogram, and is specifically divided as follows:
[0054] The first average brightness: 0 ≤ L ≤ 60;
[0055] The second average brightness: 60 < L ≤ 100;
[0056] The third average brightness: 100 < L ≤ 160;
[0057] The fourth average brightness: 160 < L ≤ 255, where L represents the average brightness value of the image;
[0058] Two branch networks are added behind the feature map of each fusion layer of the said pyramid. One branch is used for classification and the other branch is used for regression. And each branch first performs 4 convolutions on the feature map to enhance the scribbled and non-scribbled image features respectively, and the convolution kernel size is 2×2 and the number is 128.
[0059] Furthermore, the specific implementation process of S6 is as follows:
[0060] The stroke scribbled feature and the overall font scribbled feature of the scribbled image to be processed are input into the first hypergraph neural network layer of the preset Chinese character correction model to obtain the first scribbled feature vector; wherein, the first scribbled feature vector represents the association relationship between the stroke scribbled feature and the overall font scribbled feature of the scribbled image to be processed; the preset Chinese character correction model is a model obtained based on the hypergraph neural network model;
[0061] Input the sub - features of the stroke scribble features and the sub - features of the overall font scribble features of the to - be - processed scribbled image into the second hyper - graph neural network layer of a preset Chinese character correction model to obtain a second scribble feature vector; wherein, the second feature vector characterizes the correlation between the sub - features of the stroke scribble features and the sub - features of the overall font scribble features of the to - be - processed scribbled image;
[0062] Connect the first scribble feature vector and the second scribble feature vector to obtain a to - be - processed feature vector;
[0063] Based on the preset Chinese character correction model, identify the to - be - processed feature vector to obtain the scribble type of the scribbled image, and the scribble type at least includes long - short stroke scribble, stroke - connection scribble, stroke - inclination scribble, pen - gesture scribble, whole - character position scribble, whole - character width - to - height ratio scribble, and whole - character center - of - gravity scribble;
[0064] For the scribble types in the preset handwritten Chinese character correction library, calculate the target similarity between the scribble type of the to - be - processed scribbled image and the scribble types in the preset handwritten Chinese character correction library through the cosine similarity algorithm, and set a preset similarity threshold. If the target similarity is greater than the preset similarity threshold, find the corresponding scribble type in the handwritten Chinese character correction library and then correct the Chinese characters in the scribbled image to generate standard Chinese characters. If the target similarity is less than or equal to the preset similarity threshold, find the scribble type with the nearest similarity in the handwritten Chinese character correction library to correct the Chinese characters in the scribbled image.
[0065] Furthermore, the loss function of the preset Chinese character correction model is:
[0066]
[0067] In the formula, is the loss function, k is the balance factor, is the correlation matrix aggregation parameter, is the target probability of the to - be - classified scribble type.
[0068] Compared with the prior art, the embodiments of the present invention achieve the following beneficial effects:
[0069] An embodiment of the present invention provides a method for obtaining a handwritten Chinese character image and performing noise reduction processing on it; performing image segmentation on the noise-reduced handwritten Chinese character image to extract a to-be-recognized image with Chinese characters; obtaining the bounding box of a single Chinese character in the to-be-recognized image based on template matching; establishing a training data set and a first recognition model for target handwritten Chinese characters, and training the first recognition model for target handwritten Chinese characters based on the training data set; detecting the to-be-recognized image based on the trained first recognition model for target handwritten Chinese characters to obtain a first recognition result of the handwritten Chinese characters in the to-be-recognized image, where the first recognition result includes scribbled fonts and non-scribbled fonts; correcting the scribbled fonts in the first recognition result based on a hypergraph neural network; the present invention uses an EfficientDet dual network to recognize images, fuses a feature pyramid model, processes the features of scribbled and non-scribbled handwritten Chinese characters, uses an adversarial network to generate a large number of adversarial scribbled sample data to improve the recognition effect of scribbled Chinese characters, and at the same time uses a hypergraph neural network to accurately identify the scribbled types of the recognized scribbled Chinese characters. Therefore, the recognition speed of scribbled Chinese characters is fast and the accuracy of correcting scribbled Chinese characters is high. Description of the Drawings
[0070] Figure 1 is a flowchart of a method for recognizing and correcting handwritten Chinese characters provided by an embodiment of the present invention;
[0071] Figure 2 is a schematic diagram of the process operation of the feature extraction network ResNeXt for a method for recognizing and correcting handwritten Chinese characters provided by an embodiment of the present invention. Detailed Embodiments
[0072] The present invention will be described in detail below with reference to the accompanying drawings.
[0073] Embodiment 1
[0074] An embodiment of the present invention provides a method for recognizing and correcting handwritten Chinese characters, as Figure 1 shown, including the following steps:
[0075] S1: Obtain a handwritten Chinese character image and perform noise reduction processing on it;
[0076] S2: Perform image segmentation on the noise-reduced handwritten Chinese character image to extract a to-be-recognized image with Chinese characters;
[0077] S3: Obtain the bounding box of a single Chinese character in the to-be-recognized image based on template matching;
[0078] S4: Establish a training data set and a first recognition model for target handwritten Chinese characters, and train the first recognition model for target handwritten Chinese characters based on the training data set;
[0079] S5: Detect the image to be recognized based on the trained first recognition model for target handwritten Chinese characters, and obtain the first recognition result of the handwritten Chinese characters in the image to be recognized. The first recognition result includes scribbled fonts and non-scribbled fonts;
[0080] S6: Correct the scribbled fonts in the first recognition result based on the hypergraph neural network.
[0081] Further, in step S1, the noise reduction process is Gaussian filtering, specifically:
[0082] ,
[0083] where g(x, y) is the handwritten Chinese character image, (x, y) is the handwritten Chinese character image after Gaussian filtering, u(i, j) is the weight, and D(x, y) is the size range of (N + 1)×(N + 1) centered on (x, y);
[0084] wherein the weight is specifically:
[0085]
[0086] where, (i, j) is the spatial channel proximity weight value, (i, j) is the spatial channel similarity weight value, is the input random coefficient.
[0087] Further, in step S2, the image segmentation of the handwritten Chinese character image after noise reduction is specifically: Image segmentation of the handwritten Chinese character image after noise reduction is performed based on finding the target minimum bounding rectangle, and specifically includes the following steps:
[0088] S21: Perform binarization processing on the handwritten Chinese character image after noise reduction to obtain a binarized image;
[0089] S22: Based on the contour detection algorithm, search for and draw the contours of the binary image, and then fit the drawn target contours to find the minimum bounding square of the contours;
[0090] S23: Based on the center point of the minimum bounding square being at 2 / 4 to 3 / 4 of the handwritten Chinese character image, search for the Chinese character at the center and intercept the segmented image with the central Chinese character;
[0091] S24: Correct the segmented image to obtain the image to be recognized.
[0092] Further, in step S24, the correction of the segmented image to obtain the image to be recognized is specifically:
[0093] S251: Corresponding the four corner points of the minimum circumscribed square with the four corners of the image, and then based on these four pairs of corresponding corner points, obtaining the perspective transformation matrix using the fitgeotrans function in MATLAB;
[0094] S252: Based on the perspective transformation matrix, calling the imwarp function in MATLAB to transform the segmented image into the corrected view, obtaining the image to be recognized;
[0095] S253: Outputting the image to be recognized.
[0096] Further, the specific steps of S3 are as follows:
[0097] S31: Constructing the first image template and the second image template, where the first image template is the standard image of Chinese characters, and the second image template is the standard inner frame image of Chinese characters;
[0098] S32: Sliding the first image template and the second image template on the image to be recognized respectively, and finding the position with the largest best matching value, that is, obtaining the inner and outer frames of the Chinese characters in the image to be recognized.
[0099] Further, in step S4, the specific implementation process of establishing the training data set and the first recognition model for target handwritten Chinese characters is as follows:
[0100] S41: Obtaining the scribbled and non-scribbled Chinese character image data through handwritten Chinese characters, with a part used for the test set and a part used for the training set. Training two EfficientDet network models through the training set, where the feature acquisition network of the EfficientDet network model is ResNeXt, as Figure 2 shown. The two network models obtained after training are respectively denoted as ResNeXt-A and ResNeXt-B. The Chinese character image data is defined with labels manually, including defining two labels of scribbled and non-scribbled;
[0101] S42: Dividing the EfficientDet network model in S41 into the feature extraction network ResNeXt, the feature pyramid, and the scribbled classification and regression network. Using the feature extraction network ResNeXt to extract the features of scribbled and non-scribbled images, connecting the feature pyramid to the feature extraction network ResNeXt, performing convolution on each layer of the feature pyramid, and performing classification and regression through the scribbled classification and regression network;
[0102] S43: Using the test set to perform network iterative testing on the feature extraction network ResNeXt, and outputting the recognition result. When the recognition accuracy rate reaches the set threshold, stopping the training of this model. The recognition result is used to judge whether the image is scribbled.
[0103] Further, the specific construction steps of the feature extraction network ResNeXt are as follows:
[0104] S411: Judge the images in the training set data according to the two labels of scribbled and non-scribbled. If the label of the image is a scribbled image, ResNeXt-A is used for processing. If the label of the image is a non-scribbled image, ResNeXt-B is used for processing. If it is judged that the currently input training set data image is scribbled, a deep convolutional generative adversarial network for scribbled images is constructed;
[0105] S412: Use the deep convolutional generative adversarial network to generate scribbled image training set samples for the scribbled images;
[0106] The processing flow for scribbled and non-scribbled features is as follows:
[0107] The feature acquisition network adopts the network framework of the feature extraction network ResNeXt, extracts the Chinese character features of scribbled and non-scribbled respectively, first performs superposition based on channels on the feature maps of the scribbled and non-scribbled convolutional layers in the second layer, and then performs dimensionality reduction processing through 2×2 convolution. The same operation is performed on the third convolutional layer, the fourth convolutional layer, the fifth convolutional layer, and the sixth convolutional layer;
[0108] The specific process of using the deep convolutional generative adversarial network to generate scribbled image training set samples for the scribbled images is as follows: According to the original scribbled image given in the current training, generate adversarial scribbled image samples. The specifications of the adversarial scribbled image samples are: diverse font size features, scribbled stroke features, and overall scribbled font features. Input 1 original given scribbled image, add a perturbation factor through the deep convolutional generative adversarial network to generate 200 scribbled sample images. The scribbled stroke features at least include the feature that the stroke lengths are not fixed and the feature that the relative positions of the strokes deviate. The overall scribbled font features at least include the feature that the overall font width-to-height ratio deviates and the feature that the overall font center of gravity position deviates.
[0109] Specifically, the scribbled image training set samples can also include using a data augmentation algorithm based on DropSegment.
[0110] The DropSegment method is divided into three steps: First, corner detection is performed. Then, the Chinese characters are segmented into a group of shorter stroke segments according to the detected corners. Finally, a part of the stroke segments is discarded according to the randomly selected method. Each handwritten Chinese character contains at least one stroke, and each stroke contains multiple stroke segments connected by corners. After each Chinese character passes through DropSegment, the remaining stroke segments remain in the same order in the Chinese character and are combined into a new Chinese character with writing features. The corners can be defined by the defined segmentation method.
[0111] It should be noted that when a handwritten Chinese character contains 4 strokes, and each stroke contains 3 stroke segments respectively, the total number of Chinese characters enhanced by the DropSegment algorithm is 32768. Therefore, by randomly deleting some stroke segments from Chinese characters through the DropSegment algorithm, the generated new characters can greatly expand the sample of the scribbled image training set, generate a large number of scribbled image training set samples, and improve the recognition effect of scribbled Chinese characters.
[0112] Furthermore, the specific process of concatenating the feature pyramid to the feature extraction network ResNeXt is as follows:
[0113] Sample the feature map of the third convolutional layer, that is, the feature map of the fusion of scribbled and non-scribbled, and then superimpose it on the feature map of the fusion of scribbled and non-scribbled of the second convolutional layer to obtain the first layer of the pyramid. Continue to execute this step for the fourth convolutional layer, the fifth convolutional layer, and the sixth convolutional layer. Superimpose the feature maps of every two adjacent layers on the channel dimension, and one layer of the pyramid can be obtained. Finally, a total of four layers of feature pyramids are obtained. The specific processing process of the four layers of the pyramid is as follows:
[0114] The first layer of the pyramid is used to process the first image resolution, and the brightness of the image belongs to the first-level brightness;
[0115] The second layer of the pyramid is used to process the second image resolution, and the brightness of the image belongs to the second-level brightness;
[0116] The third layer of the pyramid is used to process the third image resolution, and the brightness of the image belongs to the third-level brightness;
[0117] The fourth layer of the pyramid is used to process the fourth image resolution, and the brightness of the image belongs to the fourth-level brightness;
[0118] The resolution levels of the image are specifically divided as follows:
[0119] The first image resolution: w < 160, h < 160;
[0120] The second image resolution: 160 < w ≤ 320, 160 < h ≤ 320;
[0121] The third image resolution: 320 < w ≤ 1080, 320 < h ≤ 1080;
[0122] The fourth image resolution: 1080 < w, 1080 < h, where w represents the width of the image and h represents the height of the image;
[0123] The average brightness of the image is calculated by using the histogram to calculate the average brightness of the image, and is specifically divided as follows:
[0124] The first average brightness: 0 ≤ L ≤ 60;
[0125] Second average brightness: 60 < L ≤ 100;
[0126] Third average brightness: 100 < L ≤ 160;
[0127] Fourth average brightness: 160 < L ≤ 255, where L represents the average brightness value of the image;
[0128] Two branch networks are added after the feature map of each fusion layer of the pyramid. One branch is used for classification and the other branch is used for regression. And each branch first performs 4 convolutions on the feature map to enhance the scribbled and non-scribbled image features respectively, and the convolution kernel size is 2×2 and the number is 128.
[0129] Furthermore, the specific implementation process of S6 is as follows:
[0130] The stroke scribbling feature and the overall font scribbling feature of the scribbled image to be processed are input into the first hypergraph neural network layer of the preset Chinese character correction model to obtain the first scribbling feature vector; wherein, the first scribbling feature vector represents the association relationship between the stroke scribbling feature and the overall font scribbling feature of the scribbled image to be processed; the preset Chinese character correction model is a model obtained based on the hypergraph neural network model;
[0131] The sub-feature of the stroke scribbling feature and the sub-feature of the overall font scribbling feature of the scribbled image to be processed are input into the second hypergraph neural network layer of the preset Chinese character correction model to obtain the second scribbling feature vector; wherein, the second feature vector represents the association relationship between the sub-feature of the stroke scribbling feature and the sub-feature of the overall font scribbling feature of the scribbled image to be processed;
[0132] The first scribbling feature vector and the second scribbling feature vector are connected to obtain the feature vector to be processed;
[0133] Based on the preset Chinese character correction model, the feature vector to be processed is recognized to obtain the scribbling type of the scribbled image, and the scribbling type at least includes scribbled stroke length, scribbled stroke intersection, scribbled stroke inclination, scribbled pen gesture, scribbled whole word position, scribbled whole word width-to-height ratio, and scribbled whole word center of gravity;
[0134] For the scribbling types in the preset handwritten Chinese character correction library, the target similarity between the scribbling type of the scribbled image to be processed and the scribbling types in the preset handwritten Chinese character correction library is calculated through the cosine similarity algorithm, and a preset similarity threshold. If the target similarity is greater than the preset similarity threshold, the corresponding scribbling type in the handwritten Chinese character correction library is found to correct the Chinese characters in the scribbled image to generate standard Chinese characters. If the target similarity is less than or equal to the preset similarity threshold, the scribbling type in the handwritten Chinese character correction library with the closest similarity is found to correct the Chinese characters in the scribbled image.
[0135] It should be noted that for the method of extracting scribbled features of Chinese characters, by extracting some geometric features of Chinese characters, such as the endpoints, bifurcation points, concave and convex parts of the characters, as well as line segments and closed loops in various directions such as horizontal, vertical, and inclined, logical combination judgments are made based on the positions and mutual relationships of these features to obtain Chinese character scribbled feature information. For handwritten fonts, the Chinese character feature information is relatively complex. Therefore, in the embodiments of the present invention, by extracting the overall and local features of handwritten fonts, the recognition accuracy of handwritten Chinese character scribbled types is effectively improved.
[0136] Specifically, the cosine similarity algorithm:
[0137]
[0138] Among them, represents the scribbled type feature of the scribbled image to be processed, the scribbled type feature preset in the handwritten Chinese character correction library;
[0139] The cosine similarity algorithm can be used to calculate the target similarity of the scribbled image to be processed, the larger the value, the more similar the target similarity is to the preset similarity threshold, and vice versa;
[0140] If the target similarity is less than or equal to the preset similarity threshold, sorting is performed through the similarity value, and the one with the closest sorting value is set as the standard Chinese character corresponding to the scribbled type in the handwritten Chinese character correction library, and the scribbled Chinese character is corrected to generate a standard Chinese character.
[0141] Furthermore, the loss function of the preset Chinese character correction model is:
[0142]
[0143] In the formula, is the loss function, k is the balance factor, is the associated matrix aggregation parameter, is the target probability of the scribbled type to be classified.
[0144] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. Based on the above description, the structure required to construct such a system is obvious. In addition, the present invention is not directed to any particular programming language. It should be understood that the content of the present invention described herein can be implemented using various programming languages, and the description of the specific language above is for the purpose of disclosing the best mode of the present invention.
[0145] In the description provided herein, numerous specific details are set forth. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description.
[0146] Similarly, it should be understood that in order to streamline this disclosure and assist in understanding one or more of the various inventive aspects, in the foregoing description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all the features of a single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the invention.
[0147] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from those of the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract and drawings) can be replaced by an alternative feature providing the same, equivalent or similar purpose.
[0148] In addition, those skilled in the art will be able to understand that although some of the embodiments herein include certain features included in other embodiments but not others, the combination of features of different embodiments means that it is within the scope of the invention and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0149] Each component embodiment of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the device according to the embodiments of the present invention. The present invention can also be implemented as a device or device program (e.g., a computer program and a computer program product) for performing part or all of the methods described herein. Such a program for implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or in any other form.
Claims
1. A method for handwritten Chinese character recognition and correction, characterized in that, It includes the following steps: S1: Obtain a handwritten Chinese character image and perform noise reduction processing on it; S2: Perform image segmentation on the handwritten Chinese character image after noise reduction processing, and extract the image to be recognized with Chinese characters; S3: Obtain the bounding box of a single Chinese character in the image to be recognized based on template matching; S4: Establish a training data set and a first recognition model for target handwritten Chinese characters, and train the first recognition model for target handwritten Chinese characters based on the training data set; S5: Detect the image to be recognized based on the trained first recognition model for target handwritten Chinese characters, and obtain the first recognition result of the handwritten Chinese characters in the image to be recognized. The first recognition result includes scribbled fonts and non-scribbled fonts; S6: Correct the scribbled fonts in the first recognition result based on a hypergraph neural network; In step S4, the specific implementation process of establishing the training data set and the first recognition model for target handwritten Chinese characters is as follows: S41: Obtain scribbled and non-scribbled Chinese character image data through handwritten Chinese characters. A part is used as a test set, and a part is used as a training set. Two EfficientDet network models are trained through the training set. The feature acquisition network of the EfficientDet network model is ResNeXt. The two network models obtained after training are respectively denoted as ResNeXt-A and ResNeXt-B. The Chinese character image data is defined with labels manually, including two labels defined as scribbled and non-scribbled; S42: Divide the EfficientDet network model in S41 into a feature extraction network ResNeXt, a feature pyramid, and a scribbled classification and regression network. Use the feature extraction network ResNeXt to extract the features of scribbled and non-scribbled images, connect the feature pyramid to the feature extraction network ResNeXt, perform convolution on each layer of the feature pyramid, and perform classification and regression through the scribbled classification and regression network; S43: Use the test set to perform network iterative testing on the feature extraction network ResNeXt, and output the recognition result. When the recognition accuracy reaches the set threshold, stop the training of this model. The recognition result is used to judge whether the image is scribbled.
2. The method for handwritten Chinese character recognition and correction according to claim 1, characterized in that, In step S1, the noise reduction processing is Gaussian filtering processing, specifically: , where g(x, y) is the handwritten Chinese character image, (x, y) is the handwritten Chinese character image after Gaussian filtering, u(i, j) is the weight, and D(x, y) is the size range of (N + 1) × (N + 1) centered on (x, y); Among them, the weights are specifically: In the formula, (i, j) is the spatial channel proximity weight value, (i, j) is the spatial channel similarity weight value, is the input random coefficient.
3. A method for handwritten Chinese character recognition and correction according to claim 2, characterized in that, In step S2, the image segmentation of the handwritten Chinese character image after noise reduction processing is specifically: Based on finding the minimum bounding rectangle of the target, perform image segmentation on the handwritten Chinese character image after noise reduction processing. The specific steps are as follows: S21: Perform binarization processing on the handwritten Chinese character image after noise reduction processing to obtain a binarized image; S22: Based on the contour detection algorithm, search for and draw the contours of the binary image, then fit the drawn target contours, and find the minimum bounding square of the contours; S23: Based on the center point of the minimum bounding square being at 2 / 4 to 3 / 4 of the handwritten Chinese character image, search for the Chinese character at the center and intercept the segmented image with the central Chinese character; S24: Correct the segmented image to obtain the image to be recognized.
4. A method for handwritten Chinese character recognition and correction according to claim 3, characterized in that, In step S24, the correction of the segmented image to obtain the image to be recognized is specifically as follows: S251: Corresponding the four corner points of the minimum circumscribed square with the four corners of the image, and then based on these four pairs of corresponding corner points, obtaining the perspective transformation matrix using the fitgeotrans function in MATLAB; S252: Based on the perspective transformation matrix, calling the imwarp function in MATLAB to transform the segmented image into the corrected view to obtain the image to be recognized; S253: Outputting the image to be recognized.
5. A method for handwritten Chinese character recognition and correction according to claim 4, characterized in that, The specific steps of S3 are as follows: S31: Constructing a first image template and a second image template, where the first image template is the standard image of Chinese characters, and the second image template is the standard inner frame image of Chinese characters; S32: Sliding the first image template and the second image template on the image to be recognized respectively to find the position with the largest best matching value, that is, obtaining the inner and outer borders of the Chinese characters in the image to be recognized.
6. A method for handwritten Chinese character recognition and correction according to claim 1, characterized in that, The specific construction steps of the feature extraction network ResNeXt are as follows: S411: Judging the images in the training set data according to the two labels of scribbled and non-scribbled. If the label of the image is a scribbled image, it is processed by ResNeXt-A. If the label of the image is a non-scribbled image, it is processed by ResNeXt-B. If it is judged that the currently input training set data image is scribbled, a scribbled image deep convolutional generative adversarial network is constructed; S412: Using the deep convolutional generative adversarial network to generate scribbled image training set samples for the scribbled image; The processing flow for scribbled and non-scribbled features is as follows: The feature acquisition network adopts the network framework of the feature extraction network ResNeXt, extracts the Chinese character features of scribbled and non-scribbled respectively, first performs channel-based superposition on the feature maps of the scribbled and non-scribbled convolutional layers in the second layer, and then performs dimensionality reduction processing through a 2×2 convolution. The same operation is performed on the third convolutional layer, the fourth convolutional layer, the fifth convolutional layer, and the sixth convolutional layer; The specific process of using the deep convolutional generative adversarial network to generate scribbled image training set samples for the scribbled image is as follows: According to the original scribbled image given in the current training, generating adversarial scribbled image samples. The specifications of the adversarial scribbled image samples are: diverse font size features, scribbled stroke features, and overall scribbled font features. Inputting 1 original given scribbled image, adding a perturbation factor through the deep convolutional generative adversarial network to generate 200 scribbled sample images. The scribbled stroke features at least include the feature of non-fixed stroke length and the feature of deviation of the relative position of strokes. The overall scribbled font features at least include the feature of deviation of the overall font width-to-height ratio and the feature of deviation of the overall font center of gravity position.
7. A method for handwritten Chinese character recognition and correction according to claim 1, characterized in that, The specific process of connecting the feature pyramid to the feature extraction network ResNeXt is as follows: Sample the feature map of the third convolutional layer, i.e., the fused feature map of scribbled and non-scribbled, and then superimpose it on the fused feature map of scribbled and non-scribbled of the second convolutional layer to obtain the first layer of the pyramid. Continue to perform this step on the fourth convolutional layer, the fifth convolutional layer, and the sixth convolutional layer. Superimpose the feature maps of every two adjacent layers channel by channel to obtain one layer of the pyramid. Finally, a total of four layers of feature pyramids are obtained. The specific processing process of the four layers of the pyramid is as follows: The first layer of the pyramid is used to process the first image resolution, and the brightness of the image belongs to the first-level brightness; The second layer of the pyramid is used to process the second image resolution, and the brightness of the image belongs to the second-level brightness; The third layer of the pyramid is used to process the third image resolution, and the brightness of the image belongs to the third-level brightness; The fourth layer of the pyramid is used to process the fourth image resolution, and the brightness of the image belongs to the fourth-level brightness; The resolution levels of the images are specifically divided as follows: The first image resolution: w < 160, h < 160; The second image resolution: 160 < w ≤ 320, 160 < h ≤ 320; The third image resolution: 320 < w ≤ 1080, 320 < h ≤ 1080; The fourth image resolution: 1080 < w, 1080 < h, where w represents the width of the image and h represents the height of the image; The average brightness of the images is calculated by using the histogram method, and is specifically divided as follows: The first average brightness: 0 ≤ L ≤ 60; The second average brightness: 60 < L ≤ 100; The third average brightness: 100 < L ≤ 160; The fourth average brightness: 160 < L ≤ 255, where L represents the average brightness value of the image; Two branch networks are added after the feature map of each fusion layer of the pyramid. One branch is used for classification, and one branch is used for regression. And each branch first performs 6 convolutions on the feature map to enhance the image features of scribbled and non-scribbled respectively, and the convolution kernel size is 3×3 and the number is 64.
8. A method for handwritten Chinese character recognition and correction according to claim 1, characterized in that, The specific implementation process of S6 is as follows: Input the stroke scribble feature and the overall font scribble feature of the scribbled image to be processed into the first hypergraph neural network layer of the preset Chinese character correction model to obtain the first scribble feature vector; wherein, the first scribble feature vector represents the association relationship between the stroke scribble feature and the overall font scribble feature of the scribbled image to be processed; the preset Chinese character correction model is a model obtained based on the hypergraph neural network model; Input the sub-feature of the stroke scribble feature and the sub-feature of the overall font scribble feature of the scribbled image to be processed into the second hypergraph neural network layer of the preset Chinese character correction model to obtain the second scribble feature vector; wherein, the second scribble feature vector represents the association relationship between the sub-feature of the stroke scribble feature and the sub-feature of the overall font scribble feature of the scribbled image to be processed; Connect the first scribble feature vector and the second scribble feature vector to obtain the feature vector to be processed; Identifying the to-be-processed feature vector based on the preset Chinese character correction model to obtain the scribbled type of the scribbled image, where the scribbled type at least includes scribbled stroke length, scribbled stroke connection, scribbled stroke inclination, scribbled pen movement, scribbled whole-character position, scribbled whole-character width-to-height ratio, and scribbled whole-character center of gravity; For the scribbled types in the preset handwritten Chinese character correction library, calculate the target similarity between the scribbled type of the to-be-processed scribbled image and the scribbled types in the preset handwritten Chinese character correction library through the cosine similarity algorithm, and preset a similarity threshold. If the target similarity is greater than the preset similarity threshold, find the corresponding scribbled type in the handwritten Chinese character correction library and then correct the Chinese characters in the scribbled image to generate standard Chinese characters. If the target similarity is less than or equal to the preset similarity threshold, find the scribbled type in the handwritten Chinese character correction library with the closest similarity to correct the Chinese characters in the scribbled image.
9. A method for handwritten Chinese character recognition and correction according to claim 8, characterized in that, The loss function of the preset Chinese character correction model is: In the formula, is the loss function, k is the balance factor, is the correlation matrix aggregation parameter, is the target probability of the scribble type to be classified.
Citation Information
Patent Citations
handwritten Chinese recognition method based on deep learning
CN109800763A
Text review and error correction system based on natural language processing
CN116341525A
Cited By
Handwritten text recognition method and system based on image-structure multi-mode mutual learning
CN121259846A