Handwritten Chinese character recognition and error correction method based on artificial intelligence
By constructing various non-standard data sample sets and graph convolutional neural networks, the problem of low accuracy and insufficient recognition of untrained and non-standard Chinese characters in existing handwritten Chinese character recognition models is solved, and fast and accurate error correction of non-standard Chinese characters is achieved.
Patent Information
- Application Number
- CN202511276387.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-11-14
AI Technical Summary
Existing handwritten Chinese character recognition models have low accuracy for untrained handwritten Chinese characters, insufficient generalization ability, inability to recognize non-standard Chinese characters, and limited error correction capabilities.
An AI-based Chinese character recognition and error correction method is constructed. By establishing multiple non-standard data sample sets, noise perturbation preprocessing is performed, a Chinese character recognition attention model is constructed, and graph convolutional neural networks are used for initial classification and refined type determination. Finally, graph convolutional neural networks are used to correct non-standard types.
It improves the speed and accuracy of handwritten Chinese character recognition, can quickly and accurately correct non-standard Chinese characters, and enhances the model's generalization ability and recognition efficiency.
Smart Images

Figure CN120954019A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of handwritten Chinese character recognition technology, and more specifically, to an artificial intelligence-based method for handwritten Chinese character recognition and error correction. Background Technology
[0002] Handwritten Text Recognition (HTR) is a crucial research area in computer vision and pattern recognition, widely applied in document digitization, education, finance, and other fields. With the rapid development of artificial intelligence, the accuracy and efficiency of HTR have significantly improved. However, the diversity and complexity of handwritten characters still present numerous challenges. For example, differences in writing styles among individuals, the phenomenon of cursive writing, and illegible handwriting can all lead to recognition errors.
[0003] Chinese patent application CN109800763A discloses a deep learning-based method for recognizing handwritten Chinese characters. This method improves the performance of handwritten Chinese character recognition by combining P2DMN normalization, NCFE feature extraction, ADBN coarse classification, and MQDF fine classification with deep learning technology. However, the above method has low accuracy in recognizing untrained handwritten Chinese characters, weak generalization ability, and requires a large dataset for training to prevent overfitting, making it unable to provide a large amount of training data. Furthermore, it does not recognize non-standard handwritten Chinese characters, and even when recognizing standard Chinese characters, it lacks a multiple classification method for initial and fine classification of non-standard characters.
[0004] In existing technologies, the recognition of handwritten Chinese characters relies on a single neural network model, which can easily lead to a low accuracy rate in handwritten Chinese character recognition. In addition, when encountering non-standard fonts in handwritten Chinese characters, neural networks are needed for rapid and accurate correction. Summary of the Invention
[0005] The purpose of this invention is to provide an artificial intelligence-based method for recognizing and correcting handwritten Chinese characters, in order to solve the aforementioned problems existing in the prior art.
[0006] The application is as follows:
[0007] A handwritten Chinese character recognition and error correction method based on artificial intelligence, the method comprising:
[0008] S1. Establish a data sample set corresponding to Chinese characters with various non-standard types;
[0009] S2. Perform preprocessing on the data sample set, including adding noise perturbation, to obtain the sample coefficients of each non-normal type;
[0010] S3. Construct a Chinese character recognition attention model, wherein the Chinese character recognition attention model includes an input layer, an embedding layer, a convolutional fusion layer, a pooling layer, and a pyramid classification layer connected in sequence. The convolutional fusion layer includes multiple convolutional fusion units, each of which includes a depth / shallow module, a channel fusion module, and a positional attention module connected in sequence. The pyramid classification layer includes two fully connected layers connected in series. The depth / shallow module has a deep convolutional separable unit and a shallow convolutional separable unit connected in a cascaded manner. The kernel size of the deep convolutional separable unit is larger than the kernel size of the shallow convolutional separable unit. The channel fusion module has a pixel-wise convolutional unit.
[0011] S4. Optimize the hyperparameters of the Chinese character recognition attention model by combining the sample coefficients to obtain the optimized Chinese character recognition attention model;
[0012] S5. Use the preprocessed data sample set as the training set to train the optimized Chinese character recognition attention model;
[0013] S6. Acquire images of handwritten Chinese characters to be recognized, segment the image of handwritten Chinese characters to be recognized, extract the region to be recognized containing Chinese characters, use a multi-scale template matching algorithm to determine the bounding box of a single Chinese character in the region to be recognized, and generate the handwritten Chinese character image to be processed.
[0014] S7. Input the handwritten Chinese character image to be processed into the trained optimized Chinese character recognition attention model, and output the non-standard type of the handwritten Chinese character to be recognized.
[0015] S8. Based on the graph convolutional neural network, refine the type determination of non-standard types in the recognition results, and make corrections based on the determination results.
[0016] Furthermore, S2 specifically includes:
[0017] With the constraint that the feature entropy of the data sample set is greater than a preset feature entropy, Poisson noise is injected into each sample image in the data sample set:
[0018] ,
[0019] Wherein, noise represents random noise following a Poisson distribution with a mean of 0, and Y1 represents the image pixel matrix in the data sample set. This represents the image pixel matrix after adding noise, where T represents the feature entropy of the data sample set and the preset feature entropy. The probability of pixel j appearing after the i-th image in the data sample set is converted to a grayscale image is represented by N, where N represents the number of sample images in the data sample set.
[0020] Calculate the sample image coefficients for each non-standard category of Chinese characters in the data sample set:
[0021] ,
[0022] in, This represents the coefficients of the sample images for the k1th class of non-standard Chinese characters. This represents the number of sample images of the non-standard category of Chinese characters in the k1th class in the data sample set.
[0023] Furthermore, the method for constructing the embedding layer specifically includes:
[0024] Receive the input image from the input layer;
[0025] The input image is divided into multiple local image blocks of size G1×G2, where G1 and G2 represent the number of columns and rows in the X and Y directions, respectively;
[0026] The input image is embedded into a low-dimensional image patch feature map:
[0027] ,
[0028] ,
[0029] ,
[0030] in, A represents the low-dimensional image patch feature map. , These represent the dimensionality channels of the low-dimensional image patch feature map, i.e., the channels, height, and width of the embedding layer. H and W represent the number of channels, height, and width of the input image.
[0031] Furthermore, the depth-to-shallow convolution module further includes a first activation unit and a first regression unit. The depth-separable convolution unit, the shallow convolution separable unit, the first activation unit, and the first regression unit are connected sequentially. The method for constructing the depth-to-shallow module specifically includes:
[0032] Receive the low-dimensional image block feature map;
[0033] The low-dimensional image patch feature map is convolved by deep convolutional separable units and shallow convolutional separable units connected in a cascaded manner, and a convolutional feature map is output:
[0034] ,
[0035] Where (k, m) represents the two-dimensional coordinates of the low-dimensional image patch feature map. This represents the convolutional feature map. This represents the convolution kernel of the depthwise separable unit. This represents the convolution kernel of the shallow convolutional separable unit, where (i1,j1) represents the convolution kernel. In the i1th row and j1st column, (u, v) represents the convolution kernel. The u-th row and v-th column.
[0036] Furthermore, the channel fusion module also includes a one-dimensional channel unit, a two-dimensional channel unit, a second regression fusion unit, and the pixel-by-pixel convolution unit; the one-dimensional channel unit, the two-dimensional channel unit, the second regression fusion unit, and the position attention module are connected sequentially, and the channel fusion module is constructed as follows:
[0037] Receive the convolutional feature map;
[0038] The convolutional feature map is processed sequentially by the pixel-wise convolutional unit, the one-dimensional channel unit, and the two-dimensional channel unit to capture the one-dimensional and two-dimensional channel information of the convolutional feature map, thereby obtaining a one-dimensional feature map and a two-dimensional feature map, respectively.
[0039] ,
[0040] in, This refers to the one-dimensional feature map or the two-dimensional feature map. denoted as Softmax activation function, GN represents batch normalization, and GWConv() represents a pixel-wise convolutional unit, wherein the kernel size of the pixel-wise convolutional unit is 2×2×A×A;
[0041] Based on the preset channel descent coefficient, the one-dimensional feature map and the two-dimensional feature map are divided into multiple channel groups in order of gradually decreasing channel number. Self-attention is calculated for the features in each channel group to obtain the group attention features. Based on the deformable convolutional network, the local spatial dependency relationship of the group attention features is calculated to obtain the spatial dependency weight. The spatial dependency weight is weighted with the group attention features to obtain the first weighted feature.
[0042] The mutual information between adjacent one-dimensional feature maps and two-dimensional feature maps is calculated to obtain the corresponding feature correlation coefficients. Based on the feature correlation coefficients, the first weighted feature is downsampled and passed to obtain a downsampled feature. The first weighted feature is then upsampled and passed to obtain an upsampled feature. The downsampled feature and the upsampled feature are combined to obtain a second weighted feature.
[0043] Multi-head attention operations are performed on the second weighted feature in both the channel dimension and the spatial dimension to obtain channel attention weights and spatial attention weights. The channel attention weights are weighted with the channel information of the second weighted feature to obtain channel information weighted features, and the spatial attention weights are weighted with the spatial information of the second weighted feature to obtain spatial information weighted features.
[0044] The second regression fusion unit combines the weighted features of the channel information with the weighted features of the spatial information to generate fused features.
[0045] Furthermore, the positional attention module includes pixel-wise convolutional units, and the specific construction method of the positional attention module includes:
[0046] The location attention module performs global feature extraction on the fused features to obtain an attention-based global feature map with the same size as the low-dimensional image patch feature map.
[0047] ,
[0048] in, The attention feature is represented by GAtt(), which represents the positional attention module.
[0049] Furthermore, the pyramid classification layer includes a pyramid classification sublayer and a pyramid regression sublayer;
[0050] The specific methods for constructing the pyramid classification layer include:
[0051] The global attention feature map is classified by a pyramid classification sub-layer, and various non-standard types of Chinese characters are output.
[0052] The pyramid classification sub-layer output results are regressed and confirmed by the pyramid regression sub-layer, and the final non-standard types of Chinese characters are output. The non-standard types include at least stroke breakage, component misalignment, loose structure, disproportion, center of gravity shift and stroke order error.
[0053] The original images of various non-standard Chinese character input data samples are found. The latent space feature vectors of the original images are extracted by the encoder module of the variational autoencoder. Random perturbation and interpolation operations are applied to the feature vectors in the latent space to generate diverse non-standard feature vectors. After the non-standard feature vectors are reconstructed by the decoder, Chinese character image samples with non-standard stroke features and overall structural features are output.
[0054] The pyramid classification sublayer performs 5 convolutions on the global attention feature map to enhance the image features of non-standard Chinese characters, and the convolution kernel size is 2×2 with a number of 32.
[0055] Further, in step S6, the handwritten Chinese character image to be recognized is segmented into regions, and the regions containing the Chinese characters to be recognized are extracted, including:
[0056] Adaptive binarization is performed on the image of the handwritten Chinese character to be recognized to generate a binary image;
[0057] An edge detection algorithm is used to locate the contours of Chinese characters in a binary image, and the minimum bounding rectangle of each contour is calculated.
[0058] Based on the spatial distribution characteristics of Chinese characters, the outlines of Chinese characters in the central region of the image are selected;
[0059] Geometric correction is performed on the selected Chinese character outline area to obtain a standardized region to be recognized.
[0060] Furthermore, the geometric correction specifically includes:
[0061] Establish the mapping relationship between the corner points of the minimum bounding rectangle and the corner points of the standard rectangle;
[0062] Calculate the image transformation matrix based on the perspective transformation algorithm;
[0063] Apply an image transformation matrix to perform an affine transformation on the Chinese character region;
[0064] Output the corrected, standardized Chinese character images.
[0065] Furthermore, the specific implementation process of S8 is as follows:
[0066] The non-standard stroke features and overall structural features of the handwritten Chinese character image to be processed are input into the first graph convolutional network layer of the preset Chinese character correction model to obtain the first non-standard feature vector; wherein, the first non-standard feature vector represents the topological relationship between the non-standard stroke features and overall structural features of the handwritten Chinese character image to be processed; the preset Chinese character correction model is a correction model constructed based on a multi-layer graph convolutional network;
[0067] The local deformation features and global layout features of the handwritten Chinese character image to be processed are input into the second graph convolutional network layer of the preset Chinese character correction model to obtain the second non-normal feature vector; wherein, the second non-normal feature vector represents the spatial constraint relationship between the local deformation features and global layout features of the handwritten Chinese character image to be processed.
[0068] The first non-normalized feature vector and the second non-normalized feature vector are concatenated to obtain the concatenated corrected feature vector.
[0069] Based on the preset Chinese character correction model, the splicing correction feature vector is analyzed to obtain the second non-standard type of the handwritten Chinese character;
[0070] The preset standard Chinese character correction library stores correction templates corresponding to various second non-standard types. The matching degree between the second non-standard type of the handwritten Chinese character to be processed and the non-standard type in the standard Chinese character correction library is calculated by the feature space distance metric algorithm. A matching degree threshold is set. If the matching degree is greater than the threshold, the corresponding standard correction template is called to reconstruct the non-standard Chinese character and generate a standard Chinese character. If the matching degree is less than or equal to the threshold, the correction template with the highest matching degree is selected for correction.
[0071] Compared with the prior art, the present invention achieves the following beneficial effects:
[0072] This invention establishes a data sample set corresponding to various non-standard types of Chinese characters; preprocesses the data sample set including adding noise perturbation to obtain sample coefficients for each non-standard type; constructs a Chinese character recognition attention model; trains the Chinese character recognition attention model using the preprocessed data sample set as the training set; acquires and preprocesses images of handwritten Chinese characters to be recognized to generate images of handwritten Chinese characters to be processed; inputs the images of handwritten Chinese characters to be processed into the trained Chinese character recognition attention model, outputting the non-standard types of the handwritten Chinese characters to be recognized; and corrects the non-standard types in the recognition results based on a graph convolutional neural network. This invention uses a Chinese character recognition attention model for initial classification of non-standard Chinese characters and a graph convolutional neural network for secondary fine classification, thus achieving fast recognition of non-standard Chinese characters and high accuracy in correcting them. Attached Figure Description
[0073] Figure 1 This is a flowchart illustrating an artificial intelligence-based handwritten Chinese character recognition and correction method provided in an embodiment of the present invention.
[0074] Figure 2 This is a schematic diagram of the operation of the Chinese character recognition attention model in an artificial intelligence-based handwritten Chinese character recognition and error correction method provided in an embodiment of the present invention. Detailed Implementation
[0075] The present invention will now be described in detail with reference to the accompanying drawings.
[0076] Example 1
[0077] This invention provides an artificial intelligence-based method for handwritten Chinese character recognition and error correction, the method comprising:
[0078] S1. Establish a data sample set corresponding to Chinese characters with various non-standard types;
[0079] S2. Perform preprocessing on the data sample set, including adding noise perturbation, to obtain the sample coefficients of each non-normal type;
[0080] S3. Construct a Chinese character recognition attention model, wherein the Chinese character recognition attention model includes an input layer, an embedding layer, a convolutional fusion layer, a pooling layer, and a pyramid classification layer connected in sequence. The convolutional fusion layer includes multiple convolutional fusion units, each of which includes a depth / shallow module, a channel fusion module, and a positional attention module connected in sequence. The pyramid classification layer includes two fully connected layers connected in series. The depth / shallow module has a deep convolutional separable unit and a shallow convolutional separable unit connected in a cascaded manner. The kernel size of the deep convolutional separable unit is larger than the kernel size of the shallow convolutional separable unit. The channel fusion module has a pixel-wise convolutional unit.
[0081] S4. Optimize the hyperparameters of the Chinese character recognition attention model by combining the sample coefficients to obtain the optimized Chinese character recognition attention model;
[0082] S5. Use the preprocessed data sample set as the training set to train the optimized Chinese character recognition attention model;
[0083] S6. Acquire images of handwritten Chinese characters to be recognized, segment the image of handwritten Chinese characters to be recognized, extract the region to be recognized containing Chinese characters, use a multi-scale template matching algorithm to determine the bounding box of a single Chinese character in the region to be recognized, and generate the handwritten Chinese character image to be processed.
[0084] S7. Input the handwritten Chinese character image to be processed into the trained optimized Chinese character recognition attention model, and output the non-standard type of the handwritten Chinese character to be recognized.
[0085] S8. Based on the graph convolutional neural network, refine the type determination of non-standard types in the recognition results, and make corrections based on the determination results.
[0086] Specifically, a data sample set corresponding to various non-standard types of Chinese characters is established; the data sample set is preprocessed, including the addition of noise perturbation, to obtain sample coefficients for each non-standard type; a Chinese character recognition attention model is constructed; the preprocessed data sample set is used as a training set to train the Chinese character recognition attention model; images of handwritten Chinese characters to be recognized are acquired and preprocessed to generate images of handwritten Chinese characters to be processed; the images of handwritten Chinese characters to be processed are input into the trained Chinese character recognition attention model, and the non-standard types of the handwritten Chinese characters to be recognized are output; the non-standard types in the recognition results are corrected based on a graph convolutional neural network. This invention uses a Chinese character recognition attention model to perform initial classification of non-standard Chinese characters in images and a graph convolutional neural network to perform secondary fine classification, thus achieving fast recognition of non-standard Chinese characters and high accuracy in correcting them. The Chinese character recognition attention model is as follows: Figure 2 As shown.
[0087] In the above embodiments, specifically, S2 includes:
[0088] With the constraint that the feature entropy of the data sample set is greater than a preset feature entropy, Poisson noise is injected into each sample image in the data sample set:
[0089] ,
[0090] Wherein, noise represents random noise following a Poisson distribution with a mean of 0, and Y1 represents the image pixel matrix in the data sample set. This represents the image pixel matrix after adding noise, where T represents the feature entropy of the data sample set and the preset feature entropy. The probability of pixel j appearing after the i-th image in the data sample set is converted to a grayscale image is represented by N, where N represents the number of sample images in the data sample set.
[0091] Calculate the sample image coefficients for each non-standard category of Chinese characters in the data sample set:
[0092] ,
[0093] in, This represents the coefficients of the sample images for the k1th class of non-standard Chinese characters. This represents the number of sample images of the non-standard category of Chinese characters in the k1th class in the data sample set.
[0094] In the above embodiments, specifically, the method for constructing the embedding layer includes:
[0095] Receive the input image from the input layer;
[0096] The input image is divided into multiple local image blocks of size G1×G2, where G1 and G2 represent the number of columns and rows in the X and Y directions, respectively;
[0097] The input image is embedded into a low-dimensional image patch feature map:
[0098] ,
[0099] ,
[0100] ,
[0101] in, A represents the low-dimensional image patch feature map. , These represent the dimensionality channels of the low-dimensional image patch feature map, i.e., the channels, height, and width of the embedding layer. H and W represent the number of channels, height, and width of the input image.
[0102] In the above embodiments, specifically, the depth-to-shallow convolution module further includes a first activation unit and a first regression unit, wherein the depthwise separable convolution unit, the shallow convolution separable unit, the first activation unit, and the first regression unit are connected sequentially, and the method for constructing the depth-to-shallow module specifically includes:
[0103] Receive the low-dimensional image block feature map;
[0104] The low-dimensional image patch feature map is convolved by deep convolutional separable units and shallow convolutional separable units connected in a cascaded manner, and a convolutional feature map is output:
[0105] ,
[0106] Where (k, m) represents the two-dimensional coordinates of the low-dimensional image patch feature map. This represents the convolutional feature map. This represents the convolution kernel of the depthwise separable unit. This represents the convolution kernel of the shallow convolutional separable unit, where (i1,j1) represents the convolution kernel. In the i1th row and j1st column, (u, v) represents the convolution kernel. The u-th row and v-th column.
[0107] Furthermore, the channel fusion module also includes a one-dimensional channel unit, a two-dimensional channel unit, a second regression fusion unit, and the pixel-by-pixel convolution unit; the one-dimensional channel unit, the two-dimensional channel unit, the second regression fusion unit, and the position attention module are connected sequentially, and the channel fusion module is constructed as follows:
[0108] Receive the convolutional feature map;
[0109] The convolutional feature map is processed sequentially by the pixel-wise convolutional unit, the one-dimensional channel unit, and the two-dimensional channel unit to capture the one-dimensional and two-dimensional channel information of the convolutional feature map, thereby obtaining a one-dimensional feature map and a two-dimensional feature map, respectively.
[0110] ,
[0111] in, This refers to the one-dimensional feature map or the two-dimensional feature map. denoted as Softmax activation function, GN represents batch normalization, and GWConv() represents a pixel-wise convolutional unit, wherein the kernel size of the pixel-wise convolutional unit is 2×2×A×A;
[0112] Based on the preset channel descent coefficient, the one-dimensional feature map and the two-dimensional feature map are divided into multiple channel groups in order of gradually decreasing channel number. Self-attention is calculated for the features in each channel group to obtain the group attention features. Based on the deformable convolutional network, the local spatial dependency relationship of the group attention features is calculated to obtain the spatial dependency weight. The spatial dependency weight is weighted with the group attention features to obtain the first weighted feature.
[0113] The mutual information between adjacent one-dimensional feature maps and two-dimensional feature maps is calculated to obtain the corresponding feature correlation coefficients. Based on the feature correlation coefficients, the first weighted feature is downsampled and passed to obtain a downsampled feature. The first weighted feature is then upsampled and passed to obtain an upsampled feature. The downsampled feature and the upsampled feature are combined to obtain a second weighted feature.
[0114] Multi-head attention operations are performed on the second weighted feature in both the channel dimension and the spatial dimension to obtain channel attention weights and spatial attention weights. The channel attention weights are weighted with the channel information of the second weighted feature to obtain channel information weighted features, and the spatial attention weights are weighted with the spatial information of the second weighted feature to obtain spatial information weighted features.
[0115] The second regression fusion unit combines the weighted features of the channel information with the weighted features of the spatial information to generate fused features.
[0116] In the above embodiments, specifically, the position attention module includes pixel-wise convolutional units, and the construction method of the position attention module specifically includes:
[0117] The location attention module performs global feature extraction on the fused features to obtain an attention-based global feature map with the same size as the low-dimensional image patch feature map.
[0118] ,
[0119] in, The attention feature is represented by GAtt(), which represents the positional attention module.
[0120] In the above embodiments, specifically, the pyramid classification layer includes a pyramid classification sublayer and a pyramid regression sublayer;
[0121] The specific methods for constructing the pyramid classification layer include:
[0122] The global attention feature map is classified by a pyramid classification sub-layer, and various non-standard types of Chinese characters are output.
[0123] The pyramid classification sub-layer output results are regressed and confirmed by the pyramid regression sub-layer, and the final non-standard types of Chinese characters are output. The non-standard types include at least stroke breakage, component misalignment, loose structure, disproportion, center of gravity shift and stroke order error.
[0124] The original images of various non-standard Chinese character input data samples are found. The latent space feature vectors of the original images are extracted by the encoder module of the variational autoencoder. Random perturbation and interpolation operations are applied to the feature vectors in the latent space to generate diverse non-standard feature vectors. After the non-standard feature vectors are reconstructed by the decoder, Chinese character image samples with non-standard stroke features and overall structural features are output.
[0125] The pyramid classification sublayer performs 5 convolutions on the global attention feature map to enhance the image features of non-standard Chinese characters, and the convolution kernel size is 2×2 with a number of 32.
[0126] Specifically, the image training set samples corresponding to each non-standard Chinese character can also include those generated using the DropDistortion-based data augmentation algorithm;
[0127] The DropDistortion method consists of three steps: First, corner detection is performed. Then, the Chinese character is segmented into a group of shorter stroke segments according to the detected corner points. Finally, some stroke segments are discarded using a random selection method. Each handwritten Chinese character contains at least one stroke, and each stroke contains multiple stroke segments connected by corner points. After each Chinese character undergoes DropDistortion, the remaining stroke segments maintain their original order within the character, thus synthesizing a new Chinese character containing writing features. Corner points can be defined using a defined segmentation method.
[0128] It should be noted that when a handwritten Chinese character contains 5 strokes, and each stroke contains 4 stroke segments, the total number of Chinese characters enhanced by the DropDistortion algorithm is 65536. Therefore, by randomly deleting some stroke segments from Chinese characters using the DropDistortion algorithm, the generated new characters can greatly expand the training set of non-standard Chinese character images, generate a massive training set of non-standard Chinese character images, and improve the recognition effect of non-standard Chinese characters.
[0129] In the above embodiments, specifically, step S6 involves region segmentation of the handwritten Chinese character image to be recognized, extracting the region to be recognized containing the Chinese characters, including:
[0130] Adaptive binarization is performed on the image of the handwritten Chinese character to be recognized to generate a binary image;
[0131] An edge detection algorithm is used to locate the contours of Chinese characters in a binary image, and the minimum bounding rectangle of each contour is calculated.
[0132] Based on the spatial distribution characteristics of Chinese characters, the outlines of Chinese characters in the central region of the image are selected;
[0133] Geometric correction is performed on the selected Chinese character outline area to obtain a standardized region to be recognized.
[0134] In the above embodiments, specifically, the geometric correction includes:
[0135] Establish the mapping relationship between the corner points of the minimum bounding rectangle and the corner points of the standard rectangle;
[0136] Calculate the image transformation matrix based on the perspective transformation algorithm;
[0137] Apply an image transformation matrix to perform an affine transformation on the Chinese character region;
[0138] Output the corrected, standardized Chinese character images.
[0139] In the above embodiments, the specific implementation process of S8 is as follows:
[0140] The non-standard stroke features and overall structural features of the handwritten Chinese character image to be processed are input into the first graph convolutional network layer of the preset Chinese character correction model to obtain the first non-standard feature vector; wherein, the first non-standard feature vector represents the topological relationship between the non-standard stroke features and overall structural features of the handwritten Chinese character image to be processed; the preset Chinese character correction model is a correction model constructed based on a multi-layer graph convolutional network;
[0141] The local deformation features and global layout features of the handwritten Chinese character image to be processed are input into the second graph convolutional network layer of the preset Chinese character correction model to obtain the second non-normal feature vector; wherein, the second non-normal feature vector represents the spatial constraint relationship between the local deformation features and global layout features of the handwritten Chinese character image to be processed.
[0142] The first non-normalized feature vector and the second non-normalized feature vector are concatenated to obtain the concatenated corrected feature vector.
[0143] Based on the preset Chinese character correction model, the splicing correction feature vector is analyzed to obtain the second non-standard type of the handwritten Chinese character;
[0144] The preset standard Chinese character correction library stores correction templates corresponding to various second non-standard types. The matching degree between the second non-standard type of the handwritten Chinese character to be processed and the non-standard types in the standard Chinese character correction library is calculated through the feature space distance measurement algorithm, and a matching degree threshold is set. If the matching degree is greater than the threshold, the corresponding standard correction template is called to reconstruct the non-standard Chinese character to generate a standard Chinese character; if the matching degree is less than or equal to the threshold, the correction template with the highest matching degree is selected for correction.
[0145] It should be noted that for the method of extracting non-standard stroke features of Chinese characters, by extracting some geometric features of Chinese characters, such as the bifurcation points, endpoints, concave and convex parts of Chinese characters, as well as line segments and closed loops in various directions such as horizontal, vertical, and inclined, logical combination judgments are made based on the positions and mutual relationships of these features to obtain non-standard stroke feature information. For handwritten fonts, the Chinese character feature information is relatively complex. Therefore, in the embodiments of the present invention, the topological correlation relationship between the non-standard stroke features and the overall structural features of handwritten fonts, and further the spatial constraint relationship between the local deformation features and the global layout features of handwritten Chinese character images effectively improve the secondary recognition and confirmation of non-standard stroke feature types, and further subdivide the corresponding subtypes based on the non-standard types. The subtypes are represented by the second non-standard types, and the accuracy of recognition is further improved through the confirmation of the second non-standard types.
[0146] It should be noted that the non-standard types at least include stroke breakage, component dislocation, loose structure, proportion imbalance, center of gravity deviation, and incorrect stroke order;
[0147] The second non-standard types corresponding to each type of the non-standard types include:
[0148] The second non-standard types corresponding to stroke breakage include: local interruption: discontinuity appears in the middle of the stroke (such as a horizontal stroke being broken in the middle); end missing: the ending of the stroke is not completed (such as the tip of a left-falling or right-falling stroke not emerging); virtual stroke phenomenon: the stroke is too light or intermittent, resulting in a visual break; intersection point breakage: the intersection of strokes is not connected (such as the separation of the intersection point of the "十" character);
[0149] The second non-standard types corresponding to component dislocation include: horizontal offset: the left-right position of the component is incorrect (such as the "日" in "明" being shifted to the left); vertical offset: the component is displaced up and down (such as the "子" in "字" sinking); rotation and inclination: the angle of the component deviates (such as the "口" component being inclined); mirror flip: the component is reversed left and right (such as the "女" in "好" being written as a mirror image);
[0150] The second non-standard types corresponding to loose structure include: excessive spacing: too much blank space between components (e.g., the two 'woods' in 'lin' are separated); insufficient adhesion: the strokes that should be connected do not touch (e.g., the two strokes of the character 'person' are separated); scattered center of gravity: the component layout is loose and the overall sense is weak (e.g., the triangular distribution of the three characters in 'pin' is unbalanced); disproportionate ratio: the sizes of components are不协调 (e.g., in 'xie', 'yan' is too large and'she' is too small).
[0151] The second non-standard types corresponding to disproportionate ratio include: local enlargement / reduction: the size of a single stroke or component is abnormal (e.g., the right-falling stroke of the character 'big' is too long); abnormal width-to-height ratio: the overall shape of the character is flattened or elongated (e.g., the character 'yue' is too flat); confusion between primary and secondary: secondary strokes overshadow the primary ones (e.g., the side dot of the character 'yong' is too large); crowded components: multiple components are compressed and overlapped (e.g., the components in 'ying' are stacked).
[0152] The second non-standard types corresponding to offset center of gravity include: heavier on the left and lighter on the right: the density on the left side of the character shape is too high (e.g., the right ear radical in 'du' is too small); floating up or sinking down: the overall is too high or too low (e.g., the lower part of the character 'jing' is suspended); local imbalance: a certain stroke is too heavy (e.g., the vertical stroke in 'zhong' is偏左); dynamically unstable: the center of gravity fluctuates due to continuous writing (e.g., the center of gravity of running script or cursive script is unstable).
[0153] The second non-standard types corresponding to incorrect stroke order include: writing in reverse order: violating the standard stroke order rules (e.g., writing the character 'nine' by first writing the left-falling stroke and then the horizontal fold and hook); skipping strokes: skipping key strokes (e.g., writing the character 'kou' without writing the horizontal stroke); incorrect cross order: the order of stroke crossing is chaotic (e.g., writing the character 'ten' by first writing the vertical stroke and then the horizontal stroke); improper connecting strokes: incorrect connecting strokes in running script or cursive script (e.g., the incorrect connecting stroke order of the three dots in the character 'heart').
[0154] It should be understood that the above embodiments are one or more embodiments of the present invention. Based on the present invention, there are many other embodiments and their variations. When ordinary technicians in this industry do not make pioneering innovations, the variations and modifications made through the present invention all fall within the protection scope of the present invention.
Claims
1. A handwritten Chinese character recognition and error correction method based on artificial intelligence, characterized in that, The method includes: S1. Establish a data sample set corresponding to Chinese characters with various non-standard types; S2. Perform preprocessing on the data sample set, including adding noise perturbation, to obtain the sample coefficients of each non-normal type; S3. Construct a Chinese character recognition attention model, wherein the Chinese character recognition attention model includes an input layer, an embedding layer, a convolutional fusion layer, a pooling layer, and a pyramid classification layer connected in sequence. The convolutional fusion layer includes multiple convolutional fusion units, each of which includes a depth / shallow module, a channel fusion module, and a positional attention module connected in sequence. The pyramid classification layer includes two fully connected layers connected in series. The depth / shallow module has a deep convolutional separable unit and a shallow convolutional separable unit connected in a cascaded manner. The kernel size of the deep convolutional separable unit is larger than the kernel size of the shallow convolutional separable unit. The channel fusion module has a pixel-wise convolutional unit. S4. Optimize the hyperparameters of the Chinese character recognition attention model by combining the sample coefficients to obtain the optimized Chinese character recognition attention model; S5. Use the preprocessed data sample set as the training set to train the optimized Chinese character recognition attention model; S6. Acquire images of handwritten Chinese characters to be recognized, segment the image of handwritten Chinese characters to be recognized, extract the region to be recognized containing Chinese characters, use a multi-scale template matching algorithm to determine the bounding box of a single Chinese character in the region to be recognized, and generate the handwritten Chinese character image to be processed. S7. Input the handwritten Chinese character image to be processed into the trained optimized Chinese character recognition attention model, and output the non-standard type of the handwritten Chinese character to be recognized. S8. Based on the graph convolutional neural network, refine the type determination of non-standard types in the recognition results, and make corrections based on the determination results.
2. The handwritten Chinese character recognition and error correction method based on artificial intelligence according to claim 1, characterized in that, S2 specifically includes: With the constraint that the feature entropy of the data sample set is greater than a preset feature entropy, Poisson noise is injected into each sample image in the data sample set: , Wherein, noise represents random noise following a Poisson distribution with a mean of 0, and Y1 represents the image pixel matrix in the data sample set. This represents the image pixel matrix after adding noise, where T represents the feature entropy of the data sample set and the preset feature entropy. The probability of pixel j appearing after the i-th image in the data sample set is converted to a grayscale image is represented by N, where N represents the number of sample images in the data sample set. Calculate the sample image coefficients for each non-standard category of Chinese characters in the data sample set: , in, This represents the coefficients of the sample images for the k1th class of non-standard Chinese characters. This represents the number of sample images of the non-standard category of Chinese characters in the k1th class in the data sample set.
3. The handwritten Chinese character recognition and error correction method based on artificial intelligence according to claim 1, characterized in that, The method for constructing the embedding layer specifically includes: Receive the input image from the input layer; The input image is divided into multiple local image blocks of size G1×G2, where G1 and G2 represent the number of columns and rows in the X and Y directions, respectively; The input image is embedded into a low-dimensional image patch feature map: , , , in, A represents the low-dimensional image patch feature map. , These represent the dimensionality channels of the low-dimensional image patch feature map, i.e., the channels, height, and width of the embedding layer. H and W represent the number of channels, height, and width of the input image.
4. The handwritten Chinese character recognition and error correction method based on artificial intelligence according to claim 3, characterized in that, The depth-to-shallow module further includes a first activation unit and a first regression unit. The depthwise separable convolutional unit, the shallow convolutional separable unit, the first activation unit, and the first regression unit are connected sequentially. The method for constructing the depth-to-shallow module specifically includes: Receive the low-dimensional image block feature map; The low-dimensional image patch feature map is convolved by deep convolutional separable units and shallow convolutional separable units connected in a cascaded manner, and a convolutional feature map is output: , Where (k, m) represents the two-dimensional coordinates of the low-dimensional image patch feature map. This represents the convolutional feature map. This represents the convolution kernel of the depthwise separable unit. This represents the convolution kernel of the shallow convolutional separable unit, where (i1,j1) represents the convolution kernel. In the i1th row and j1st column, (u, v) represents the convolution kernel. The u-th row and v-th column.
5. The handwritten Chinese character recognition and error correction method based on artificial intelligence according to claim 4, characterized in that, The channel fusion module further includes a one-dimensional channel unit, a two-dimensional channel unit, a second regression fusion unit, and the pixel-by-pixel convolution unit; the one-dimensional channel unit, the two-dimensional channel unit, the second regression fusion unit, and the position attention module are connected in sequence, and the channel fusion module is constructed as follows: Receive the convolutional feature map; The convolutional feature map is processed sequentially by the pixel-wise convolutional unit, the one-dimensional channel unit, and the two-dimensional channel unit to capture the one-dimensional and two-dimensional channel information of the convolutional feature map, thereby obtaining a one-dimensional feature map and a two-dimensional feature map, respectively. , in, This refers to the one-dimensional feature map or the two-dimensional feature map. denoted as Softmax activation function, GN represents batch normalization, and GWConv() represents a pixel-wise convolutional unit, wherein the kernel size of the pixel-wise convolutional unit is 2×2×A×A; Based on the preset channel descent coefficient, the one-dimensional feature map and the two-dimensional feature map are divided into multiple channel groups in order of gradually decreasing channel number. Self-attention is calculated for the features in each channel group to obtain the group attention features. Based on the deformable convolutional network, the local spatial dependency relationship of the group attention features is calculated to obtain the spatial dependency weight. The spatial dependency weight is weighted with the group attention features to obtain the first weighted feature. The mutual information between adjacent one-dimensional feature maps and two-dimensional feature maps is calculated to obtain the corresponding feature correlation coefficients. Based on the feature correlation coefficients, the first weighted feature is downsampled and passed to obtain a downsampled feature. The first weighted feature is then upsampled and passed to obtain an upsampled feature. The downsampled feature and the upsampled feature are combined to obtain a second weighted feature. Multi-head attention operations are performed on the second weighted feature in both the channel dimension and the spatial dimension to obtain channel attention weights and spatial attention weights. The channel attention weights are weighted with the channel information of the second weighted feature to obtain channel information weighted features, and the spatial attention weights are weighted with the spatial information of the second weighted feature to obtain spatial information weighted features. The second regression fusion unit combines the weighted features of the channel information with the weighted features of the spatial information to generate fused features.
6. The handwritten Chinese character recognition and error correction method based on artificial intelligence according to claim 5, characterized in that, The positional attention module includes pixel-wise convolutional units, and the specific construction method of the positional attention module includes: The location attention module performs global feature extraction on the fused features to obtain an attention-based global feature map with the same size as the low-dimensional image patch feature map. , in, The attention feature is represented by GAtt(), which represents the positional attention module.
7. The handwritten Chinese character recognition and error correction method based on artificial intelligence according to claim 6, characterized in that, The pyramid classification layer includes a pyramid classification sublayer and a pyramid regression sublayer; The specific methods for constructing the pyramid classification layer include: The global attention feature map is classified by a pyramid classification sub-layer, and various non-standard types of Chinese characters are output. The pyramid classification sub-layer output results are regressed and confirmed by the pyramid regression sub-layer, and the final non-standard types of Chinese characters are output. The non-standard types include at least stroke breakage, component misalignment, loose structure, disproportion, center of gravity shift and stroke order error. The original images of various non-standard Chinese character input data samples are found. The latent space feature vectors of the original images are extracted by the encoder module of the variational autoencoder. Random perturbation and interpolation operations are applied to the feature vectors in the latent space to generate diverse non-standard feature vectors. After the non-standard feature vectors are reconstructed by the decoder, Chinese character image samples with non-standard stroke features and overall structural features are output. The pyramid classification sublayer performs 5 convolutions on the global attention feature map to enhance the image features of non-standard Chinese characters, and the convolution kernel size is 2×2 with a number of 32.
8. The handwritten Chinese character recognition and error correction method based on artificial intelligence according to claim 1, characterized in that, In step S6, the handwritten Chinese character image to be recognized is segmented into regions, and the region to be recognized containing the Chinese characters is extracted, including: Adaptive binarization is performed on the image of the handwritten Chinese character to be recognized to generate a binary image; An edge detection algorithm is used to locate the contours of Chinese characters in a binary image, and the minimum bounding rectangle of each contour is calculated. Based on the spatial distribution characteristics of Chinese characters, the outlines of Chinese characters in the central region of the image are selected; Geometric correction is performed on the selected Chinese character outline area to obtain a standardized region to be recognized.
9. The handwritten Chinese character recognition and error correction method based on artificial intelligence according to claim 8, characterized in that, The geometric correction specifically includes: Establish the mapping relationship between the corner points of the minimum bounding rectangle and the corner points of the standard rectangle; Calculate the image transformation matrix based on the perspective transformation algorithm; Apply an image transformation matrix to perform an affine transformation on the Chinese character region; Output the corrected, standardized Chinese character images.
10. The handwritten Chinese character recognition and error correction method based on artificial intelligence according to claim 7, characterized in that, The specific implementation process of S8 is as follows: The non-standard stroke features and overall structural features of the handwritten Chinese character image to be processed are input into the first graph convolutional network layer of the preset Chinese character correction model to obtain the first non-standard feature vector; wherein, the first non-standard feature vector represents the topological relationship between the non-standard stroke features and overall structural features of the handwritten Chinese character image to be processed; the preset Chinese character correction model is a correction model constructed based on a multi-layer graph convolutional network; The local deformation features and global layout features of the handwritten Chinese character image to be processed are input into the second graph convolutional network layer of the preset Chinese character correction model to obtain the second non-normal feature vector; wherein, the second non-normal feature vector represents the spatial constraint relationship between the local deformation features and global layout features of the handwritten Chinese character image to be processed. The first non-normalized feature vector and the second non-normalized feature vector are concatenated to obtain the concatenated corrected feature vector. Based on the preset Chinese character correction model, the splicing correction feature vector is analyzed to obtain the second non-standard type of the handwritten Chinese character; The preset standard Chinese character correction library stores correction templates corresponding to various second non-standard types. The matching degree between the second non-standard type of the handwritten Chinese character to be processed and the non-standard type in the standard Chinese character correction library is calculated by the feature space distance metric algorithm. A matching degree threshold is set. If the matching degree is greater than the threshold, the corresponding standard correction template is called to reconstruct the non-standard Chinese character and generate a standard Chinese character. If the matching degree is less than or equal to the threshold, the correction template with the highest matching degree is selected for correction.
Citation Information
Patent Citations
handwritten Chinese recognition method based on deep learning
CN109800763A
Handwriting model training method, handwriting recognition method, device and apparatus and medium
CN108985442A
Irregular character recognition device and method based on deep learning
CN110427938A
Image classification method
CN115222998A
Handwritten Chinese character recognition and correction method
CN119323794A