Chinese character recognition method based on stroke tree representation
Through the Chinese character recognition method based on stroke tree representation, the sequence representation of the stroke tree and the deep learning model extract feature and prediction structure information is solved, and the problem of insufficient robustness of Chinese character recognition in the prior art is achieved, and a higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202111578044.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-22
AI Technical Summary
The existing Chinese character recognition method based on deep learning is not robust enough to deal with occlusion, fuzzy and zero sample problems, especially the stroke-based recognition method has poor processing capabilities for fuzzy samples.
The Chinese character recognition method based on stroke tree representation is adopted, and the Chinese characters are disassembled into sequence representations of stroke tree, and the feature and prediction structure information are extracted using residual convolution neural network and Transformer decoder, and the final prediction result is calculated in combination with weighted editing distance.
The robustness of Chinese character recognition is improved, especially when dealing with occlusion, blur and zero sample problems, and the recognition accuracy is significantly improved.
Smart Images

Figure CN116363670B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of Chinese character recognition, and in particular relates to a Chinese character recognition method based on stroke tree representation. Background Art
[0002] In recent years, deep learning technology has developed rapidly and has been applied to various practical scenarios. In the field of optical character recognition, Chinese character recognition methods based on deep learning are widely adopted by the industry. Chinese character recognition methods based on deep learning can be divided into three categories: whole-word-based recognition methods, radical-based recognition methods, and stroke-based recognition methods. Most whole-word-based recognition methods use convolutional neural networks (CNNs) to extract character features, and combine them with traditional manual features to feed into a character classifier to obtain character predictions. Radical-based recognition methods regard Chinese characters as a structural combination of several radicals, that is, characters are decomposed into sequences of radicals and structures. Whole-word-based and radical-based recognition methods rely more on the overall features of characters, so these two types of recognition methods are not robust to occluded and blurred samples, and cannot handle character zero-sample or radical zero-sample problems (characters or radicals in the test set have never appeared in the training set). Although stroke-based recognition methods can solve the zero-sample problem, the one-to-many relationship between strokes and characters brings additional difficulties to recognition, and stroke-based recognition methods are less robust to blurred character samples. Summary of the invention
[0003] In order to solve the above problems, a Chinese character recognition method based on stroke tree representation with stronger robustness is provided. The present invention adopts the following technical solutions:
[0004] The present invention provides a Chinese character recognition method based on stroke tree representation, which is used to predict a Chinese character to be recognized according to a plurality of label Chinese characters in a dictionary, and is characterized by comprising:
[0005] Step S1, decomposing the Chinese characters in the label according to a predetermined decomposition rule to obtain a sequence representation of a stroke tree corresponding to the Chinese characters in the label, recorded as a label sequence;
[0006] Step S2, assigning a corresponding weight to each element in the tag sequence according to the level of the element in the corresponding stroke tree;
[0007] Step S3, extracting the image features of the Chinese characters to be recognized by using a residual convolutional neural network;
[0008] Step S4, predicting the structure information and position information of each radical of the Chinese character to be recognized by using a Transformer-based radical decoder;
[0009] Step S5, multiplying the image feature and each corresponding pixel in the radical position information to obtain each radical feature information of the Chinese character to be recognized;
[0010] Step S6, using a Transformer-based stroke decoder to decode the radical feature information to obtain the radical stroke information;
[0011] Step S7, combining the structural information and the information of each radical stroke into a sequence representation of the stroke tree of the Chinese character to be recognized, recorded as a prediction sequence, and assigning a corresponding weight to each element of the prediction sequence according to the level of the element in the corresponding stroke tree;
[0012] Step S8, calculating the weighted edit distance between the prediction sequence and each of the label sequences, and taking the label Chinese character with the smallest weighted edit distance as the final prediction result.
[0013] The Chinese character recognition method based on stroke tree representation provided by the present invention may also have such a technical feature, wherein the label sequence is represented as In the formula, G r Represents structural information, G s represents stroke sequence information, and T represents the number of elements in the label sequence.
[0014] The Chinese character recognition method based on stroke tree representation provided by the present invention may also have such a technical feature, wherein in step S2 and step S7, the weight distribution is performed according to the following formula: W = α k , where W is the weight, k represents the level of the element in the tree structure of the stroke tree, and α represents the attenuation coefficient of the weight allocation.
[0015] The Chinese character recognition method based on stroke tree representation provided by the present invention may also have such a technical feature, wherein in step S8, the formula for calculating the weighted edit distance is:
[0016]
[0017] In the formula, ED(·) represents the function of calculating the edit distance, L(·) represents the function of calculating the string length, and M i represents the i-th element in the prediction sequence M, G j represents the jth element in the tag sequence GT, W i M represents the weight of the i-th element in M, W j gt Represents the weight of the j-th element in GT.
[0018] The Chinese character recognition method based on stroke tree representation provided by the present invention may also have such a technical feature, wherein the residual convolutional neural network is ResNet34.
[0019] Function and Effect of the Invention
[0020] According to the Chinese character recognition method based on stroke tree representation of the present invention, based on the characteristic that Chinese characters can be decomposed hierarchically, the structural information of Chinese characters is fully considered and Chinese characters are represented in the form of stroke trees, thereby further alleviating the one-to-many problem of sequence representation to Chinese characters on the basis of the prior art; and because the method of the present invention gives a higher weight to the radical structure when calculating the distance, the candidate predicted Chinese characters are closer to the label Chinese characters as a whole, thereby improving the prediction accuracy. Furthermore, because the representation of the stroke tree combines the representation advantages of the radical level and the stroke level, the neural network model has stronger robustness for samples with occlusion and blur. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flow chart of a Chinese character recognition method based on stroke tree representation in an embodiment of the present invention;
[0022] Figure 2 is a schematic diagram of a disassembly of a Chinese character stroke tree representation in an embodiment of the present invention;
[0023] Figure 3 is a schematic diagram of weight distribution represented by a stroke tree in an embodiment of the present invention;
[0024] Figure 4 is a schematic diagram of the structure of a residual convolutional neural network in an embodiment of the present invention;
[0025] Figure 5 is a schematic diagram of structural information and radical position information of Chinese characters in an embodiment of the present invention;
[0026] Figure 6 is a schematic diagram of the principle of calculating radical feature information in an embodiment of the present invention;
[0027] Figure 7 is a schematic diagram of the principle of decoding radical stroke information in an embodiment of the present invention;
[0028] Figure 8 is a schematic diagram of the principle of combining structural information and stroke information in an embodiment of the present invention;
[0029] Fig. 9 4 is a flow chart of calculating the weighted edit distance in an embodiment of the present invention. DETAILED DESCRIPTION
[0030] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the Chinese character recognition method based on stroke tree representation of the present invention is specifically described below in conjunction with embodiments and drawings.
[0031] <Example>
[0032] Figure 1 It is a flow chart of a Chinese character recognition method based on stroke tree representation in an embodiment of the present invention.
[0033] like Figure 1 As shown, in this embodiment, all the labeled Chinese characters in the dictionary are first decomposed into stroke tree representations, and then based on the labeled Chinese characters in the dictionary, the trained neural network model is used to predict the character image to be recognized. The Chinese character recognition method based on stroke tree representation in this embodiment specifically includes the following steps:
[0034] Step S1, decomposing the Chinese character to be recognized according to a predetermined decomposition rule to obtain a sequence representation of a stroke tree corresponding to the Chinese character (hereinafter referred to as a stroke tree sequence).
[0035] In this embodiment, the dictionary (also called the Chinese alphabet) contains a plurality of labeled Chinese characters, and the Chinese character to be recognized is predicted based on the plurality of labeled Chinese characters in the dictionary, that is, the labeled Chinese character with the highest similarity to the Chinese character to be recognized is found in the dictionary as the prediction of the Chinese character to be recognized. The Chinese character to be recognized is an image of the Chinese character.
[0036] According to the characteristic that Chinese characters can be decomposed hierarchically (Chinese characters can be decomposed into radicals, and radicals can be decomposed into strokes), all the tagged Chinese characters in the dictionary are decomposed according to the predetermined decomposition rules. In this embodiment, the predetermined decomposition rules are specifically: a Chinese character can be decomposed into a tree with the radical structure of the whole character as the root node and the decomposed radicals as leaf nodes; if the decomposed radical can be further decomposed according to the standard, then the leaf node where the radical is located is expanded into a subtree with the subdivision structure of the radical as the root node and the corresponding subdivision radical as the leaf node; and so on, until each radical is decomposed to be indivisible, and the hierarchical expanded radical tree of the Chinese character is obtained; further, the stroke information of the radical corresponding to the leaf node in the above radical tree is used as the element of the leaf node to obtain the stroke tree representation of the Chinese character.
[0037] The stroke tree of each Chinese character can be represented as a unique sequence corresponding to it, that is, the pre-order traversal sequence of the stroke tree. Specifically: the sequence corresponding to the stroke tree of the label Chinese character in the dictionary is represented as Among them G r Represents structural information, G s Represents the stroke sequence information, and T represents the number of elements in the stroke tree representation.
[0038] For the convenience of description, hereinafter, the sequence representation G of the stroke tree of the Chinese characters in the label will be abbreviated as the label sequence G.
[0039] Figure 2 It is a schematic diagram of the decomposition of the Chinese character stroke tree representation in the embodiment of the present invention.
[0040] As Figure 2 shown, in this embodiment, the Chinese character "You" in the label is decomposed into a tree structure representation of a radical structure and a radical stroke sequence, where the non-leaf nodes represent the structure of the character or a certain radical of the character, and the leaf nodes represent the stroke sequence of a certain radical in the character.
[0041] Step S2: For each element G in the label sequence G i , assign a corresponding weight W according to the level of the element G i in the corresponding stroke tree.
[0042] In this embodiment, a certain weight is assigned to each element G in the label sequence G i (1 ≤ i ≤ T), and the weight assignment method is as follows: W = α k , where k represents the level of the element in the tree structure, and α represents the attenuation coefficient of weight assignment. Thus, the label sequence G of the Chinese characters in the label has a corresponding weight W gt .
[0043] Figure 3 It is a schematic diagram of weight assignment of the stroke tree representation in the embodiment of the present invention.
[0044] As Figure 3 shown, a certain weight is assigned to each node of the tree structure. And the higher the layer where the node is located, the higher the weight assigned to the node. As Figure 3 shown, the weight of the root node is empirically selected as 0.5, and the weight of the lower layer is 1 / 2 of the upper layer.
[0045] Step S3: Use the residual convolutional neural network to extract the image feature F of the Chinese character to be recognized.
[0046] Figure 4 It is a schematic diagram of the residual convolutional neural network structure in the embodiment of the present invention.
[0047] As Figure 4 shown, in this embodiment, the residual convolutional neural network "ResNet34" is used as a feature extractor to extract the image feature F of the Chinese character to be recognized (i.e., the character image). And as Figure 4 known, the original residual convolutional neural network "ResNet34" downsamples the input image to 1 / 32 of the original size. In order to retain more character information, only the first pooling operation is used in the feature extraction process of the present invention.
[0048] Step S4: Use the radical decoder based on Transformer to predict the structural information S of the Chinese character to be recognized r and the position information P of each radical.
[0049] Among them, the radical structure information can be expressed as The radical structure information P adopts Figure 2 the structural categories shown for supervision, and the radical position information P uses the relative position of the radical in the Chinese character to be recognized (that is, the path category from the root node of the corresponding stroke tree of the Chinese character to each radical leaf node) as supervision.
[0050] In this embodiment, each radical position information P is the attention map generated by the radical decoder based on Transformer when decoding the radical relative position category.
[0051] Figure 5 is a schematic diagram of the structural information and radical position information of the Chinese character to be recognized in the embodiment of the present invention.
[0052] As Figure 5 shown, the extracted image feature F is sent into the radical decoder based on Transformer to obtain the structural information S of the Chinese character to be recognized r , as Figure 5 shown in the circular area in, and obtain the radical position information P of the Chinese character to be recognized, that is Figure 5 the part marked in gray in the Chinese character in.
[0053] Step S5: Multiply each corresponding pixel in the extracted image feature F and each radical position information P to obtain the radical feature information F of the Chinese character to be recognized radical .
[0054] Figure 6 is a schematic diagram of the principle of calculating the radical feature information in the embodiment of the present invention.
[0055] As Figure 6 shown, in this embodiment, the feature of the Chinese character "you" is multiplied by the position information of the radical "ren" pixel by pixel to obtain the feature information of only the radical "ren".
[0056] Step S6: Use the stroke decoder based on Transformer to decode each radical feature information F radical to obtain the stroke information of each radical where n is the number of leaf nodes in the tree representation of the character.
[0057] Similarly, the feature information of the radical structure is also decoded by this decoder.
[0058] Among them, if the node is a radical structure, the stroke combination sequences of all leaf nodes in the subtree with this radical structure as the root node in the "radical-stroke decoder" are used as supervision information; if the node is a radical node, the stroke sequence corresponding to this radical is used as supervision information. Since the supervision of the radical structure is only to make the attention area cover all radical areas of the subtree with it as the root node as much as possible and is not used for subsequent calculations, the stroke sequence prediction of the radical structure is not reflected in the stroke information.
[0059] Figure 7 is a schematic diagram of the principle of decoding radical stroke information in an embodiment of the present invention.
[0060] As Figure 7 shown, the feature information of each radical or radical structure of the Chinese character to be recognized is respectively sent into the stroke decoder based on Transformer to obtain the stroke information corresponding to each radical or each radical structure. For example: the radical "亻" can obtain the stroke predictions of "丿" and "丨" through the stroke decoder.
[0061] Step S7, merge the structure information S r and the stroke information S s of each radical into the sequence representation M = {M1, M2,..., M m+n} of the predicted stroke tree of the Chinese character to be recognized, and assign weights to the elements in the predicted sequence M in the same weight assignment method as above.
[0062] For the convenience of narration, the sequence representation M of the predicted stroke tree of the Chinese character to be recognized is also simply recorded as the predicted sequence M.
[0063] It can be seen from the above steps S1 and S2 that the predicted sequence M also has the corresponding weight W M .
[0064] Figure 8 is a schematic diagram of the principle of merging structure information and stroke information in an embodiment of the present invention.
[0065] As Figure 8 shown, fuse the prediction of the structure information and the prediction of the stroke information to obtain the predicted sequence M of the Chinese character to be recognized, where Figure 8 the dotted arrow in indicates the source of the structure information or the stroke information in the sequence representation of the predicted stroke tree.
[0066] Step S8, calculate the weighted edit distance between the predicted sequence M and each label sequence G in the dictionary, and use the Chinese character of the label with the smallest weighted edit distance as the final prediction result.
[0067] In this embodiment, the state transition equation of the "weighted edit distance" is as follows:
[0068]
[0069] In the formula, ED(·) represents the edit distance function, L(·) represents the string length function, and D[i,j] represents the weighted edit distance between 1 to i elements in the prediction sequence M and 1 to j elements in the label sequence G in the dictionary. Therefore, the distance between the prediction sequence M and the label sequence G in the dictionary is D[m+n,T].
[0070] Fig. 9 4 is a flow chart of calculating the weighted edit distance in an embodiment of the present invention.
[0071] like Fig. 9 As shown, in this embodiment, the process of calculating the above-mentioned "weighted edit distance" specifically includes the following steps:
[0072] Initialize D[0,i](0≤i≤L(G)) and D[i,0](0≤i≤L(M)).
[0073] That is, the weights of the first i elements of the label sequence G are accumulated as the distance between the first i elements of the label sequence G and the first zero elements of the prediction sequence M, or the weights of the first i elements of the prediction sequence M are accumulated as the distance between the first i elements of the prediction sequence M and the first zero elements of G.
[0074] D[1,1] to D[3,3] are calculated according to the above formula (1).
[0075] That is, D[i,j] takes the smallest final distance value among the following three cases:
[0076] (1) The sum of D[i-1,j] and the weight of the i-th element in M;
[0077] (2) The sum of D[i,j-1] and the weight of the j-th element in G;
[0078] (3)D[i-1,j-1] and element M i and element G j The sum of the weighted edit distances of .
[0079] Among them, the element M i and element G j The calculation formula of the weighted distance is:
[0080]
[0081] For example, Fig. 9 As shown, when calculating D[2,2], it is known that D[1,2]=0.25, D[2,1]=0.25, D[1,1]=0, so D[2,2] is selected D[2,1]+Wj GT = 0.5 and The smallest distance is taken as the weighted edit distance between the first two elements in the prediction sequence M and the first two elements in each label sequence G in the dictionary.
[0082] Finally, D[m+n,T] is the "weighted edit distance" between the two stroke tree sequences. The Chinese character with the smallest weighted edit distance in the dictionary is taken as the final prediction result.
[0083] The recognition of character images can be completed through the above process. Compared with the previous recognition methods based on whole words, radicals and strokes, the stroke tree representation used in this embodiment is more robust to occlusion, blur and zero-sample problems. The experimental results are shown in Table 1. Taking the Chinese character recognition accuracy as the evaluation indicator, experiments were conducted on occlusion, blur, and character and radical zero-sample data sets. The results show that the representation based on the stroke tree is more robust to character samples such as occlusion and blur, and the recognition accuracy of the method in this embodiment is significantly higher than that of the recognition method in the prior art.
[0084] Table 1 Comparison of recognition accuracy of different recognition methods
[0085]
[0086] Example Function and Effect
[0087] According to the Chinese character recognition method based on stroke tree representation provided by this embodiment, based on the characteristic that Chinese characters can be decomposed hierarchically, the structural information of Chinese characters is fully considered and Chinese characters are represented in the form of stroke trees, thereby further alleviating the one-to-many problem of sequence representation to Chinese characters on the basis of the prior art; and because the method of this embodiment gives a higher weight to the radical structure when calculating the distance, the candidate predicted characters are closer to the label Chinese characters as a whole, thereby improving the prediction accuracy. Furthermore, because the representation of the stroke tree combines the representation advantages of the radical level and the stroke level, the neural network model has stronger robustness for samples with occlusion and blur.
[0088] The above embodiments are only used to illustrate specific implementation modes of the present invention, and the present invention is not limited to the description scope of the above embodiments.
Claims
1. A Chinese character recognition method based on stroke tree representation, used to predict a Chinese character to be recognized based on a plurality of label Chinese characters in a dictionary, characterized in that: include: Step S1, decomposing the Chinese characters in the label according to a predetermined decomposition rule to obtain a sequence representation of a stroke tree corresponding to the Chinese characters in the label, recorded as a label sequence; Step S2, assigning a corresponding weight to each element in the tag sequence according to the level of the element in the corresponding stroke tree; Step S3, extracting the image features of the Chinese characters to be recognized by using a residual convolutional neural network; Step S4, predicting the structure information and position information of each radical of the Chinese character to be recognized by using a Transformer-based radical decoder; Step S5, multiplying the image feature and each corresponding pixel in the radical position information to obtain each radical feature information of the Chinese character to be recognized; Step S6, using a Transformer-based stroke decoder to decode the radical feature information to obtain the radical stroke information; Step S7, combining the structural information and the information of each radical stroke into a sequence representation of the stroke tree of the Chinese character to be recognized, recorded as a prediction sequence, and assigning a corresponding weight to each element of the prediction sequence according to the level of the element in the corresponding stroke tree; Step S8, calculating the weighted edit distance between the prediction sequence and each of the label sequences, and taking the label Chinese character with the smallest weighted edit distance as the final prediction result.
2. The Chinese character recognition method based on stroke tree representation according to claim 1, characterized in that: in, The tag sequence is represented as In the formula, G r Represents structural information, G s represents stroke sequence information, and T represents the number of elements in the label sequence.
3. The Chinese character recognition method based on stroke tree representation according to claim 1, characterized in that: in, In step S2 and step S7, the weights are allocated according to the following formula: W=α k In the formula, W is the weight, k represents the level of the element in the tree structure of the stroke tree, and α represents the attenuation coefficient of weight allocation.
4. The Chinese character recognition method based on stroke tree representation according to claim 1, characterized in that: in, In step S8, the formula for calculating the weighted edit distance is: In the formula, ED(·) represents the function of calculating the edit distance, L(·) represents the function of calculating the string length, and M i represents the i-th element in the prediction sequence M, G j represents the jth element in the tag sequence GT, W i M represents the weight of the i-th element in M, Represents the weight of the j-th element in GT.
5. The Chinese character recognition method based on stroke tree representation according to claim 1, characterized in that: in, The residual convolutional neural network is ResNet34.
Citation Information
Patent Citations
Chinese word vector modeling method
CN109992783A
Free-type Chinese-character enter method using keypad and its device
CN1243982A