A small sample ancient character recognition method based on structural features and evolution patterns
By adopting a neural network model of structural features and evolutionary patterns in the small sample ancient text recognition method, the problems of overfitting and feature extraction in ancient text recognition are solved, and high recognition rate and short-term learning are achieved, which is suitable for ancient text recognition tasks.
Patent Information
- Application Number
- CN202411372569.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-09-29
AI Technical Summary
The existing small sample learning methods have problems such as overfitting, incomplete feature extraction and low recognition accuracy in ancient characters recognition tasks, especially the failure to effectively utilize the structural features and evolution patterns of Chinese characters.
A small sample ancient text recognition method based on structural features and evolution mode is adopted. Through the feature extraction module, character classification module and comparison learning module in the neural network model, the structural features and evolution mode of ancient text are extracted and utilized to perform data augmentation and comparison learning to improve the recognition accuracy.
On the data sets CCED and Oracle-241, the recognition rate is better than the existing related methods, the data enhancement effect is significant, the training time is short, and the feature extraction is innovative, which can effectively solve the problem of scarcity of ancient text data.
Smart Images

Figure CN119206739B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of Chinese character recognition, and in particular relates to a small sample ancient Chinese character recognition method based on structural features and evolution patterns. Background Art
[0002] Ancient character recognition is a challenging task. Traditional computer vision tasks usually require a large amount of training data, which is difficult to meet in ancient character recognition tasks. For newly unearthed artifacts or documents, the ancient characters in them often lack sufficient references, making it difficult to obtain a large amount of annotated data. In addition, for some scenes of ancient characters, the amount of data available is very small. Small sample learning can effectively train a model with good generalization performance using only a small amount of data, which can solve the problem well.
[0003] Existing small sample learning methods can be roughly divided into three categories: optimization-based methods, model-based methods, and metric-based methods. Optimization-based methods seek good initialization parameters so that the model can quickly adapt to new tasks. When faced with complex ancient character recognition tasks, the limited number of samples may lead to overfitting. Model-based methods aim to quickly update parameters on a small number of samples by designing the model structure and directly establishing a mapping function between input values and predicted values. However, due to the complex structure and variable fonts of ancient characters, it is difficult to design a model that can effectively capture these features. Metric-based methods learn a metric space in which similar data points are close to each other and dissimilar data points are far away. Existing metric-based small sample methods are usually designed for general datasets such as miniImageNet, which are very different from ancient character datasets. In particular, ancient characters have special structural features, with a total of 13 structures, and special evolution patterns, which are not in general datasets such as miniImageNet. Therefore, the existing metric-based small sample methods do not effectively utilize the structural features and evolution patterns of Chinese characters, resulting in poor performance in Chinese character recognition tasks. Summary of the invention
[0004] In view of the above-mentioned deficiencies in the prior art, the small sample ancient character recognition method based on structural features and evolution patterns provided by the present invention solves the problems of overfitting in the field of ancient character recognition, incomplete feature extraction and low recognition accuracy due to the small amount of data in the existing small sample learning.
[0005] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: a small sample ancient character recognition method based on structural features and evolution patterns, comprising the following steps:
[0006] S1, collect ancient character images, build a data set, and perform data enhancement processing on it;
[0007] S2, use the data-enhanced data set to train the neural network and obtain a small sample ancient character recognition model;
[0008] The small sample ancient character recognition model includes a feature extraction module, a character classification module and a contrastive learning module; wherein the feature extraction module is used to extract the feature map of the character image in the data set; the character classification module is used to perform self-attention enhancement on the feature map, and perform an average pooling operation on the enhanced feature map containing local features to obtain global features, and then evaluate the similarity between the character images based on the global features and the local features to determine the character image category; the contrastive learning module is used to perform contrastive learning on the global features of the character image after data enhancement according to the class label to obtain discriminative features;
[0009] S3. Input the target ancient character image into a small sample ancient character recognition model, and then output the ancient character recognition result.
[0010] Furthermore, in the step S1, the data enhancement processing includes color jittering, slight rotation, slight shearing and slight scaling.
[0011] Furthermore, in step S2, in the small sample ancient character recognition model, the method by which the feature extraction module extracts global features and local features is specifically:
[0012] S2-A1, extracting a feature map containing local features of the input character image through a convolutional neural network;
[0013] S2-A2, divide the feature map containing local features into uniform grids to determine the corresponding local spatial features;
[0014] S2-A3, use the global average pooling operation to aggregate local spatial features to obtain global features.
[0015] Furthermore, in step S2, in the small sample ancient character recognition model, the method by which the character classification module evaluates the similarity between character images is specifically as follows:
[0016] S2-B1, perform self-attention enhancement on each type of feature map in the support set to obtain a self-attention enhanced feature map;
[0017] S2-B2, perform self-attention enhancement on each query image in the query set to obtain an enhanced query feature map;
[0018] S2-B3, calculate the global similarity and local similarity of the self-attention enhanced feature map and the enhanced query feature map respectively;
[0019] S2-B4, based on the global similarity and the local similarity, determine the similarity between the query image and various character images in the support set, and then determine the character category of the query image;
[0020] The data-enhanced dataset is divided into a support set and a query set according to the ancient character categories identified by the classification task.
[0021] Furthermore, the step S2-B1 includes the following sub-steps:
[0022] B11. Take the average value of each class feature map in the support set to obtain the class feature map;
[0023] B12. Use the self-attention mechanism to process the class feature map to obtain the class attention feature map;
[0024] B13. Combine the class feature map and the class attention feature map to obtain the self-attention enhanced feature map.
[0025] Furthermore, in the step S2-B3, the global similarity sim between the self-attention enhanced feature map and the enhanced query feature map g for:
[0026]
[0027] The local similarity sim at each position in the self-attention enhanced feature map and the enhanced query feature map pos for:
[0028] sim pos =α·cos(A ij ,B ij )+β·cos(A ij ,R ij )
[0029] In the formula, f i It means that the features of each class in the support set are averaged, and a prototype corresponding to each class is obtained. surface
[0030] The features of the i-th query image are enhanced by the self-attention mechanism and then globally pooled. represents the class enhancement feature for the support image of category cj, CosineSimilarity(·) represents the cosine similarity, and sim(·) represents the similarity measure;
[0031] A ij and B ij denote the enhanced query feature map corresponding to the i-th query image and the c-th j Self-attention enhanced feature map, R ij Indicates Aij The corresponding boundary region features.
[0032] Furthermore, in step S2-B4, the i-th query image and the c-th image in the support set j The similarity sim of the character image is expressed as:
[0033] sim=λ·sim g +sim l
[0034] In the formula, sim g represents the global similarity, λ represents the scaling weight of the global similarity, sim l Represents local similarity.
[0035] Furthermore, during the training of the small sample ancient character recognition model, the loss function L of the character classification module is CE for:
[0036]
[0037] Where N Q represents the number of query images, C represents the number of classes, and y qi represents the true label value of the i-th class of the q-th image, p qi It represents the probability that the qth image belongs to the i-th category predicted by the classification recognition module.
[0038] Furthermore, in the training process of the small sample ancient character recognition model, the loss function of the contrastive learning module is:
[0039]
[0040] In the formula, 2N yi Indicates that the label is y i Number of images, 1 cond ∈{0,1} means that the value is 1 when the condition is met, i≠j means taking different samples, y i =y j Indicates that the labels of the two samples are the same, N indicates the number of samples in the support set, and l ij represents the loss between samples i and j, f i represents the eigenvector of the i-th sample, and τ represents the scalar temperature parameter.
[0041] The beneficial effects of the present invention are:
[0042] (1) The small sample recognition method for ancient characters proposed in this invention has better recognition rates than existing related methods on the CCED and Oracle-241 datasets;
[0043] (2) The data enhancement effect used in the present invention has obvious performance advantages in the field of ancient character recognition compared with the existing conventional data enhancement methods;
[0044] (3) The present invention is a short-time learning method suitable for ancient character recognition tasks, and the training time is relatively short;
[0045] (4) The present invention takes into account both global and local features, as well as the structural characteristics and evolution laws of Chinese characters in the recognition of ancient Chinese characters, and is innovative in extracting features from ancient characters;
[0046] (5) The present invention uses a data set with a smaller amount of data, achieves excellent performance, and effectively solves the problem of scarcity of ancient character data. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 This is a flow chart of the small sample ancient character recognition method based on structural features and evolution patterns in the present invention.
[0048] Figure 2 This is a small sample ancient character recognition framework diagram in the present invention. DETAILED DESCRIPTION
[0049] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0050] The embodiment of the present invention provides a small sample ancient character recognition method based on structural features and evolution patterns, such as Figure 1-2 As shown, the following steps are included:
[0051] S1, collect ancient character images, build a data set, and perform data enhancement processing on it;
[0052] S2, use the data-enhanced data set to train the neural network and obtain a small sample ancient character recognition model;
[0053] The small sample ancient character recognition model includes a feature extraction module, a character classification module and a contrastive learning module; wherein the feature extraction module is used to extract the feature map of the character image in the data set; the character classification module is used to perform self-attention enhancement on the feature map, and perform an average pooling operation on the enhanced feature map containing local features to obtain global features, and then evaluate the similarity between the character images based on the global features and the local features to determine the character image category; the contrastive learning module is used to perform contrastive learning on the global features of the character image after data enhancement according to the class label to obtain discriminative features;
[0054] S3. Input the target ancient character image into a small sample ancient character recognition model, and then output the ancient character recognition result.
[0055] In step S1 of the embodiment of the present invention, the data enhancement processing includes color jittering, slight rotation, slight shearing and slight scaling.
[0056] Specifically, in this embodiment, in order to alleviate the impact of data scarcity, there are N samples of data sets in this embodiment, and two different data sets are obtained through two data enhancement strategies, so the data set in this embodiment is expanded to 2N samples. From the evolution process of Chinese characters, each component of Chinese characters only changes in its original area, such as the original Chinese characters in the upper left corner of the area right a certain structure, the structure after evolution is still in the upper left corner. Therefore, some data enhancement operations may not be suitable for ancient characters. Conventional data enhancement operations include rotation, flipping, cropping and color jittering, etc. In the data enhancement strategy in this embodiment, we abandon the operations that may destroy the character structure such as flipping and cropping. On the contrary, the above-mentioned softer way is selected in this embodiment, such as color jittering, slight rotation, slight shearing and slight scaling operations, and in this embodiment, the intensity of these operations is limited in the parameter setting to ensure that the final output is reasonable for the characters. By combining these operations, the above-mentioned data enhancement strategy is tailored for ancient characters.
[0057] In step S2 of the embodiment of the present invention, in the small sample ancient character recognition model, the method of extracting global features and local features by the feature extraction module is specifically as follows:
[0058] S2-A1, extracting a feature map containing local features of the input character image through a convolutional neural network;
[0059] S2-A2, divide the feature map containing local features into uniform grids to determine the corresponding local spatial features;
[0060] S2-A3, use the global average pooling operation to aggregate local spatial features to obtain global features.
[0061] Specifically, in this embodiment, the feature extraction module in this embodiment is a CNNs network structure, wherein the feature map containing local spatial features extracted from the input character image is several continuous convolution blocks, and the formed local spatial features are expressed as Z=φ(x).
[0062] For the extracted feature map, if its size is too small, key local details may be ignored, and if its size is too large, it may not capture the complete features. Since Chinese characters have 13 structures, in order to extract local structural information (for example, left-right is two parts, left-middle-right is three parts), a segmentation method is adopted in this embodiment to split each character image into a 6×6 grid. This grid size enables us to capture fine-grained local details and retain the overall character structure while maintaining uniformity.
[0063] Finally, the local spatial features are aggregated by using the global pooling operation to obtain the global features.
[0064] In step S2 of the embodiment of the present invention, in the small sample ancient character recognition model, the method by which the character classification module evaluates the similarity between character images is specifically as follows:
[0065] S2-B1, perform self-attention enhancement on each type of feature map in the support set to obtain a self-attention enhanced feature map;
[0066] S2-B2, perform self-attention enhancement on each query image in the query set to obtain an enhanced query feature map;
[0067] S2-B3, calculate the global similarity and local similarity of the self-attention enhanced feature map and the enhanced query feature map respectively;
[0068] S2-B4, based on the global similarity and the local similarity, determine the similarity between the query image and various character images in the support set, and then determine the character category of the query image;
[0069] Among them, the data set after data enhancement is divided into a support set and a query set according to the ancient character categories recognized by the classification task; specifically, for the data set after data enhancement, N classes are randomly selected from all categories according to the classification task, and then K images are selected from each class as the support set, and the remaining certain number that have not been selected are used as the query set. In the learning process, the features of the images in the support set are first learned, and then it is determined which class each image in the query set corresponds to.
[0070] Step S2-B1 of this embodiment includes the following sub-steps:
[0071] B11. Take the average value of each class feature map in the support set to obtain the class feature map;
[0072] For example, for each class c in the support set i , take the average value Zc of the feature map of this type i , get the class feature map Zc i ;
[0073] B12. Use the self-attention mechanism to process the class feature map to obtain the class attention feature map;
[0074] For example, for the class feature map Zc i Use the self-attention mechanism to obtain the class attention feature map
[0075] B13. Combine the class feature map and the class attention feature map to obtain a self-attention enhanced feature map;
[0076] For example, Zc i and Combined, we get the self-attention enhanced feature map
[0077] In step S2-B3 of this embodiment, since a classifier that considers global and local similarities is required, in order to correctly classify, it is necessary to simultaneously calculate the global similarity and local similarity between the query image and each class.
[0078] For global similarity, the features of each class in the support set are averaged to obtain a prototype f for each class. i , and then calculate the global similarity sim of the self-attention enhanced feature map and the enhanced query feature map g for:
[0079]
[0080] For local similarity, the local similarity sim at each position in the self-attention enhanced feature map and the enhanced query feature map pos for:
[0081] sim pos =α·cos(A ij ,B ij )+β·cos(A ij ,R ij )
[0082] In the formula, f i It means that the features of each class in the support set are averaged, and a prototype corresponding to each class is obtained. It represents the features of the i-th query image enhanced by the self-attention mechanism and then globally pooled. represents the class enhancement feature for the support image of category cj, CosineSimilarity(·) represents the cosine similarity, and sim(·) represents the similarity measure;
[0083] sim pos Represents the local similarity of each position of the query image, A ij and B ij denote the enhanced query feature map corresponding to the i-th query image and the c-th j Self-attention enhanced feature map, R ij Indicates A ij Corresponding to the boundary area features, α and β represent the scaling weights. The center position naturally has a higher weight than the surrounding positions. In addition, the local similarity between the query image and each class is sim l is the average of the local similarities of all positions.
[0084] In this embodiment, in order to reduce the influence of the boundary position, the cth j The boundary area features of the self-attention-enhanced feature map are first filled with the same value as the adjacent boundary. ij The boundary of the region is then considered as a whole at the 3×3 position, and the global average pooling operation is used to obtain the boundary region features of the region.
[0085] In step S2-B4 of this embodiment, based on the above steps, the i-th query image and the c-th image in the support set are obtained. j The similarity sim of the character image is expressed as:
[0086] sim=λ·sim g +sim l
[0087] In the formula, sim g represents the global similarity, λ represents the scaling weight of the global similarity, sim l Represents local similarity.
[0088] In this embodiment, during the training of the small sample ancient character recognition model, the loss function L of the character classification module is CE for:
[0089]
[0090] Where N Q represents the number of query images, C represents the number of classes, and y qi represents the true label value of the i-th class of the q-th image, p qi It represents the probability that the qth image belongs to the i-th category predicted by the classification recognition module.
[0091] In this embodiment, based on the processing of the above-mentioned character classification module, global features and local features are used when classifying characters, and by considering structural features and inherent patterns in the evolution process, the similarity of characters can be evaluated more comprehensively.
[0092] In this embodiment, in the small sample ancient character recognition model, contrastive learning is used as an auxiliary task. In the above data enhancement part, 2N new images are obtained using the data enhancement strategy of N samples in the support set, and each enhanced image has the same class label as the original image. For any two enhanced images, if their class labels are the same, they are considered to be positive pairs; otherwise, they are considered to be negative pairs. Based on this, the loss function of the contrastive learning module is:
[0093]
[0094] In the formula, 2N yi Indicates that the label is y i The number of images, 1 cond ∈{0,1} means that the value is 1 when the condition is met, i≠j means taking different samples, y i =y j Indicates that the labels of the two samples are the same, N indicates the number of samples in the support set, and l ij represents the loss between samples i and j, f i represents the eigenvector of the i-th sample, and τ represents the scalar temperature parameter.
[0095] In this embodiment, the goal of contrastively learning the representation vectors of positive pairs is to be closer in the representation space, while the distance between negative pairs increases.
[0096] The present invention uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
[0097] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.
Claims
1. A small sample ancient character recognition method based on structural features and evolution patterns, characterized in that: The following steps are involved: S1, collect ancient character images, build a data set, and perform data enhancement processing on it; S2, use the data-enhanced data set to train the neural network and obtain a small sample ancient character recognition model; The small sample ancient character recognition model includes a feature extraction module, a character classification module and a contrast learning module; The feature extraction module is used to extract the feature map of the character image in the data set; the character classification module is used to perform self-attention enhancement on the feature map, and perform average pooling operation on the enhanced feature map containing local features to obtain global features, and then evaluate the similarity between the character images based on the global features and local features to determine the character image category; the contrastive learning module is used to perform contrastive learning on the global features of the character image after data enhancement according to the class label to obtain discriminative features; S3, inputting the target ancient character image into a small sample ancient character recognition model, and then outputting the ancient character recognition result; In S1, the data enhancement processing includes color jitter, slight rotation, slight shearing and slight scaling; In S2, in the small sample ancient character recognition model, the method for the character classification module to evaluate the similarity between character images specifically comprises the following steps: S2-B1, perform self-attention enhancement on each type of feature map in the support set to obtain a self-attention enhanced feature map; S2-B2, perform self-attention enhancement on each query image in the query set to obtain an enhanced query feature map; S2-B3, calculate the global similarity and local similarity of the self-attention enhanced feature map and the enhanced query feature map respectively; S2-B4, based on the global similarity and the local similarity, determine the similarity between the query image and various character images in the support set, and then determine the character category of the query image; The data augmented data set is divided into a support set and a query set according to the ancient character categories identified by the classification task; In S2-B3, the global similarity between the self-attention enhanced feature map and the enhanced query feature map for: The local similarity of each position in the self-attention enhanced feature map and the enhanced query feature map for: In the formula, It means that the features of each class in the support set are averaged, and a prototype corresponding to each class is obtained. Indicates i The features of the query image are enhanced using the self-attention mechanism and then globally pooled. It means that the target category obtained by the self-attention enhanced feature map is cj The class enhancement features of the support image, represents the cosine similarity, represents a similarity measure; and Respectively represent i The enhanced query feature map corresponding to the query image and the first cj Self-attention-enhanced feature map, express The corresponding boundary region features.
2. The small sample ancient character recognition method based on structural features and evolution patterns according to claim 1 is characterized in that: In S2, in the small sample ancient character recognition model, the method for extracting the global features and local features is specifically as follows: S2-A 1. Extracting a feature map containing local features of the input character image through a convolutional neural network; S2-A 2. Perform uniform grid segmentation on the feature map containing local features to determine the corresponding local spatial features; S2-A 3. Use the global average pooling operation to aggregate local spatial features to obtain global features.
3. The small sample ancient character recognition method based on structural features and evolution patterns according to claim 1 is characterized in that: The S2-B1 comprises the following sub-steps: B11. Take the average value of each class feature map in the support set to obtain the class feature map; B12. Use the self-attention mechanism to process the class feature map to obtain the class attention feature map; B13. Combine the class feature map and the class attention feature map to obtain the self-attention enhanced feature map.
4. The small sample ancient character recognition method based on structural features and evolution patterns according to claim 1 is characterized in that: In S2-B4, the i-th query image and the i-th image in the support set c j Similarity of character-like images It is expressed as: In the formula, represents the global similarity, represents the scaling weight of the global similarity, Represents local similarity.
5. The small sample ancient character recognition method based on structural features and evolution patterns according to claim 1 is characterized in that: During the training of the small sample ancient character recognition model, the loss function of the character classification module is for: In the formula, represents the number of query images, C represents the number of classes, Indicates q The image i The true label value of each class, Represents the first q The image belongs to i The probability of the class.
6. The small sample ancient character recognition method based on structural features and evolution patterns according to claim 1 is characterized in that: In the training process of the small sample ancient character recognition model, the loss function of the contrastive learning module is: In the formula, Indicates that the label is The number of images, Indicates that when the condition is met, the value is 1. Indicates taking different samples, Indicates that the labels of the two samples are the same. represents the number of samples in the support set, represents the loss between samples i and j, represents the feature vector of the i-th sample, Represents a scalar temperature parameter.
Citation Information
Patent Citations
Small sample remote sensing image scene classification method based on discrimination enhancement
CN116740565A
Cross-modal pedestrian re-identification method based on inter-modal common semantic learning
CN118711217A