An automatic scoring system and method for the writing standardization of handwritten Chinese character images

By combining templates and machine learning methods, dictionary and feature matching modules are constructed, and the problems of insufficient rule limitations and objectivity in the existing technology are solved, and the automation, accuracy and adaptability of handwritten Chinese character scores are realized.

CN119540966BActive Publication Date: 2025-07-22INNER MONGOLIA NORMAL UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411598425.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-07-22
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

The existing handwritten Chinese character scoring technology has the problems of relying on rules to limit the difficulty in meeting the needs of huge Chinese character types, conflicts of rules, lack of objectivity and repeated evaluations, especially the rules and template-based methods are not effective when dealing with varied handwritten styles.

Method used

Combining the template-based and machine learning methods, a dictionary containing standard handwritten Chinese character templates is constructed, the Chinese character features are obtained through the feature extraction module, the feature matching module is used for matching, and the similarity is calculated by combining multiple linear regression and BP neural network to achieve automatic scoring.

Benefits of technology

Improves the objectivity and accuracy of the score, reduces rule conflicts and repeated evaluations, adapts to varied handwriting styles, and provides interpretable scoring results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119540966B_ABST
    Figure CN119540966B_ABST
Patent Text Reader

Abstract

The present invention discloses an automatic scoring system and method for the writing standardization of handwritten Chinese character images, which relates to the technical field of image data processing. It includes a dictionary construction module for constructing a dictionary containing standard handwritten Chinese character templates; a feature extraction module for obtaining handwritten Chinese character pictures and extracting feature information from the handwritten Chinese character pictures; a feature matching module for performing feature matching according to the feature information of the feature extraction module and the dictionary construction module to determine which handwritten Chinese character in the standard handwritten Chinese character template dictionary the handwritten Chinese character picture belongs to. By combining the template-based method and the machine learning-based method, and using the similarity index between the character features of the work and the character features of the standard copybook as training data to train the machine learning model, it not only ensures the interpretability of the function implementation but also solves the problem that different machine learning models need to be trained for Chinese characters of different styles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and more specifically, to an automatic scoring system and method for the writing standardization of handwritten Chinese character images. Background Art

[0002] The handwritten Chinese character standardization scoring technology refers to an algorithm in which a computer gives scores to the writing standardization of each Chinese character in a handwritten calligraphy work through an algorithm. The more standardized the writing, the higher the score.

[0003] The existing Chinese character scoring technologies are mainly divided into rule-based, template-based, and deep learning-based methods;

[0004] The rule-based method uses low integer value coding to represent the stroke features and character features of Chinese characters, and judges the correctness of strokes by manually selecting thresholds to verify the stroke actions input by users online. By formulating corresponding Chinese character feature extraction rules and judgment rules for various strokes, the Chinese character writing level is obtained by judging the classification area to which the features belong. At the same time, a decision tree is used to train the microscopic features of strokes, realizing stroke standardization judgment from both macroscopic and microscopic aspects;

[0005] The template-based method starts from imitation, uses a set of standard templates as a reference, and performs similarity analysis with the Chinese characters to be scored, so as to achieve the purpose of scoring the quality of students' handwritten Chinese characters. Combining the line spacing, character spacing of calligraphy works, and projections in the horizontal and vertical directions to enrich the features of the composition layout, so as to achieve a more accurate evaluation of the works; at the same time, it is divided into two methods: similarity evaluation and difference evaluation;

[0006] The deep learning-based method is a method that uses a computer to simulate the learning process of humans. This method continuously optimizes the algorithm performance by learning a large amount of handwritten work data and corresponding scores, and has strong learning ability and adaptability;

[0007] However, in the actual use process, the rule-based Chinese character quality evaluation method is simple to operate and easy to implement. However, due to the huge number and various types of Chinese characters, it is difficult to meet the evaluation needs of all Chinese characters simply relying on the limitations of rules. At the same time, as the number of rules increases, there will be contradictions and conflicts between rules, which will lead to the re-setting of rules. In addition, the styles of handwritten Chinese characters are diverse and difficult to be constrained by rules. Therefore, the results of this type of method are often not ideal;

[0008] The template-based method relies on the analysis and extraction of Chinese character features. At present, most research mainly uses calligraphy experts with many years of practical experience to determine the features of Chinese characters and the importance of features, lacking objectivity. At the same time, there is a correlation between Chinese character features, and it is easy to have the problem of repeated evaluation when manually assigning weight values to features. Summary of the Invention

[0009] To solve the above problems, the present invention provides an automatic scoring system and method for the writing standardization of handwritten Chinese character images.

[0010] The present invention provides an automatic scoring system for the writing standardization of handwritten Chinese character images, including:

[0011] A dictionary construction module, which is used to construct a dictionary containing standard handwritten Chinese character templates;

[0012] A feature extraction module, which is used to obtain a handwritten Chinese character picture and extract feature information from the handwritten Chinese character picture;

[0013] A feature matching module, which is used to perform feature matching according to the feature information of the feature extraction module and the dictionary construction module, and judge which feature of the corresponding Chinese character in the standard handwritten Chinese character template dictionary the feature in the handwritten Chinese character picture belongs to;

[0014] A feature similarity calculation module; the feature similarity calculation module is used to calculate the similarity between the feature information of the feature extraction module and the corresponding handwritten Chinese character in the standard handwritten Chinese character template dictionary in the dictionary construction module according to the judgment result of the feature matching module;

[0015] A scoring module, which is used to give a writing standardization score for the handwritten Chinese character in the handwritten Chinese character picture according to the result of the feature similarity calculation module.

[0016] Preferably, the specific working mode of the feature extraction module is as follows:

[0017] The feature extraction module includes a macro-level extraction unit, a meso-level extraction unit, and a micro-level extraction unit;

[0018] The macro-level extraction unit is used to extract the whole-character features in the handwritten Chinese character image, specifically:

[0019] Calculate the centroid coordinates of the handwritten Chinese character image, use the binary processing result of the image, count the coordinates of each pixel, calculate the centroid position, and complete the extraction of the Chinese character space position;

[0020] By calculating the circumscribed rectangle of the Chinese character in the handwritten Chinese character image, obtain the width and height of the Chinese character, and complete the extraction of the Chinese character size;

[0021] Calculate the width-to-height ratio of the Chinese character in the handwritten Chinese character image, that is, the ratio of the width to the height of the Chinese character, and complete the extraction of the Chinese character width-to-height ratio;

[0022] Perform thinning processing on the handwritten Chinese character image, extract the skeleton structure of the Chinese character, and complete the extraction of the Chinese character skeleton;

[0023] Divide the handwritten Chinese character image into a nine - grid, count the pixel distribution in each small grid, and complete the extraction of the pixel distribution within the nine - grid;

[0024] By scanning the outer contour of the Chinese character in the handwritten Chinese character image, count the area of the outer contour of the Chinese character, and complete the extraction of the outer contour features;

[0025] Scan the skeleton image of the Chinese character in the handwritten Chinese character image, count the area of the internal structure, and complete the extraction of the inner contour features;

[0026] The medium - level extraction unit is used to extract the component features in the handwritten Chinese character image;

[0027] The micro - level extraction unit is used to extract the stroke features in the handwritten Chinese character image.

[0028] Preferably, the specific working steps of the medium - level extraction unit are as follows:

[0029] Segment the Chinese character in the handwritten Chinese character image, calculate the centroid coordinates of each component, determine the relative position of the component centroid and the Chinese character centroid, and complete the extraction of the component spatial position;

[0030] Calculate the circumscribed rectangle of the Chinese character in the handwritten Chinese character image to obtain the bounding box of the Chinese character, calculate the circumscribed rectangle of each component to obtain the bounding box of the component, and calculate the ratio of the area of the component circumscribed rectangle to the area of the Chinese character circumscribed rectangle to complete the extraction of the component size;

[0031] Calculate the width and height of the circumscribed rectangle of each component to complete the extraction of the component width - to - height ratio; by calculating the component width - to - height ratio, evaluate whether the shape of the component is appropriate.

[0032] Preferably, the specific working steps of the micro - level extraction unit are as follows:

[0033] Calculate the centroid coordinates of the Chinese character in the handwritten Chinese character image as a reference point, segment the Chinese character into strokes, calculate the centroid coordinates of each stroke, and determine the relative position of the stroke centroid and the Chinese character centroid to complete the extraction of the stroke spatial position;

[0034] Thin the handwritten Chinese character image, extract the skeleton of the stroke, determine the starting point and ending point of each stroke, calculate the relative position of the stroke endpoint and the Chinese character centroid, and complete the extraction of the stroke endpoints to evaluate the position of the stroke endpoints;

[0035] Determine the length of the stroke by the number of pixel points of the stroke skeleton in the handwritten Chinese character image, traverse the skeleton of each stroke, count the number of pixel points, and complete the extraction of the stroke length:

[0036] Connect the starting points and ending points of each stroke of the Chinese character in the handwritten Chinese character image, calculate the slope of the connection line, and complete the extraction of the stroke slope.

[0037] Preferably, the specific working mode of the feature matching module is as follows:

[0038] Start: Initiate the process of the feature matching algorithm;

[0039] Construct a bipartite graph: Create a bipartite graph for feature matching in the model;

[0040] Initialize the vertex labels: Set initial labels for the nodes in the bipartite graph;

[0041] Hungarian algorithm: Apply the Hungarian algorithm to find the maximum matching in the bipartite graph;

[0042] Yes / No judgment: Check whether a perfect matching is found;

[0043] If yes, end the process;

[0044] If no, proceed to the next step;

[0045] Modify the feasible vertex labels: Adjust the vertex labels to find a better matching;

[0046] End: Complete the process of the feature matching algorithm.

[0047] Preferably, the working steps of the feature similarity calculation module include a multiple linear regression unit and a BP neural network unit, both of which are used to calculate the similarity between the feature information of the feature extraction module and the corresponding handwritten Chinese characters in the standard handwritten Chinese character template dictionary:

[0048] The multiple linear regression unit is used to calculate the similarity through a multiple linear regression model, specifically:

[0049] Normalize the historical feature information of the feature extraction module;

[0050] For each handwritten Chinese character image, construct a historical feature vector containing all the extracted feature values;

[0051] Take the historical feature vector as the independent variable and the historical similarity score between the handwritten Chinese character image and the standard handwritten Chinese character template dictionary as the dependent variable, and construct a multiple linear regression model as the similarity score model;

[0052] Use the previous handwritten Chinese character images and their similarity scores to train the multiple linear regression model;

[0053] For a new handwritten Chinese character image, use the trained model to construct a new feature vector from the feature information of the feature extraction module, input the new feature vector into the multiple linear regression model, and calculate the similarity score W of the multiple linear regression unit with the standard handwritten Chinese character template dictionary;

[0054] Output the calculated similarity score W of the multiple linear regression unit.

[0055] Preferably, the working steps of the BP neural network unit include:

[0056] For a new handwritten Chinese character image, perform a similarity score using the trained model;

[0057] Feedback loop: Compare the prediction result with the expert evaluation to establish a feedback loop.

[0058] Preferably, the working steps of the BP neural network unit include:

[0059] Design a BP neural network model, including an input layer, a hidden layer, and an output layer;

[0060] Select the activation functions for the hidden layer and the output layer;

[0061] Determine the loss function and the optimizer;

[0062] Use the previous handwritten Chinese character images and their similarity scores as a data set, and divide the data set into a training set, a validation set, and a test set;

[0063] Train the model, use the training set data, and monitor the performance on the validation set;

[0064] Use the similarity between the feature information of the feature extraction module and the corresponding handwritten Chinese characters in the standard handwritten Chinese character template dictionary in the dictionary construction module as a feature vector to input into the BP neural network model, and output the similarity score E of the BP neural network unit.

[0065] Preferably, the specific working method of the scoring module is as follows:

[0066] Obtain the similarity score W of the multiple linear regression unit and the similarity score E of the BP neural network unit;

[0067] Obtain the variance R of all similarity scores before the current time period output by the multiple linear regression unit;

[0068] Obtain the variance T of all similarity scores before the current time period output by the BP neural network unit;

[0069] According to the formula Calculate and obtain the confidence level C1 of the multiple linear regression unit;

[0070] According to the formula calculate and obtain the confidence level C2 of the BP neural network unit;

[0071] According to the formula calculate and obtain the final similarity score S.

[0072] The present invention also proposes a method for automatically scoring the writing standardization of a handwritten Chinese character image, including the following steps:

[0073] Step 1: Construct a dictionary containing standard handwritten Chinese character templates, obtain handwritten Chinese character pictures, and extract feature information from the handwritten Chinese character pictures;

[0074] Step 2: Perform feature matching according to the feature information of the feature extraction module and the dictionary construction module, and judge which feature of the corresponding Chinese character in the standard handwritten Chinese character template dictionary the feature in the handwritten Chinese character picture belongs to;

[0075] Step 3: According to the judgment result of the feature matching module, calculate the similarity between the feature information of the feature extraction module and the standard handwritten Chinese character template dictionary corresponding to the handwritten Chinese character in the dictionary construction module, and give the writing standardization score of the handwritten Chinese character in the handwritten Chinese character picture according to the result of the feature similarity calculation module.

[0076] Beneficial effects: Combine the template-based method and the machine learning-based method, use the similarity index between the Chinese character features of the work and the Chinese character features of the standard copybook as training data to train the machine learning model, which not only ensures the interpretability of its function implementation, but also gets rid of the problem that different machine learning models need to be trained for Chinese characters of different styles;

[0077] It realizes inputting a picture of a handwritten Chinese character work, obtaining a scoring result, and quantifying the Chinese character feature information, which can be used as the data basis for other downstream tasks and can be applied to simulate the distribution range of real teaching evaluation scores. Description of the Drawings

[0078] Figure 1 is a schematic diagram of the feature extraction module of the present invention;

[0079] Figure 2 is a schematic diagram of the macro-level extraction unit of the present invention;

[0080] Figure 3 is a schematic diagram of the middle-level extraction unit of the present invention;

[0081] Figure 4 is a schematic diagram of the micro-level extraction unit of the present invention;

[0082] Figure 5 is a flowchart of the whole of the present invention;

[0083] Figure 6 It is the flowchart of the feature matching module of the present invention. Specific embodiments

[0084] Application scenario: In the actual use process, the rule-based Chinese character quality evaluation method is simple to operate and easy to implement. However, due to the huge number and complex types of Chinese characters, it is very difficult to meet the evaluation requirements for all Chinese characters simply relying on rule restrictions. At the same time, as the number of rules increases, there will be conflicts between rules, leading to the need to re-set the rules. In addition, the styles of handwritten Chinese characters vary greatly and it is difficult to use rules to constrain them. Therefore, the results of this type of method are often not ideal.

[0085] The template-based method relies on the analysis and extraction of Chinese character features. Currently, most studies mainly use calligraphy experts with many years of practical experience to determine the features of Chinese characters and the importance of the features, lacking objectivity. At the same time, there is a correlation between Chinese character features, and it is easy to have the problem of repeated evaluation when manually assigning weight values to features.

[0086] As Figures 1 to 6 shown: An automatic scoring system for the writing standardization of handwritten Chinese character images includes:

[0087] A dictionary construction module, which is used to construct a dictionary containing standard handwritten Chinese character templates. It should be noted that in this embodiment, the standard handwritten Chinese character templates can be collected through publicly available standard handwritten Chinese character samples on the network or in cooperation with calligraphy experts, and a dictionary of standard handwritten Chinese character templates can be constructed based on these writing samples. It should be noted that the main core features of the standard handwritten Chinese character template dictionary are the feature information at the macro level, middle level, and micro level of Chinese characters, and the feature information at the three levels of the standard handwritten Chinese character template dictionary is close to the standards of the staff in this field.

[0088] A feature extraction module, which is used to obtain a handwritten Chinese character picture and extract feature information from the handwritten Chinese character picture. It should be noted that in this embodiment, the feature information includes the feature information at the macro level, middle level, and micro level.

[0089] Specifically, at the macro level: it includes the spatial position, size, width-to-height ratio, skeleton, pixel distribution within the nine-square grid, outer contour features, and inner contour features of the Chinese character.

[0090] At the middle level: it includes the spatial position, size, and width-to-height ratio of components.

[0091] At the micro level: it includes the spatial position, endpoints, length, and slope of strokes.

[0092] A feature matching module, which is used to perform feature matching based on the feature information of the feature extraction module and the dictionary construction module, and determine which handwritten Chinese character in the standard handwritten Chinese character template dictionary the handwritten Chinese character picture belongs to;

[0093] A feature similarity calculation module; the feature similarity calculation module is used to calculate the similarity between the feature information of the feature extraction module and the corresponding handwritten Chinese character in the standard handwritten Chinese character template dictionary of the dictionary construction module according to the judgment result of the feature matching module;

[0094] A scoring module, which is used to give a writing standardization score for the handwritten Chinese character in the handwritten Chinese character picture according to the result of the feature similarity calculation module.

[0095] Through the above modules, it is possible to process a large number and variety of Chinese characters, adapt to diverse handwritten styles, avoid conflicts and re - setting between rules caused by increased rule complexity. Through automated feature extraction and matching, the solution reduces the dependence on the personal experience and subjective judgment of calligraphy experts, improves the objectivity and accuracy of scoring, enhances the understanding and evaluation of the relevance between Chinese character features, and reduces the problem of repeated evaluation that may occur in manual feature weight assignment;

[0096] It should be noted that the handwritten Chinese character standardization scoring technology refers to an algorithm in which a computer gives a score for the writing standardization of each Chinese character in a handwritten calligraphy work through an algorithm. The more standardized the writing, the higher the score. The main methods are divided into the rule - based Chinese character quality evaluation method and the template - based method;

[0097] However, in actual use, the rule - based Chinese character quality evaluation method is simple to operate and easy to implement. However, due to the huge number and complex variety of Chinese characters, it is difficult to meet the evaluation requirements for all Chinese characters solely relying on rule restrictions; at the same time, as the number of rules increases, there will be conflicts between rules, leading to the re - setting of rules; in addition, the styles of handwritten Chinese characters are diverse and difficult to be restricted by rules. Therefore, the results of this type of method are often not ideal;

[0098] The template - based method relies on the analysis and extraction of Chinese character features. At present, most research mainly uses calligraphy experts with many years of practical experience to determine the features of Chinese characters and the importance of features, lacking objectivity. At the same time, there is a correlation between Chinese character features, and manual weight assignment of features is prone to the problem of repeated evaluation.

[0099] As an optional embodiment: The specific working mode of the feature extraction module is as follows:

[0100] The feature extraction module includes a macro - level extraction unit, a meso - level extraction unit, and a micro - level extraction unit;

[0101] The macro-level extraction unit is used to extract the whole-character features in the handwritten Chinese character image, specifically as follows:

[0102] Calculate the centroid coordinates of the handwritten Chinese character image. Using the binary processing result of the image, count the coordinates of each pixel and calculate the centroid position to complete the extraction of the spatial position of the Chinese character. It should be noted that the centroid position can be used to describe whether the position of the Chinese character in the Tian character grid is shifted.

[0103] By calculating the circumscribed rectangle of the Chinese character in the handwritten Chinese character image, obtain the width and height of the Chinese character to complete the extraction of the Chinese character size. According to the ratio of the area of the circumscribed rectangle to the area of the Tian character grid, judge whether the writing area of the Chinese character is reasonable.

[0104] Calculate the width-to-height ratio of the Chinese character in the handwritten Chinese character image, that is, the ratio of the width to the height of the Chinese character, to complete the extraction of the width-to-height ratio of the Chinese character. This ratio can be used to judge the degree of height, width, and narrowness of the Chinese character, which affects the beauty of the character.

[0105] Perform thinning processing on the handwritten Chinese character image to extract the skeleton structure of the Chinese character and complete the extraction of the Chinese character skeleton. Commonly used thinning algorithms include the Zhang-Suen thinning algorithm, which can effectively retain the structural features of the Chinese character.

[0106] Divide the handwritten Chinese character image into a nine-grid, count the pixel distribution in each small grid to complete the extraction of the pixel distribution within the nine-grid. By calculating the proportion of white pixels (strokes) in each small grid to the total area, judge whether the distribution of each part of the Chinese character is symmetrical.

[0107] By scanning the outer contour of the Chinese character in the handwritten Chinese character image, count the area of the outer contour of the Chinese character to complete the extraction of the outer contour features. The elastic grid technology can be used to improve the robustness to character deformation and ensure the accuracy of the outer contour features.

[0108] Scan the skeleton image of the Chinese character in the handwritten Chinese character image, count the area of the internal structure to complete the extraction of the inner contour features. By comparing the areas of the first and second passes through the strokes, extract the inner contour features of the Chinese character.

[0109] The middle-level extraction unit is used to extract the component features in the handwritten Chinese character image.

[0110] The micro-level extraction unit is used to extract the stroke features in the handwritten Chinese character image.

[0111] It should be noted that by extracting the center-of-gravity coordinates, bounding rectangle, aspect ratio, skeleton structure, pixel distribution within the nine-square grid, and outer contour features of Chinese characters, it is possible to evaluate whether the overall structure and layout of Chinese characters are standardized. For example, if the center of gravity of a Chinese character is offset, it may mean that the structure of the character is unstable during writing; if the aspect ratio of the Chinese character is unreasonable, it may affect the beauty and recognition of the character.

[0112] As an optional embodiment: The specific working steps of the mesoscopic-level extraction unit are as follows:

[0113] Segment the Chinese characters in the handwritten Chinese character image, calculate the center-of-gravity coordinates of each component, determine the relative position of the component center of gravity and the Chinese character center of gravity, and complete the extraction of the component spatial position; to evaluate whether the component position is offset;

[0114] Calculate the bounding rectangle of the Chinese characters in the handwritten Chinese character image to obtain the bounding box of the Chinese characters, calculate the bounding rectangle of each component, obtain the bounding box of the component, and calculate the ratio of the area of the component bounding rectangle to the area of the Chinese character bounding rectangle to complete the extraction of the component size; to evaluate whether the component size is reasonable;

[0115] Calculate the width and height of the bounding rectangle of each component to complete the extraction of the component aspect ratio; by calculating the component aspect ratio, to evaluate whether the shape of the component is appropriate.

[0116] It should be noted that through component segmentation, calculation of component center-of-gravity coordinates, ratio of component bounding rectangle areas, and extraction of component aspect ratios, it is possible to evaluate whether the positions, sizes, and shapes of the various components in Chinese characters conform to writing specifications. For example, if the relative position of the center of gravity of a component and the center of gravity of the Chinese character is offset, it may mean that the writing position of the component is inaccurate; if the aspect ratio of the component does not match the standard ratio, it may mean that the component is written too wide or too narrow.

[0117] As an optional embodiment: The specific working steps of the microscopic-level extraction unit are as follows:

[0118] Calculate the center-of-gravity coordinates of the Chinese characters in the handwritten Chinese character image as a reference point, segment the strokes of the Chinese characters, calculate the center-of-gravity coordinates of each stroke, determine the relative position of the stroke center of gravity and the Chinese character center of gravity, and complete the extraction of the stroke spatial position; to evaluate whether the stroke position is offset;

[0119] Thin the handwritten Chinese character image, extract the skeleton of the strokes, determine the starting point and ending point of each stroke, calculate the relative position of the stroke endpoints and the Chinese character center of gravity, and complete the extraction of the stroke endpoints to evaluate the positions of the stroke endpoints;

[0120] Determine the length of the stroke by the number of pixel points of the stroke skeleton in the handwritten Chinese character portrait, traverse the skeleton of each stroke, count the number of pixel points, and complete the extraction of the stroke length:

[0121] Connect the starting point and ending point of each stroke in the handwritten Chinese character image, calculate the slope of the connecting line, and complete the stroke slope extraction to determine the direction of the stroke;

[0122] It should be noted that through stroke segmentation, stroke center of gravity coordinate calculation, stroke endpoint extraction, stroke length extraction and stroke slope extraction, it is possible to evaluate whether the position, length and direction of each stroke in a Chinese character are standardized. For example, if the starting point and ending point of a stroke are inaccurate, it may mean that the starting or ending position of the stroke is wrong; if the stroke slope is abnormal, it may mean that the inclination of the stroke does not meet the writing standards.

[0123] As an optional embodiment: the specific working mode of the above feature matching module is as follows:

[0124] Start: Start the process of feature matching algorithm;

[0125] Construct a bipartite graph: Create a bipartite graph for feature matching in the model;

[0126] Initialize top labels: set initial labels for nodes in the bipartite graph;

[0127] Hungarian Algorithm: Apply the Hungarian algorithm to find the maximum matching in a bipartite graph;

[0128] Yes / No Decision: Checks whether a perfect match was found;

[0129] If yes, the process ends;

[0130] If no, go to the next step;

[0131] Modify the feasible top mark: adjust the top mark to find a better match;

[0132] End: Complete the process of feature matching algorithm.

[0133] It should be noted that, for example, we have an image of a handwritten Chinese character, and our goal is to use the feature matching module to determine whether this handwritten Chinese character matches the corresponding character in the standard template;

[0134] For example, the Chinese character in the handwritten Chinese character image is "秀", and the feature matching module compares the features of the handwritten "秀" with the standard "秀" character template in the dictionary;

[0135] Compare the spatial position features (center of gravity positions) of the "禾" and "乃" components of the handwritten "秀" with the corresponding components in the standard template to see if they are aligned; if the center of gravity of the handwritten "禾" component is biased to the left, while it is centered in the standard template, the matching degree will be reduced;

[0136] The strokes in the handwritten "xiu", such as the horizontal stroke of "he" and the left-falling stroke of "nai", have their endpoints, lengths, and slopes matched with the strokes in the standard template. For example, if the starting point of the left-falling stroke of the handwritten "xiu" does not match the position of the starting point in the standard template, then this feature matching will fail.

[0137] The feature matching module will output a matching result indicating the degree of matching between the handwritten "xiu" and the standard template. If all features are highly consistent with the standard template, then the matching is successful, and we can determine that the handwritten Chinese character "xiu" matches the standard template.

[0138] Suppose the center of gravity of the "he" component in the handwritten "xiu" is shifted, while the size of the "nai" component is appropriate. Then the feature matching module will identify the mismatch of the "he" component and may require the user to adjust the position of the "he" component to better match the standard template.

[0139] Through this example, we can see how the feature matching module works and how it helps to identify the degree of matching between handwritten Chinese characters and the standard template. This process is automated and can quickly match and identify a large number of handwritten Chinese characters.

[0140] As an optional embodiment, the working steps of the feature similarity calculation module include a multiple linear regression unit and a BP neural network unit, both of which are used to calculate the similarity between the feature information of the feature extraction module and the corresponding handwritten Chinese characters in the standard handwritten Chinese character template dictionary of the dictionary construction module.

[0141] The multiple linear regression unit is used to calculate the similarity through a multiple linear regression model, specifically as follows:

[0142] Normalize the historical feature information of the feature extraction module. It should be noted that in this embodiment, Min-MaxScaler or StandardScaler can be used.

[0143] For each handwritten Chinese character image, construct a historical feature vector containing all the extracted feature values. It should be noted that the historical feature vector is the historical feature vector of each handwritten Chinese character image collected previously, and these feature vectors are used to train the model.

[0144] Take the historical feature vector as the independent variable and the historical similarity score between the handwritten Chinese character image and the standard handwritten Chinese character template dictionary as the dependent variable, and construct a multiple linear regression model as the similarity score model. It should be noted that in this embodiment, the historical similarity score between the handwritten Chinese character image and the standard handwritten Chinese character template dictionary is obtained through the corresponding system and expert evaluation.

[0145] Train a multiple linear regression model using the handwritten Chinese character images before use and their similarity scores.

[0146] For a new handwritten Chinese character image, use the trained model to construct a new feature vector from the feature information of the feature extraction module, input the new feature vector into the multiple linear regression model, and calculate the similarity score W of the multiple linear regression unit with the standard handwritten Chinese character template dictionary.

[0147] Output the calculated similarity score W of the multiple linear regression unit.

[0148] As an optional embodiment: The working steps of the BP neural network unit include: It should be noted that although the multiple linear regression unit can obtain the similarity score, this method highly depends on the quality of the previous data, and the quality of the data determines the quality of the model. In practical applications, due to the limitation of the data quantity, it is difficult to accurately predict the internal law of the data, and this technical solution can solve the above defects.

[0149] For a new handwritten Chinese character image, use the trained model to perform similarity scoring.

[0150] Feedback loop: Compare the prediction result with the expert evaluation to establish a feedback loop.

[0151] As an optional embodiment: The working steps of the BP neural network unit include:

[0152] Design a BP neural network model, including an input layer, a hidden layer, and an output layer; It should be noted that for the input layer: the number of nodes is equal to the number of features, for the hidden layer: at least one hidden layer, and the number of nodes can be determined according to the problem complexity and the data volume, and usually needs to be adjusted through experiments, for the output layer: a single node, outputting the similarity score.

[0153] Select the activation functions for the hidden layer and the output layer; It should be noted that in this embodiment, for the hidden layer: usually use ReLU or tanh as the activation function, for the output layer: use a linear activation function because the similarity score is a continuous value.

[0154] Determine the loss function and the optimizer; In this embodiment, for the loss function: mean squared error (MSE) or root mean squared error (RMSE), for the optimizer: Adam or SGD (stochastic gradient descent).

[0155] Use the previous handwritten Chinese character images and their similarity scores as the data set, and divide the data set into a training set, a validation set, and a test set; Usually, the ratio is 70% for the training set, 15% for the validation set, and 15% for the test set.

[0156] Train the model, use the training set data, and monitor the performance on the validation set.

[0157] The newly extracted feature information of the feature extraction module is used as a feature vector and input into the BP neural network model to output the similarity score E of the BP neural network unit.

[0158] For example: We have an image of a handwritten Chinese character, and the following are the feature data extracted from the macroscopic, mesoscopic, and microscopic levels:

[0159] Macroscopic level features:

[0160] Spatial position (center coordinates): (0.3, 0.5);

[0161] Size (width, height): (0.8, 0.6);

[0162] Aspect ratio: 1.33;

[0163] Skeleton: The proportion of the number of skeleton pixels in the total area is 0.15;

[0164] Pixel distribution within the nine-square grid: [0.1, 0.2, 0.3, 0.4, 0.5, 0.4, 0.3, 0.2, 0.1];

[0165] Outer contour feature: The proportion of the outer contour area in the total area is 0.2;

[0166] Inner contour feature: The proportion of the inner contour area in the total area is 0.1;

[0167] Mesoscopic level features:

[0168] Component spatial position (center coordinates): (0.4, 0.6);

[0169] Component size (area of the circumscribed rectangle): 0.7;

[0170] Component aspect ratio: 1.2;

[0171] Microscopic level features:

[0172] Stroke spatial position (center of gravity coordinates): (0.5, 0.7);

[0173] Stroke endpoints: Starting point (0.2, 0.3), ending point (0.8, 0.9);

[0174] Stroke length: 1.5;

[0175] Stroke slope: The slope of the line connecting the starting point and the ending point is calculated to be 0.75;

[0176] Input layer: The number of input layer nodes is 15, corresponding to the above 15 features;

[0177] Hidden layer: Assume that we set two hidden layers. The first hidden layer has 20 nodes and the second hidden layer has 10 nodes;

[0178] Output layer: The number of nodes in the output layer is 1, and the similarity score is output;

[0179] Activation function: The ReLU activation function is used for the hidden layer, and the linear activation function is used for the output layer;

[0180] Loss function: Mean Squared Error (MSE);

[0181] Optimizer: Adam optimizer;

[0182] Assume that we have 100 images of handwritten Chinese characters with their labeled feature data, as well as the similarity scores with the standard templates; we divide these data into a training set (70 samples), a validation set (15 samples), and a test set (15 samples);

[0183] Use the training set data to train the BP neural network model, and monitor the model performance on the validation set to prevent overfitting; during the training process, we adjust parameters such as the learning rate and batch size to obtain the best model performance;

[0184] For a new image of a handwritten Chinese character, we extract the same 15 features, construct a feature vector, and input it into the trained BP neural network model to calculate the similarity score with the standard handwritten Chinese character template;

[0185] The model outputs the predicted similarity score. For example, 0.85 indicates that the similarity between the handwritten Chinese character image and the standard template is 85%.

[0186] As an optional embodiment: The specific working method of the scoring module is as follows:

[0187] Obtain the similarity score W of the multiple linear regression unit and the similarity score E of the BP neural network unit;

[0188] Obtain the variance R of all similarity scores before the current time period output by the multiple linear regression unit;

[0189] Obtain the variance T of all similarity scores before the current time period output by the BP neural network unit;

[0190] According to the formula Calculate and obtain the confidence level C1 of the multiple linear regression unit;

[0191] According to the formula Calculate and obtain the confidence level C2 of the BP neural network unit;

[0192] According to the formula Calculate to obtain the final similarity score S. It should be noted that the confidence level can dynamically reflect the latest performance of the model, increase the diversity of model fusion, comprehensively consider the confidence level and the correlation between models, and adjust the weights of the models in the final score, so that the final similarity score is more accurate and robust.

[0193] The present invention also proposes a method for automatically scoring the writing standardization of handwritten Chinese character images, including the following steps:

[0194] Step 1: Construct a dictionary containing standard handwritten Chinese character templates, obtain handwritten Chinese character pictures, and extract feature information from the handwritten Chinese character pictures;

[0195] Step 2: Perform feature matching according to the feature information of the feature extraction module and the dictionary construction module, and determine which feature of the corresponding Chinese character in the standard handwritten Chinese character template dictionary the feature in the handwritten Chinese character picture belongs to;

[0196] Step 3: According to the judgment result of the feature matching module, calculate the similarity between the feature information of the feature extraction module and the corresponding handwritten Chinese character in the standard handwritten Chinese character template dictionary of the dictionary construction module, and give the writing standardization score of the handwritten Chinese character in the handwritten Chinese character picture according to the result of the feature similarity calculation module.

[0197] Working principle

[0198] Construct a dictionary containing standard handwritten Chinese character templates, obtain handwritten Chinese character pictures, and extract feature information from the handwritten Chinese character pictures; perform feature matching according to the feature information of the feature extraction module and the dictionary construction module, and determine which handwritten Chinese character in the standard handwritten Chinese character template dictionary the handwritten Chinese character picture belongs to; according to the judgment result of the feature matching module, calculate the similarity between the feature information of the feature extraction module and the corresponding handwritten Chinese character in the standard handwritten Chinese character template dictionary of the dictionary construction module, and give the writing standardization score of the handwritten Chinese character in the handwritten Chinese character picture according to the result of the feature similarity calculation module. The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be pointed out that for those of ordinary skill in the art in this technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as the protection scope of this template.

Claims

1. An automatic scoring system for the writing standardization of handwritten Chinese character images, characterized in that, Including: A dictionary construction module for constructing a dictionary containing standard handwritten Chinese character templates; A feature extraction module for obtaining a handwritten Chinese character image and extracting feature information from the handwritten Chinese character image; A feature matching module for performing feature matching based on the feature information of the feature extraction module and the dictionary construction module, and determining which feature of the corresponding Chinese character in the standard handwritten Chinese character template dictionary the feature in the handwritten Chinese character image belongs to; A feature similarity calculation module; The feature similarity calculation module is used to calculate the similarity between the feature information of the feature extraction module and the corresponding handwritten Chinese character in the standard handwritten Chinese character template dictionary according to the judgment result of the feature matching module; A scoring module for giving a writing standardization score of the handwritten Chinese character in the handwritten Chinese character image according to the result of the feature similarity calculation module; The working steps of the feature similarity calculation module include a multiple linear regression unit and a BP neural network unit, both of which are used to calculate the similarity between the feature information of the feature extraction module and the corresponding handwritten Chinese character in the standard handwritten Chinese character template dictionary: For a new handwritten Chinese character image, using the trained model, construct a new feature vector from the feature information of the feature extraction module, input the new feature vector into the multiple linear regression model, and calculate the similarity score W with the multiple linear regression unit of the standard handwritten Chinese character template dictionary; Output the calculated similarity score W of the multiple linear regression unit; Use the similarity between the feature information of the feature extraction module and the corresponding handwritten Chinese character in the standard handwritten Chinese character template dictionary in the dictionary construction module as a feature vector to input into the BP neural network model, and output the similarity score E of the BP neural network unit; The specific working mode of the scoring module is as follows: Obtain the similarity score W of the multiple linear regression unit and the similarity score E of the BP neural network unit; Obtain the variance R of all similarity scores before the current time period output by the multiple linear regression unit; Obtain the variance T of all similarity scores before the current time period output by the BP neural network unit; According to the formula , calculate to obtain the confidence level of the multiple linear regression unit ; According to the formula , calculate to obtain the confidence level of the BP neural network unit ; According to the formula , the final similarity score S is calculated and obtained.

2. The automatic scoring system for the writing standardization of handwritten Chinese character images according to claim 1, characterized in that, The specific working mode of the feature extraction module is as follows: The feature extraction module includes a macro-level extraction unit, a meso-level extraction unit, and a micro-level extraction unit; The macro-level extraction unit is used to extract the whole-character features in the handwritten Chinese character image, specifically: Calculate the centroid coordinates of the handwritten Chinese character image, use the binary processing result of the image, count the coordinates of each pixel, calculate the centroid position, and complete the extraction of the Chinese character spatial position; Obtain the width and height of the Chinese character by calculating the circumscribed rectangle of the Chinese character in the handwritten Chinese character image, and complete the extraction of the Chinese character size; Calculate the width-to-height ratio of the Chinese character in the handwritten Chinese character image, that is, the ratio of the width to the height of the Chinese character, and complete the extraction of the Chinese character width-to-height ratio; Perform thinning processing on the handwritten Chinese character image, extract the skeleton structure of the Chinese character, and complete the extraction of the Chinese character skeleton; Divide the handwritten Chinese character image into a nine - grid, count the pixel distribution in each small grid, and complete the extraction of the pixel distribution within the nine - grid; By scanning the outer contour of the Chinese character in the handwritten Chinese character image, count the area of the outer contour of the Chinese character, and complete the extraction of the outer contour features; Scan the skeleton image of the Chinese character in the handwritten Chinese character image, count the area of the internal structure, and complete the extraction of the inner contour features; The medium - level extraction unit is used to extract the component features in the handwritten Chinese character image; The micro - level extraction unit is used to extract the stroke features in the handwritten Chinese character image.

3. The automatic scoring system for the writing standardization of handwritten Chinese character images according to claim 2, characterized in that, The specific working steps of the medium - level extraction unit are as follows: Segment the components of the Chinese character in the handwritten Chinese character image, calculate the centroid coordinates of each component, determine the relative position between the component centroid and the Chinese character centroid, and complete the extraction of the component spatial position; Calculate the circumscribed rectangle of the Chinese character in the handwritten Chinese character image to obtain the bounding box of the Chinese character. Calculate the circumscribed rectangle for each component to obtain the bounding box of the component. Calculate the ratio of the area of the component circumscribed rectangle to the area of the Chinese character circumscribed rectangle, and complete the extraction of the component size; Calculate the width and height of the circumscribed rectangle of each component, and complete the extraction of the component width - to - height ratio; By calculating the component width - to - height ratio, evaluate whether the shape of the component is appropriate.

4. The automatic scoring system for the writing standardization of handwritten Chinese character images according to claim 2, wherein The specific working steps of the micro - level extraction unit are as follows: Calculate the centroid coordinates of the Chinese character in the handwritten Chinese character image as a reference point, segment the strokes of the Chinese character, calculate the centroid coordinates of each stroke, determine the relative position between the stroke centroid and the Chinese character centroid, and complete the extraction of the stroke spatial position; Thin the handwritten Chinese character image, extract the skeleton of the stroke, determine the starting point and ending point of each stroke, calculate the relative position between the stroke endpoints and the Chinese character centroid, and complete the extraction of the stroke endpoints to evaluate the position of the stroke endpoints; Determine the length of the stroke by the number of pixel points of the stroke skeleton in the handwritten Chinese character image. Traverse the skeleton of each stroke and count the number of pixel points to complete the extraction of the stroke length; Connect the starting point and ending point of each stroke in the handwritten Chinese character image, calculate the slope of the connection line, and complete the extraction of the stroke slope.

5. The automatic scoring system for the writing standardization of handwritten Chinese character images according to claim 1, characterized in that, The specific working mode of the feature matching module is as follows: Start: Initiate the process of the feature matching algorithm; Construct a bipartite graph: Create a bipartite graph for feature matching in the model; Initialize the top labels: Set the initial labels for the nodes in the bipartite graph; Hungarian algorithm: Apply the Hungarian algorithm to find the maximum matching in the bipartite graph; Yes / No judgment: Check whether a perfect match is found; If yes, the process ends; If no, go to the next step; Modify the feasible top labels: Adjust the top labels to find a better match; End: Complete the process of the feature matching algorithm.