A multimodal method for evaluating the quality of Chinese character handwriting and related equipment.

By using a multimodal Chinese character writing quality evaluation method, pen tip data and writing result images are obtained, and an alignment model between stroke nodes and character structure diagrams is constructed. This solves the problems of insufficient reliability and interpretability of existing evaluation methods, and realizes precise positioning and multifaceted evaluation of Chinese character writing quality.

CN122090462APending Publication Date: 2026-05-26CENT SOUTH UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2026-04-27
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing methods for evaluating the quality of Chinese character writing are insufficient in terms of reliability and interpretability. Static character shape analysis is difficult to reflect the dynamic behavioral characteristics during the writing process, and trajectory analysis methods lack a clear correspondence with character structure, affecting the accuracy and interpretability of the evaluation results.

Method used

By acquiring pen tip data and writing result images at multiple time points during the Chinese character writing process, we define stroke nodes, construct character structure diagrams, establish alignment models between stroke nodes and structure nodes, extract feature representations of the Chinese character writing process, and conduct multimodal evaluation.

Benefits of technology

It enables precise localization of handwriting quality issues within the character shape, improving the interpretability and reliability of the evaluation. By combining information from images and pen tip data, it enhances the multifaceted nature and information richness of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090462A_ABST
    Figure CN122090462A_ABST
Patent Text Reader

Abstract

This application relates to the field of handwriting evaluation technology, and provides a multimodal method and related equipment for evaluating the quality of Chinese character handwriting. The method includes: acquiring pen tip data at multiple time points during the Chinese character writing process and obtaining writing result images; defining multiple stroke nodes based on all pen tip data; extracting multiple structural nodes from the writing result images and constructing a character structure diagram based on all structural nodes; establishing an alignment model based on all stroke nodes and the character structure diagram; extracting feature representations of the Chinese character writing process based on the alignment model, and evaluating based on these feature representations to obtain the handwriting quality evaluation result. The method of this application can improve the reliability and interpretability of handwriting quality evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of handwriting evaluation technology, and in particular to a multimodal method and related equipment for evaluating the quality of Chinese character handwriting. Background Technology

[0002] With the continuous development of electronic writing devices, intelligent writing terminals, and digital teaching systems, objective and automated evaluation of handwriting quality has been widely applied in handwriting instruction, ability assessment, human-computer interaction, and intelligent education. Analyzing the writing process and results to provide learners with timely and accurate feedback has become an important technical means to improve the efficiency and quality of handwriting training. Therefore, constructing reliable handwriting quality evaluation methods has significant practical implications.

[0003] Currently, research on the quality of Chinese character handwriting mainly focuses on analysis methods based on static character images or rule-based modeling methods based on writing trajectories. Image-based methods typically evaluate handwriting based on features such as character outlines, structural proportions, or stroke shapes; while trajectory-based methods emphasize statistical analysis of writing paths, speed variations, and pressure information. Some studies have also attempted to combine machine learning models to model handwriting data in order to achieve automatic scoring or classification.

[0004] However, the above methods still have certain limitations in practical applications: on the one hand, static character shape analysis is difficult to reflect the dynamic behavioral characteristics during the writing process; on the other hand, the trajectory analysis method lacks a clear correspondence with the character structure, making it difficult to locate the specific position of quality problems within the character shape, thus affecting the interpretability of the evaluation results. Therefore, it is evident that the current evaluation of Chinese character writing quality has poor reliability and interpretability. Summary of the Invention

[0005] This application provides a multimodal method and related equipment for evaluating the quality of Chinese character writing, which can solve the problems of poor reliability and interpretability in the evaluation of Chinese character writing quality.

[0006] Firstly, this application provides a multimodal method for evaluating the quality of Chinese character handwriting, which includes: Acquire pen tip data at multiple time points during the writing process of Chinese characters, and obtain images of the writing results; Multiple stroke nodes are defined based on all pen tip data; Multiple structural nodes are extracted from the writing result image, and a character structure graph is constructed based on all structural nodes. In the character structure graph, multiple nodes correspond one-to-one with multiple structural nodes, and the edges between nodes represent the connection relationship between the two corresponding structural nodes. An alignment model is established based on all stroke nodes and character structure diagrams; the alignment model is used to describe the correspondence between stroke nodes and structure nodes. The writing process of Chinese characters is extracted based on the alignment model, and the writing quality is evaluated based on the feature representation to obtain the writing quality evaluation result.

[0007] Optionally, multiple stroke nodes can be defined based on all pen tip data, including: For each time point, the pen tip state is determined based on the pen tip data at that time point; the pen tip state can be a contact state or a non-contact state. The time point corresponding to the contact state is used as the writing time point, and a stroke node is defined based on the pen tip data corresponding to every two adjacent writing time points.

[0008] Optionally, the pen tip state at a given time point can be determined based on the pen tip data at that time point, including: Through the formula: ; Calculate the first Pen nib state at a specific point in time ; in, This indicates that the pen tip is in a non-contact state. This indicates that the pen tip is in contact. Indicates the first The state of the pen tip at a specific point in time. Indicates the first Pen tip pressure data at each time point Indicates the threshold for pen placement. Indicates the pen lifting threshold. , Indicates the number of time points; The stroke nodes are: ; ; ; in, Indicates the first The tuple corresponding to each stroke node Indicates the first A sequence vector of stroke nodes Indicates the first The attribute vector of each stroke node. Indicates the first The two-dimensional coordinates of the pen tip in the pen tip data at each time point The x-axis is... The vertical axis is , Indicates the first At a certain point in time, Indicates the first Each stroke node corresponds to a writing time point. Indicates the first Another writing time point corresponding to each stroke node Indicates the first The duration of each stroke node Indicates the first The trajectory length of each stroke node Indicates the first Average pressure of each stroke node Indicates the first Average pen tip speed of each stroke node Indicates the first Standard deviation of pen tip speed for each stroke node , This indicates the number of stroke nodes.

[0009] Optionally, based on all stroke nodes and character structure diagrams, an alignment model is established, including: For each stroke node, calculate the coverage probability between the stroke node and each edge in the character structure graph; Align all stroke nodes and all edges in the glyph structure graph according to all coverage probabilities to obtain the alignment distribution between each stroke node and each edge. By integrating all alignment distributions into a single dataset, we obtain the alignment model.

[0010] Optionally, calculate the coverage probability between stroke nodes and each edge in the character structure graph, including: Through the formula: ; Calculate the first The stroke node and the first The probability of coverage between edges ; in, Indicates the first The stroke node and the first The average distance between the edges This is the distance smoothing coefficient. Indicates the first The stroke node and the first The average distance between the edges This represents the set of edges in a character structure diagram.

[0011] Optionally, align all stroke nodes and all edges in the glyph structure graph according to all coverage probabilities to obtain the alignment distribution between each stroke node and each edge, including: Each stroke node is feature-encoded to obtain a temporal feature representation of each stroke node; Each edge in the character structure graph is feature-encoded to obtain the structural feature representation of each edge; For each stroke node, the initial alignment distribution between the stroke node and each edge is calculated based on the temporal feature representation of the stroke node and the structural feature representation of each edge. Calculate the loss function value based on all coverage probabilities and all initial alignment distributions; The initial alignment distribution is updated based on the loss function value to obtain the alignment distribution between each stroke node and each edge.

[0012] Optionally, based on the temporal feature representation of the stroke nodes and the structural feature representation of each edge, the initial alignment distribution between the stroke nodes and each edge is calculated, including: Through the formula: ; ; Calculate the first The stroke node and the first Initial alignment distribution between edges ; in, Indicates the first The stroke node and the first Alignment score between each edge Indicates the first The stroke node and the first Alignment score between each edge Indicates the first Temporal feature representation of each stroke node Indicates the first Structural feature representation of each edge; The loss function value is calculated based on all coverage probabilities and all initial alignment distributions, including: Through the formula: ; Calculate the loss function value ; in, Indicates the first A vector consisting of all coverage probabilities corresponding to each stroke node. Indicates the first The vector consisting of all the initial alignment distributions corresponding to each stroke node.

[0013] Optionally, feature representations of the Chinese character writing process can be extracted based on the alignment model, including: Based on the alignment model, stroke topology preservation constraints, mechanical consistency constraints, feature consistency constraints, and quality sensitivity constraints are constructed. Under the constraints of stroke topology preservation, mechanical consistency, feature consistency, and quality sensitivity, features are extracted from the Chinese character writing process to obtain a feature representation of the Chinese character writing process.

[0014] Secondly, this application provides a multimodal Chinese character writing quality evaluation device, comprising: The acquisition module is used to acquire pen tip data at multiple time points during the writing process of Chinese characters and to acquire images of the writing results. Define the module to define multiple stroke nodes based on all pen tip data; The extraction module is used to extract multiple structural nodes from the writing result image and construct a character structure graph based on all structural nodes. In the character structure graph, multiple nodes correspond one-to-one with multiple structural nodes, and the edges between nodes represent the connection relationship between the two corresponding structural nodes. A module is established to create an alignment model based on all stroke nodes and character structure diagrams; the alignment model is used to describe the correspondence between stroke nodes and structure nodes. The evaluation module is used to extract feature representations of the Chinese character writing process based on the alignment model, and to evaluate the writing quality based on these feature representations.

[0015] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the aforementioned multimodal Chinese character writing quality evaluation method.

[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned multimodal Chinese character writing quality evaluation method.

[0017] The above-mentioned solution in this application has the following beneficial effects: In the embodiments of this application, pen tip data at multiple time points during the writing process of Chinese characters is acquired, and the writing result image is obtained. Then, multiple stroke nodes are defined based on all pen tip data, and multiple structural nodes are extracted from the writing result image. A character structure diagram is constructed based on all structural nodes. Then, an alignment model is established based on all stroke nodes and the character structure diagram. Finally, feature representations of the Chinese character writing process are extracted based on the alignment model, and evaluation is performed based on the feature representations to obtain the writing quality evaluation result. Specifically, evaluating Chinese character writing based on the writing result image and pen tip data considers both the resulting image and the pen tip motion information during the writing process. This allows the dynamic behavioral features during the writing process to be mapped to the character structure space, enabling precise localization of writing quality issues within the character shape. This improves the interpretability of the writing quality evaluation. Furthermore, evaluating Chinese character writing quality based on multimodal data combines information from both image and pen tip data, increasing information richness and evaluation multifacetedness, effectively improving the reliability of the Chinese character writing quality evaluation.

[0018] Other beneficial effects of this application will be described in detail in the following detailed description section. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart of a multimodal Chinese character writing quality evaluation method provided in an embodiment of this application; Figure 2 A schematic diagram of the structure of a multimodal Chinese character writing quality evaluation device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0021] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0022] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0023] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0024] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0025] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0026] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0027] To address the issues of poor reliability and interpretability in existing Chinese character writing quality evaluation methods, this application provides a multimodal Chinese character writing quality evaluation method. This method acquires pen tip data at multiple time points during the Chinese character writing process and obtains the writing result image. Then, it defines multiple stroke nodes based on all pen tip data, extracts multiple structural nodes from the writing result image, and constructs a character structure diagram based on all structural nodes. Next, it establishes an alignment model based on all stroke nodes and the character structure diagram. Finally, it extracts feature representations of the Chinese character writing process based on the alignment model and performs evaluation based on these feature representations to obtain the writing quality evaluation result. Specifically, evaluating Chinese character writing based on the writing result image and pen tip data considers both the resulting image and the pen tip's motion information during writing, enabling the dynamic behavioral features during writing to be mapped to the character structure space. This allows for precise localization of writing quality issues within the character structure, improving the interpretability of the writing quality evaluation. Furthermore, using multimodal data for Chinese character writing quality evaluation combines information from both image and pen tip data, increasing information richness and evaluation multifacetedness, effectively improving the reliability of the Chinese character writing quality evaluation.

[0028] The following is an exemplary description of the multimodal Chinese character writing quality evaluation method provided in this application.

[0029] like Figure 1 As shown, the multimodal Chinese character writing quality evaluation method provided in this application includes the following steps: Step 11: Obtain pen tip data at multiple time points during the Chinese character writing process, and obtain the writing result image.

[0030] The above pen tip data includes the pen tip's two-dimensional coordinates (which can be coordinates in a coordinate system with the paper as a two-dimensional plane during the writing of Chinese characters), pen tip pressure, etc. The above writing result image is the image of the Chinese characters after writing.

[0031] In some embodiments of this application, pen tip data can be acquired using devices such as pressure sensors and positioning systems, and images of the writing results can be acquired using devices such as cameras.

[0032] Step 12: Define multiple stroke nodes based on all pen tip data.

[0033] The stroke nodes mentioned above are used to describe each stroke of a Chinese character during the writing process.

[0034] In some embodiments of this application, the step of defining multiple stroke nodes based on all pen tip data includes: The first step is to determine the pen tip status at each time point based on the pen tip data at that time point.

[0035] The pen tip states described above refer to both contact and non-contact states. Contact state means the pen tip is in contact with the paper, while non-contact state means the pen tip is not in contact with the paper.

[0036] Specifically, through the formula: ; Calculate the first Pen nib state at a specific point in time .

[0037] in, This indicates that the pen tip is in a non-contact state. This indicates that the pen tip is in contact. Indicates the first The state of the pen tip at a specific point in time. Indicates the first Pen tip pressure data at each time point Indicates the threshold for pen placement. Indicates the pen lifting threshold. , Indicates the number of time points.

[0038] The second step is to use the time point corresponding to the contact state as the writing time point, and define a stroke node based on the pen tip data corresponding to every two adjacent writing time points.

[0039] The stroke nodes are: ; ; ; in, Indicates the first The tuple corresponding to each stroke node Indicates the first A sequence vector of stroke nodes Indicates the first The attribute vector of each stroke node. Indicates the first The two-dimensional coordinates of the pen tip in the pen tip data at each time point The x-axis is... The vertical axis is , Indicates the first At a certain point in time, Indicates the first Each stroke node corresponds to a writing time point. Indicates the first Another writing time point corresponding to each stroke node Indicates the first The duration of each stroke node Indicates the first The trajectory length of a stroke node, denotes the average pressure of the th stroke node, denotes the average tip speed of the th stroke node, denotes the standard deviation of the tip speed of the , denotes the number of stroke nodes.

[0040] It should be noted that after defining the stroke nodes, the sequential relationship between strokes can be further constructed: when the end time of stroke node is earlier than the start time of stroke node , that is, satisfying: ; when, a directed connection relationship is established between them: ; Thus, a stroke topological structure is constructed: ; where: ; Since the above connection relationship strictly follows the chronological order, the constructed stroke topological structure is a directed acyclic graph, and its topological sorting result can reflect the true writing order of strokes.

[0041] Step 13, extract multiple structure nodes from the writing result image, and construct a glyph structure diagram based on all the structure nodes.

[0042] The above structure nodes are the endpoints, intersection points, bifurcation points, etc. of Chinese characters. Multiple nodes in the above glyph structure diagram correspond to multiple structure nodes one by one, and the edges between the nodes represent the connection relationship between the corresponding two structure nodes.

[0043] It should be noted that models such as YOLO can be used to extract structure nodes from the writing result image. If two structure nodes are connected by a Chinese character stroke, it is considered that there is a connection relationship between the two structure nodes. For example, for the stroke "-", its two endpoints are two structure nodes respectively, and they are connected by a stroke, so there is a connection relationship. Generate a corresponding node for each structure node. If there is a connection relationship between two structure nodes, then generate an edge between the corresponding two nodes to obtain the glyph structure diagram.

[0044] For example, image preprocessing operations are performed sequentially on the written image, including grayscale conversion, binarization, and denoising, to eliminate background interference and highlight the main character area. Based on this, a skeletonization operation is performed on the preprocessed image, converting the original character area with a certain stroke width into a single-pixel-wide centerline structure, thereby extracting the geometric skeleton of the character. Subsequently, structural analysis is performed on the skeleton image to identify endpoints, intersections, and bifurcation points in the skeleton, using continuous skeleton paths as edges to construct the character structure diagram.

[0045] Step 14: Establish an alignment model based on all stroke nodes and character structure diagrams.

[0046] The alignment model described above is used to describe the correspondence between stroke nodes and structure nodes.

[0047] In some embodiments of this application, the steps of establishing an alignment model based on all stroke nodes and glyph structure diagrams include: The first step is to calculate the coverage probability between each stroke node and each edge in the character structure graph.

[0048] Specifically, through the formula: ; Calculate the first The stroke node and the first The probability of coverage between edges .

[0049] in, Indicates the first The stroke node and the first The average distance between each edge (i.e., the average distance between the two-dimensional coordinates of the pen tip corresponding to the stroke node and the structural node corresponding to the edge). This is the distance smoothing coefficient. Indicates the first The stroke node and the first The average distance between the edges This represents the set of edges in a character structure diagram.

[0050] The second step is to align all stroke nodes and all edges in the character structure graph according to all coverage probabilities, so as to obtain the alignment distribution between each stroke node and each edge.

[0051] First, feature encoding is performed on each stroke node to obtain the temporal feature representation of each stroke node.

[0052] For example, the expression for feature encoding of stroke nodes is as follows: ,in, Indicates the first Temporal feature representation of each stroke node Indicates the first A sequence vector of stroke nodes This represents a timing coding function.

[0053] Then, feature encoding is performed on each edge in the character structure graph to obtain the structural feature representation of each edge.

[0054] For example, the expression for feature encoding of each edge in the glyph structure graph is as follows: ,in, Indicates the first Structural features of each edge Represents the structure encoding function. This indicates that the structure encoding function is used to encode the first... Each edge is encoded.

[0055] Then, for each stroke node, the initial alignment distribution between the stroke node and each edge is calculated based on the temporal feature representation of the stroke node and the structural feature representation of each edge.

[0056] Specifically, through the formula: ; ; Calculate the first The stroke node and the first Initial alignment distribution between edges .

[0057] in, Indicates the first The stroke node and the first Alignment score between each edge Indicates the first The stroke node and the first Alignment score between each edge Indicates the first Temporal feature representation of each stroke node Indicates the first Structural features of each edge.

[0058] Then, the loss function value is calculated based on all coverage probabilities and all initial alignment distributions.

[0059] Specifically, through the formula: ; Calculate the loss function value .

[0060] in, Indicates the first A vector consisting of all coverage probabilities corresponding to each stroke node. Indicates the first The vector consisting of all the initial alignment distributions corresponding to each stroke node.

[0061] Finally, the initial alignment distribution is updated based on the loss function value to obtain the alignment distribution between each stroke node and each edge.

[0062] For example, the parameters in the formula for calculating the initial alignment distribution are updated using algorithms such as gradient descent (e.g., the parameters in the function for feature encoding of stroke nodes and edges), and the stroke nodes and character structure graph are substituted into the updated formula to recalculate the initial alignment distribution. The loss function value is then calculated again using the above formula until the loss function value is less than the preset loss function threshold, and the alignment distribution is output.

[0063] The third step is to integrate all the alignment distributions into one dataset to obtain the alignment model.

[0064] It should be noted that before proceeding to this step, in order to unify the coordinate data of the stroke nodes with the coordinates of the character structure, the coordinate data of the stroke nodes needs to be mapped to the image coordinate system. Specifically, this is done through a coordinate mapping function. Map the coordinate data to the image coordinate system, where, Represents the two-dimensional coordinates of the pen tip. This represents the two-dimensional coordinates of the pen tip mapped to the image coordinate system. This is a coordinate mapping function used to describe the mapping relationship between the coordinates of the writing trajectory and the coordinates of the character structure. Through this mapping, each two-dimensional coordinate point of the pen tip in the original writing trajectory is transformed into the coordinate space corresponding to the character structure diagram. After the above processing: the writing trajectory is transformed from the original writing coordinate system to the character structure coordinate system; the character structure diagram and the writing trajectory are in a unified spatial reference frame, thus enabling subsequent analysis of the spatial positional relationship of stroke nodes in the character structure under a unified coordinate system.

[0065] After the spatial integration of the writing trajectory and the character structure is completed, the writing trajectory points and the character structure diagram are now in a unified coordinate system. However, it is still not possible to directly determine the specific corresponding position of each stroke node in the character structure. It is necessary to establish a clear correspondence between the stroke nodes formed by continuous writing and their specific structural positions in the character structure diagram, so that subsequent analysis of writing quality can be performed at the character structure level.

[0066] Specifically, the mapping function from trajectory points to structural edges is defined as follows: ; in, Representing the trajectory points and structural edges The minimum Euclidean distance between them. Through the above mapping function, the originally continuous trajectory points can be transformed into discrete assignment results of the edges of the character structure, thereby realizing the projection of the writing trajectory onto the character structure space.

[0067] After obtaining the structural mapping results at the trajectory point level, the mapping results are further aggregated at the stroke node level. For the first... The time interval corresponding to each stroke node is: The structural edges corresponding to all trajectory points within the time interval are set-based to construct the covering set of the stroke node in the character structure. : ; The structural coverage set is used to characterize the range of structural paths involved in the stroke in the character structure diagram. Through this process, stroke nodes that originally exist in the form of a time sequence are transformed into structural units with definite coverage positions in the character structure space.

[0068] Step 15: Extract feature representations of the Chinese character writing process based on the alignment model, and evaluate based on the feature representations to obtain writing quality evaluation results.

[0069] The above writing quality evaluation results are used to evaluate the writing quality of the Chinese character writing process. They can be scored, and the higher the score, the better the writing quality.

[0070] In some embodiments of this application, the steps of extracting feature representations of the Chinese character writing process based on the alignment model and evaluating the writing quality based on these feature representations to obtain writing quality evaluation results include: The first step is to construct stroke topology preservation constraints, mechanical consistency constraints, feature consistency constraints, and quality sensitivity constraints based on the alignment model.

[0071] Specifically, stroke topology preservation constraints for: ; in, Represents the set of stroke nodes. As weight, For the first Temporal feature representation of each stroke node For the first Temporal feature representation of each stroke node.

[0072] Mechanical consistency constraints for: ; in, Indicates the first The end index of each stroke node in the writing time series. Indicates the first The starting index of each stroke node in the writing time series. This represents the weighting coefficient of the speed smoothing term. Indicates the first Pen tip pressure data at each time point Indicates the first The pressure on the pen tip at a specific point in time. Indicates the first The pressure on the pen tip at a specific point in time. Indicates the first Pen tip speed at a given time point Indicates the first Pen tip speed at a given time point Indicates the first The pen tip speed at each point in time.

[0073] Feature Consistency Constraints for: ; ; in, Indicates the first Weighted aggregation features of stroke nodes For the alignment model, the first The stroke node and the first Alignment distribution between the edges.

[0074] Quality-sensitive constraints for: ; in, This represents the preset sorting interval hyperparameter. Indicates the first Quality score of each stroke node. Indicates the first The quality score of the degraded writing sample corresponding to each stroke node.

[0075] The second step involves extracting features from the Chinese character writing process under constraints of stroke topology preservation, mechanical consistency, feature consistency, and quality sensitivity, thereby obtaining a feature representation of the Chinese character writing process.

[0076] The above features are calculated based on the alignment distribution between stroke nodes and edges and the structural features of the edges.

[0077] For example, the feature representation is not directly obtained from single-modal information, but is obtained through joint optimization calculation based on the temporal feature representation of stroke nodes and the structural feature representation of character structure edges under structural constraints. Specifically, firstly, for the... The writing sequence of each stroke node. Encode it to obtain its temporal feature representation. At the same time, for each structural edge in the character structure diagram Feature encoding is performed to obtain the corresponding structural feature representation. Subsequently, based on the alignment distribution between stroke nodes and structural edges... We then perform weighted aggregation on the features of each structural edge to obtain the structural aggregated representation (i.e., the initial feature representation) corresponding to the stroke node: ; Based on this, and combining stroke topology preservation constraints, mechanical consistency constraints, cross-modal consistency constraints, and quality ranking constraints, the temporal feature representation of stroke nodes is obtained. Joint optimization is performed to enhance the ability to discriminate differences in handwriting quality while maintaining structural consistency and dynamic behavior patterns. The feature representation obtained after the above training can be denoted as the... Feature representation of each stroke node In a preferred embodiment, the... It can be represented by time series characteristics With structural aggregation representation Further fusion yields representations that can be obtained through splicing, weighted summation, or a combination of representations transformed by a mapping network. Feature Representation The main characteristics are as follows: 1) The dynamic behavior of the stroke node during the writing process, including trajectory changes, speed patterns, pressure changes, and rhythm stability; 2) The corresponding position of the stroke node in the character structure space and its relationship with adjacent structural units; 3) Whether the stroke node has any abnormal patterns related to writing quality, such as sudden changes in local speed, excessive pressure fluctuations, trajectory deviations, or abnormal local pauses; 4) The degree of consistency between the stroke node and the overall character structure.

[0078] For example, a comprehensive loss function is constructed based on stroke topology preservation constraints, mechanical consistency constraints, feature consistency constraints, and quality sensitivity constraints. The temporal feature representations of stroke nodes and the initial feature representations are then updated according to this comprehensive loss function to obtain the feature representations of the Chinese character writing process and the final temporal feature representation of each stroke node. (Comprehensive Loss Function) for: ; in, , , , , , All are preset weighting coefficients. For the occlusion loss function: ; in, Indicates a time window. This represents the x-coordinate of the pen tip after the masking operation. This represents the vertical coordinate of the pen tip after the masking operation. This indicates the pen tip pressure after the masking operation.

[0079] The parameters of the formulas and models used to calculate the feature representation are updated using gradient descent and other methods (such as the parameters in the functions for feature encoding of stroke nodes and edges). The stroke nodes and character structure graph are then substituted into the updated formulas to recalculate the temporal feature representation of the stroke nodes. Based on the temporal feature representation of the stroke nodes, the feature representation is recalculated again, and the value of the comprehensive loss function is calculated using the above comprehensive loss function expression. When the value of the comprehensive loss function is less than the preset value, the feature representation of the Chinese character writing process is output.

[0080] Through the joint training described above, the output feature representation simultaneously includes: dynamic behavior information of the writing process, glyph structure constraint information, stroke topological consistency information, and sensitivity to changes in writing quality. This feature representation serves as the core input for subsequent evaluation, thus forming a complete closed-loop technical path consisting of "structural modeling - alignment modeling - self-supervised learning - quality evaluation".

[0081] The third step is to evaluate based on feature representation to obtain the writing quality evaluation results.

[0082] For example, a handwriting quality evaluation model can be further constructed to perform discriminative calculations on the feature representations, thereby obtaining the handwriting quality evaluation results. Specifically, for the first... Feature representation of each stroke node The input features can be fed into a quality assessment network for feature transformation and discrimination. This quality assessment network can be a fully connected neural network, a multilayer perceptron, a convolutional neural network, a graph neural network, or other model structures capable of classification, regression, or multi-label prediction. In a preferred embodiment, the quality assessment network uses a multilayer fully connected network as the prediction head to map the input features layer by layer and output evaluation results related to handwriting quality. For stroke-level quality assessment, the feature representation of a single stroke node can be... Inputting the stroke-level quality prediction head yields the corresponding stroke quality output. The output can be either a single score or a multi-dimensional label vector, used to characterize the quality of the stroke in terms of beginning, movement, ending, speed control, pressure control, and stability. If a multi-label evaluation is used, it can be represented as: ; in, This represents a stroke-level quality prediction model. Indicates the first The quality evaluation results for each stroke node. For overall structural quality evaluation, the feature representations of all stroke nodes can be analyzed. Further aggregation is performed to form a holistic character-level feature representation. The aggregation method can be average pooling, weighted summation, attention aggregation, or graph-level readout operation. Subsequently, the overall character-level feature representation is input into the overall quality prediction head to obtain the overall handwriting quality evaluation result. ; ; in, Represents the feature aggregation function, This represents the overall quality prediction model. This indicates the overall writing quality evaluation result of the Chinese character sample.

[0083] In a preferred embodiment, the handwriting quality evaluation result includes at least one of the following forms: overall quality score; stroke-level quality label; structure-level quality label; multi-label quality evaluation result.

[0084] It is worth mentioning that evaluating Chinese character writing based on the written result image and pen tip data takes into account both the written result image and the pen tip motion information during the writing process. This allows the dynamic behavioral characteristics during the writing process to be mapped to the character structure space, enabling precise localization of writing quality issues within the character shape. This is beneficial for improving the interpretability of writing quality evaluation. Furthermore, evaluating Chinese character writing quality based on multimodal data can combine information from both image and pen tip data, improving information richness and multifaceted evaluation, and effectively enhancing the reliability of Chinese character writing quality evaluation.

[0085] Furthermore, this application has the following advantages: This application constructs a directed acyclic topology of strokes to structurally model the sequential relationship between strokes during the writing process. This can more realistically and accurately reflect the actual writing order and the organizational form of writing behavior, thereby avoiding the problem of missing stroke order information caused by relying solely on static character shape analysis.

[0086] This application spatially registers the writing trajectory with the offline character structure and establishes a mapping relationship between stroke nodes and character structure, enabling dynamic behavioral characteristics during the writing process to be mapped to the character structure space. This allows for the precise location of writing quality issues within the character structure, which is beneficial for improving the relevance and practicality of the evaluation results.

[0087] This application introduces a self-supervised learning mechanism under structural constraints. Through methods such as occlusion reconstruction, cross-modal consistency constraints, and quality-sensitive learning, it can effectively extract feature representations that are sensitive to changes in handwriting quality even when there is a lack of or only a small amount of manually labeled data. This reduces the cost of manual labeling and improves the model's generalization ability in different populations and application scenarios.

[0088] This application can simultaneously output stroke-level quality evaluation results and overall structure-level quality evaluation results, and realize the correlation analysis between local problems and overall problems through structural mapping relationship. This allows the evaluation results to not only reflect the quality of writing, but also clearly indicate the location and type of problems, thereby significantly improving the interpretability and feedback of writing quality evaluation results.

[0089] The following is an exemplary description of the multimodal Chinese character writing quality evaluation device provided in this application.

[0090] like Figure 2 As shown, this application provides a multimodal Chinese character handwriting quality evaluation device 200, which includes: The acquisition module 201 is used to acquire pen tip data at multiple time points during the writing of Chinese characters and to acquire images of the writing results. Define module 202, which is used to define multiple stroke nodes based on all pen tip data; The extraction module 203 is used to extract multiple structural nodes from the writing result image and construct a character structure graph based on all structural nodes; multiple nodes in the character structure graph correspond one-to-one with multiple structural nodes, and the edges between nodes represent the connection relationship between the two corresponding structural nodes. Module 204 is established to create an alignment model based on all stroke nodes and character structure diagrams; the alignment model is used to describe the correspondence between stroke nodes and structure nodes. Evaluation module 205 is used to extract feature representations of the Chinese character writing process based on the alignment model, and to evaluate based on the feature representations to obtain writing quality evaluation results.

[0091] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0092] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0093] like Figure 3 As shown, an embodiment of this application provides a terminal device, wherein the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 3 The diagram shows only one processor, a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 executes the computer program D102 to implement the steps in any of the above method embodiments.

[0094] Specifically, when the processor D100 executes the computer program D102, it acquires pen tip data at multiple time points during the Chinese character writing process and obtains the writing result image. Then, it defines multiple stroke nodes based on all pen tip data, extracts multiple structural nodes from the writing result image, and constructs a character structure diagram based on all structural nodes. Next, it establishes an alignment model based on all stroke nodes and the character structure diagram. Finally, it extracts feature representations of the Chinese character writing process based on the alignment model and evaluates the writing quality based on these feature representations to obtain the writing quality evaluation result. This evaluation of Chinese character writing based on the writing result image and pen tip data considers both the resulting image and the pen tip's motion information during writing. This allows the dynamic behavioral features during writing to be mapped to the character structure space, enabling precise localization of writing quality issues within the character structure. This improves the interpretability of the writing quality evaluation. Furthermore, evaluating Chinese character writing quality based on multimodal data combines information from both image and pen tip data, increasing information richness and evaluation multifacetedness, effectively enhancing the reliability of the Chinese character writing quality evaluation.

[0095] The processor D100 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0096] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may be an external storage device of the terminal device D10, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device D10. Furthermore, the memory D101 may include both internal and external storage units of the terminal device D10. The memory D101 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory D101 can also be used to temporarily store data that has been output or will be output.

[0097] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0098] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.

[0099] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a multimodal Chinese character writing quality evaluation method device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0100] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0101] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0102] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention.

Claims

1. A multi-modal Chinese character writing quality evaluation method, characterized in that, include: Acquire pen tip data at multiple time points during the writing process of Chinese characters, and obtain images of the writing results; Multiple stroke nodes are defined based on all pen tip data; Multiple structural nodes are extracted from the writing result image, and a character structure diagram is constructed based on all structural nodes; In the character structure diagram, multiple nodes correspond one-to-one with multiple structural nodes, and the edges between nodes represent the connection relationship between the two corresponding structural nodes. An alignment model is established based on all stroke nodes and the character structure diagram; the alignment model is used to describe the correspondence between stroke nodes and structure nodes; The feature representation of the Chinese character writing process is extracted based on the alignment model, and the writing quality evaluation result is obtained based on the feature representation.

2. The Chinese character writing quality evaluation method according to claim 1, characterized by, The definition of multiple stroke nodes based on all pen tip data includes: For each time point, the pen tip state at that time point is determined based on the pen tip data at that time point; the pen tip state is a contact state and a non-contact state. The time point corresponding to the contact state is used as the writing time point, and a stroke node is defined based on the pen tip data corresponding to every two adjacent writing time points.

3. The Chinese character writing quality evaluation method according to claim 2, characterized by, Determining the pen tip state at the time point based on the pen tip data at the time point includes: Through the formula: ; calculating the state of the nib at a point in time ; wherein, represents a non-contact state of the pen tip, represents a contact state of the pen tip, represents a pen tip state at a time point, represents a pen tip pressure in the pen tip data at a time point, represents a pen down threshold value, represents a pen up threshold value, , represents a number of time points; The stroke node is: ; ; ; wherein, denotes a pair of the th stroke node, denotes a sequence vector of the th stroke node, denotes an attribute vector of the th stroke node, denotes a pen tip two-dimensional coordinate in pen tip data of the th time point, is the horizontal coordinate, is the vertical coordinate, denotes the th time point, denotes a writing time point corresponding to the th stroke node, denotes another writing time point corresponding to the th stroke node, denotes a duration of the th stroke node, denotes a trajectory length of the th stroke node, denotes an average pressure of the th stroke node, denotes an average pen tip speed of the th stroke node, denotes a pen tip speed standard deviation of the th stroke node, , denotes the number of stroke nodes.

4. The Chinese character writing quality evaluation method according to claim 3, characterized by, The alignment model established based on all stroke nodes and the character structure diagram includes: For each stroke node, calculate the coverage probability between the stroke node and each edge in the character structure graph; Align all stroke nodes and all edges in the glyph structure graph according to all coverage probabilities to obtain the alignment distribution between each stroke node and each edge. By integrating all alignment distributions into a single dataset, we obtain the alignment model.

5. The Chinese character writing quality evaluation method according to claim 4, characterized by, The calculation of the coverage probability between the stroke node and each edge in the character structure graph includes: Through the formula: ; Computing a coverage probability between a first stroke node and a second edge Computing a coverage probability between a first stroke node and a second edge Computing a coverage probability between a first stroke node and a second edge ​ wherein, denotes the average distance between the th stroke node and the th edge, is a distance smoothing coefficient, denotes the average distance between the th stroke node and the th edge, denotes the set of edge numbers in the glyph structure graph.

6. The Chinese character writing quality evaluation method according to claim 5, characterized by, The step of aligning all stroke nodes and all edges in the glyph structure graph according to all coverage probabilities to obtain the alignment distribution between each stroke node and each edge includes: Each stroke node is feature-encoded to obtain a temporal feature representation of each stroke node; Each edge in the character structure graph is feature-encoded to obtain a structural feature representation of each edge; For each stroke node, based on the temporal feature representation of the stroke node and the structural feature representation of each edge, the initial alignment distribution between the stroke node and each edge is calculated; Calculate the loss function value based on all coverage probabilities and all initial alignment distributions; The initial alignment distribution is updated based on the loss function value to obtain the alignment distribution between each stroke node and each edge.

7. The Chinese character writing quality evaluation method according to claim 6, characterized by, The calculation of the initial alignment distribution between the stroke nodes and each edge, based on the temporal feature representation of the stroke nodes and the structural feature representation of each edge, includes: Through the formula: ; ; computing an initial alignment distribution between the first stroke node and the second edge ;​​ wherein, represents an alignment score between the th stroke node and the th edge, represents an alignment score between the th stroke node and the th edge, represents a timing feature representation of the th stroke node, represents a structural feature representation of the th edge; The calculation of the loss function value based on all coverage probabilities and all initial alignment distributions includes: Through the formula: ; Calculate the loss function value ; in, Indicates the first A vector consisting of all coverage probabilities corresponding to each stroke node. Indicates the first The vector consisting of all the initial alignment distributions corresponding to each stroke node.

8. The method for evaluating the quality of Chinese character writing according to claim 1, characterized in that, The step of extracting the feature representation of the Chinese character writing process based on the alignment model includes: Based on the alignment model, stroke topology preservation constraints, mechanical consistency constraints, feature consistency constraints, and quality sensitivity constraints are constructed. Under the constraints of stroke topology preservation, mechanical consistency, feature consistency, and quality sensitivity, feature extraction is performed on the Chinese character writing process to obtain the feature representation of the Chinese character writing process.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the multimodal Chinese character writing quality evaluation method as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the multimodal Chinese character writing quality evaluation method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Chinese character writing quality intelligent evaluation method

    CN111340810A

  • Writing quality evaluation method based on Internet

    CN112634262A

  • Kinematic and morpometric analysis of digitized handwriting tracings

    US20170109566A1