A calligraphy single character intelligent evaluation method based on large model fine tuning

CN122821569APending Publication Date: 2026-09-25UNIV OF SHANGHAI FOR SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610923579.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,现有的通用多模态模型缺乏针对书法领域的专门训练,难以直接用于书法评价任务

Benefits of technology

1、本发明的评价方法,在该智能化书法评价方法中,主要包含用户终端与云端服务两大主体。用户通过终端采集并上传书法作品图像,系统在不依赖人工专家参与的前提下,自动完成书法字检测定位与区域选择;云端基于自建书法数据集,通过LoRA方式微调大模型,并使用微调后的大模型对书法字进行评价。在无需人工干预、无需人工标注打分的情况下,对单字的字形、笔法、构图进行结构化评价,并输出星级、优缺点及综合评价。实验与应用表明,所提出的书法评价框架可高效应用于移动端与小程序环境,能够实现快速准确的自动评价,于书法领域而言其使用场景具有一定的广泛性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821569A_ABST
    Figure CN122821569A_ABST
Patent Text Reader

Abstract

The application discloses a calligraphy single character intelligent evaluation method based on large model fine tuning. The application provides an automatic intelligent evaluation scheme in view of the problems that the existing calligraphy evaluation relies on artificial experience, the evaluation dimension is single, and the evaluation lacks interpretability. The method constructs a calligraphy data set containing calligraphy single character images and structured evaluation labels, and fine tunes a pre-trained multi-modal large model by using a low-rank adaptation method. After a user uploads a calligraphy image to be evaluated, a single character positioning model detects the boundary box of each single character and cuts each single character. The single character image is input into the fine-tuned multi-modal large model, and evaluation text containing character shape, stroke, composition and comprehensive evaluation is output. Finally, the structured analysis module extracts the star rating of each dimension, the advantage description and the improvement suggestion, and generates a structured evaluation result. The application realizes multi-dimensional and interpretable automatic evaluation of calligraphy single characters, and can be widely applied to calligraphy teaching assistance and work review scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and artificial intelligence, specifically to an intelligent evaluation method for single characters in calligraphy based on large model fine-tuning. Background Technology

[0002] Calligraphy, as an important component of traditional Chinese art, is traditionally evaluated by professional calligraphy teachers who make comprehensive judgments based on multiple dimensions, including character structure, brushwork, and overall composition. However, this process is often time-consuming and highly dependent on the evaluator's professional level, requiring them to possess profound calligraphic skills and rich experience in evaluation. Consequently, traditional evaluation methods struggle to ensure consistency of judgment standards when faced with different evaluators, and are also difficult to implement efficiently in large-scale teaching or large-scale artwork review.

[0003] Existing automated calligraphy evaluation methods are mainly based on traditional image processing techniques or small-scale deep learning models. They typically only analyze the structure of characters, lacking a comprehensive understanding of brushstroke details and overall composition. Consequently, they struggle to evaluate individual characters across multiple dimensions, resulting in a lack of holistic assessment. Furthermore, most of these methods primarily output numerical scores, lacking explanatory descriptions of the evaluation results and failing to generate evaluation texts with pedagogical guidance value. Therefore, they fall short of meeting the practical needs of calligraphy teaching feedback and users' intuitive understanding.

[0004] In recent years, with the rapid development of deep learning and large-scale pre-training techniques, multimodal large models have made significant progress in image understanding and text generation, demonstrating a remarkable ability to organically integrate visual information with linguistic expression. These models can process both image and text data simultaneously, exhibiting good performance in tasks such as cross-modal alignment, content description, and quality assessment. However, existing general-purpose multimodal models lack specialized training for the calligraphy domain, making them difficult to directly apply to calligraphy evaluation tasks. Furthermore, existing methods typically lack a unified format for output results, hindering subsequent parsing and user understanding.

[0005] Therefore, how to utilize multimodal large models to achieve multidimensional, structured, and interpretable evaluation of individual calligraphic characters has become an urgent technical problem to be solved. Summary of the Invention

[0006] Purpose of the invention: To address the problems existing in the prior art, this invention proposes an intelligent evaluation method for calligraphy characters based on large model fine-tuning. By constructing a structured labeled dataset of single-character images and corresponding evaluations and efficiently fine-tuning a multimodal large model, the method achieves automatic scoring and evaluation of calligraphy characters in multiple dimensions such as character shape, brushstrokes, and composition.

[0007] Technical solution: The intelligent evaluation method for calligraphy single characters based on large model fine-tuning of the present invention has the following specific steps: Step 1: Construct a calligraphy dataset for model training by collecting publicly available calligraphy copybooks and creating a self-built dataset. This dataset contains no fewer than 5,000 different calligraphy character images, along with structured evaluation annotations for each character image. The annotations are constructed by at least seven reviewers with calligraphy backgrounds, who independently score and comment on each character to improve the objectivity and reliability of the data annotations. The evaluation information is uniformly constructed using a pre-defined structured evaluation format, including character shape evaluation units, brushstroke evaluation units, composition evaluation units, and an overall evaluation unit. Step 2: Based on the calligraphy dataset, use the Low-Rank Adaptation (LoRA) method to train the pre-trained multimodal large model. Fine-tuning was performed by introducing low-rank incremental parameters into the model weights. The weight matrix of the multimodal large model is adjusted to obtain a fine-tuned large model. Its internal weight matrix is ​​updated as follows: ,

[0008] in, This represents the weight matrix of the original pre-trained large model, with low-rank increment parameters. Composed of two low-rank matrices and Multiplication constitutes the matrix. This is a dimension reduction projection matrix, whose function is to map the original features from a high-dimensional space to a low-rank subspace; the matrix This is a matrix used for dimension upscaling and reconstruction, used to restore the features of the low-rank subspace to the original weight space. Through this decomposition method, only the following steps are needed: and By updating a small number of parameters, the weight matrix can be modified. Fine-tuning. and satisfy:

[0009] in, The rank of the low-rank decomposition is used to control the size of the new parameters and the model complexity. Step 3: Input the entire calligraphy image The image is fed into a pre-trained character detection model to perform character detection and localization operations, thereby obtaining the calligraphy image. The bounding boxes corresponding to all individual calligraphic characters in the text. :

[0010] in, The number of calligraphy characters detected. and Representing the first The coordinates of the top left and bottom right corners of each single character area; Step 4: Based on the bounding box coordinate set described in Step 3 For the original image After cropping, a collection of calligraphy character images is obtained:

[0011] Among them, the Single character image The dimensions are:

[0012]

[0013] in, and Representing images respectively Length and width; Step 5: Extract the calligraphy character images cropped in Step 4. Input into the fine-tuning large model described in step 2 middle, The input single-character image is analyzed from multiple dimensions, and the output is the original evaluation text that includes character shape, strokes, composition, and comprehensive evaluation. ; Step 6: Extract the original evaluation text from Step 5. The input is fed into the evaluation parsing module, which performs paragraph segmentation and semantic localization through regularization matching to obtain the text subsets corresponding to each evaluation dimension.

[0014] in , , ,and These represent the textual content corresponding to the character shape, brushstrokes, composition, and overall evaluation, respectively. This represents a paragraph splitting function based on preset keywords and regular expression rules; Based on preset evaluation rules and keyword matching strategies, in the text segments corresponding to each evaluation dimension... (in Within this section, further extraction operations are performed on the "Advantages Description" and "Improvement Suggestions" fields. The extracted text content is then processed, categorized, and summarized using the following formula:

[0015]

[0016] in, This represents a set of advantages. This represents a set of suggestions for improvement. This represents the star rating for the corresponding dimension. This represents an information extraction function based on preset keywords and text classification rules. The analyzed character shapes, strokes, and structures Figure 3 The star ratings from each dimension, descriptions of strengths, suggestions for improvement, and the overall evaluation content are integrated to generate a structured evaluation result. ; Compared with the prior art, the present invention has the following advantages: 1. The evaluation method of this invention, an intelligent calligraphy evaluation method, mainly comprises two components: a user terminal and a cloud service. Users collect and upload images of calligraphy works through the terminal. The system automatically completes the detection, localization, and region selection of calligraphy characters without relying on human experts. The cloud service, based on a self-built calligraphy dataset, fine-tunes a large model using LoRA and uses the fine-tuned model to evaluate the calligraphy characters. Without human intervention or manual annotation and scoring, the system performs a structured evaluation of the shape, brushstrokes, and composition of individual characters, outputting star ratings, strengths and weaknesses, and a comprehensive evaluation. Experiments and applications show that the proposed calligraphy evaluation framework can be efficiently applied to mobile and mini-program environments, achieving fast and accurate automatic evaluation, and its application scenarios in the field of calligraphy are quite broad.

[0017] 2. This invention, by introducing a multimodal large model for fine-tuning training, achieves comprehensive analytical capabilities for calligraphy characters across multiple dimensions, including character shape, brushstrokes, and composition. This overcomes the limitations of traditional methods that evaluate only a single feature, resulting in more comprehensive and accurate evaluation results. Furthermore, by constructing a unified structured annotation template and guiding the model to generate corresponding evaluation text, the output not only includes quantitative scores but also descriptions of strengths and suggestions for improvement, significantly enhancing the interpretability and practical value of the evaluation results.

[0018] 3. In terms of model training, the present invention uses a low-rank adaptation method to fine-tune the multimodal large model, and only updates a small number of low-rank parameters. While effectively reducing the scale of training parameters and computational complexity, the general capability of the original model is fully retained, thereby improving training stability and model generalization performance. In addition, the present invention further provides an evaluation and parsing module, which converts the natural language evaluation text generated by the model into structured data results. Such structured output facilitates subsequent data storage, statistical analysis and visual display, thereby significantly improving the engineering application capability and practical value of the system. Description of Drawings

[0019] Figure 1 is a schematic diagram of the overall process flow of the intelligent calligraphy single-character evaluation method based on large model fine-tuning in an embodiment of the present invention; Figure 2 is an example of a training dataset in the intelligent calligraphy single-character evaluation method based on large model fine-tuning in an embodiment of the present invention; Figure 3 is a schematic diagram of large model fine-tuning in the intelligent calligraphy single-character evaluation method based on large model fine-tuning in an embodiment of the present invention; Figure 4 is a schematic diagram of the generation process of calligraphy single-character images in the intelligent calligraphy single-character evaluation method based on large model fine-tuning in an embodiment of the present invention; Figure 5 shows the image of the calligraphy character "lai" used in an embodiment of the present invention and the original evaluation text output by the large model; Figure 6 is the structured evaluation text formed after processing the original evaluation text in an embodiment of the present invention. Detailed Description of Embodiments

[0020] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments specifically elaborate the intelligent calligraphy single-character evaluation method based on large model fine-tuning involved in the present invention with reference to the accompanying drawings.

[0021] The method of this embodiment comprises the following steps: Step 1: The present invention first constructs a calligraphy dataset for model training, said dataset comprising no less than 5000 different calligraphy single-character images, and structured evaluation annotation information corresponding to each calligraphy single-character image. Said annotation information is constructed in a unified evaluation format, which specifically includes a glyph evaluation unit, a brushstroke evaluation unit, a composition evaluation unit and an overall evaluation unit. Wherein, each evaluation unit further comprises three contents: a star rating, a description of advantages and suggestions for improvement, the specific example is Figure 2As shown. To ensure the professionalism and consistency of the annotation information, the structured evaluation annotation information is constructed by no fewer than seven reviewers with calligraphy professional backgrounds, each independently scoring and commenting on each calligraphy character. After all reviewers have completed their independent annotations, the multiple annotation results are integrated through consistency processing to generate the final annotation result for each calligraphy character.

[0022] Step 2: Employ a pre-trained multimodal large model As the base model, and by fine-tuning its parameter matrix using the Low-Rank Adaptation (LoRA) method, the overall process is as follows: Figure 3 As shown. Specifically, we have the model weight matrix Introduce a trainable low-rank increment parameter. To simulate the weight matrix The update amount during the fine-tuning process yields the fine-tuned weight matrix. ,satisfy:

[0023]

[0024] in For the dimension-reduced projection matrix, To reconstruct the matrix in higher dimensions, and Freeze the original weights during training. Only for matrices and Parameters are updated to achieve efficient adaptation to calligraphy evaluation tasks. This is done for the pre-trained model. After training is completed, a large trained model is obtained. .

[0025] Step 3: Obtain the calligraphy images to be evaluated uploaded or collected by the user, denoted as... Next, the calligraphy images... Processing is performed to extract Each calligraphy character image in the process is as follows: Figure 4 As shown. Specifically, the image Inputting a single-character localization model, the model detects and locates single characters in calligraphy images, obtaining a set of bounding box coordinates. :

[0026] in, The number of calligraphy characters detected. and Representing the first The coordinates of the top left and bottom right corners of each character area.

[0027] Step 4: Through For images Cropping is performed to obtain a collection of calligraphy character images. :

[0028] Among them, the Single character image height With width They are respectively:

[0029]

[0030] Step 5: Extract the single-character image With preset evaluation prompts Input them together into the fine-tuned multimodal large model , The input single-character image is analyzed from multiple dimensions, and the output is a raw evaluation text that includes character shape, strokes, composition, and comprehensive evaluation. .by Figure 5 For example, single-character images With the obtained evaluation text output like Figure 5 As shown.

[0031] Step 6: Obtain the original evaluation text Afterwards, Perform structured analysis to generate structured evaluation results. ,by Figure 5 Taking the results as an example, the generated structured evaluation results are as follows: Figure 6 As shown. The structured evaluation results are presented. The data is packaged according to a preset data format and output to the user terminal to realize the evaluation, display and feedback of individual calligraphy characters.

[0032] The role and effects of the embodiments: Most existing calligraphy evaluation technologies rely on manual experience or automatic analysis methods based on single image features to perform simple quantitative evaluation on calligraphy works, but such methods usually have problems such as single evaluation dimension and difficulty in generating interpretable evaluation results. Meanwhile, traditional methods are difficult to perform fast inference, which limits their promotion in actual teaching and application scenarios. In this embodiment, a large multimodal model is introduced and fine-tuned with structured annotation data, enabling the model to process both calligraphy image information and evaluation text information. Under the constraint of a unified evaluation template, comprehensive analysis and automatic evaluation are performed on individual calligraphy characters from multiple dimensions such as character shape, brushwork and composition, and structured evaluation results are generated. Further, the entire calligraphy image is automatically segmented through a single-character localization module, and fine-grained processing is performed with single characters as the basic unit, thereby improving the accuracy and granularity of the evaluation.

[0033] In the embodiment, first, an example calligraphy image containing multiple handwritten Chinese characters is input into a pre-trained target detection model for single-character localization processing, and the bounding box coordinates of each individual calligraphy character are obtained, including the position information of the top-left corner and the bottom-right corner. According to the obtained bounding box coordinates of each individual calligraphy character, character-by-character cropping is performed on the original calligraphy image, and the sub-image region corresponding to each single character is extracted, so as to obtain a series of independent calligraphy single-character images, facilitating subsequent character-by-character analysis and evaluation. Subsequently, taking the character "lai" obtained after cropping as an example, this single-character image and preset evaluation prompt information are input together into the fine-tuned large multimodal model for processing. After comprehensively analyzing the input information, the large multimodal model outputs a result containing character shape, brushwork and compos Figure 3 tion, star ratings for multiple dimensions as well as an evaluation text with corresponding advantage descriptions and suggestions for improvement. Finally, regularization extraction and segmentation are performed to carry out regularization processing and segmentation extraction on the evaluation text, forming a structured calligraphy single-character evaluation result.

[0034] Compared with existing calligraphy evaluation methods, the method of this embodiment realizes the transformation from single scoring to "multi-dimensional scoring + structured text evaluation" by fine-tuning the large multimodal model, making the evaluation results more comprehensive and interpretable. By introducing a unified structured annotation template and an evaluation parsing mechanism, the output results have good standardization and parsability, which is convenient for subsequent data processing and application. In addition, the present invention adopts the low-rank adaptation method for model fine-tuning, which maintains model performance while reducing training costs, improves the practicability and scalability of the system, and thus has good application prospects in the field of calligraphy teaching assistance and automatic evaluation.

[0035] The above embodiment is a preferred embodiment of the present invention, and is not used to limit the protection scope of the present invention.

Claims

1. A method for intelligent evaluation of single characters in calligraphy based on large model fine-tuning, characterized in that, include: S1. Construct a self-built calligraphy dataset, which includes multiple calligraphy character images and structured evaluation annotation information corresponding to each calligraphy character image. The structured evaluation annotation information includes character shape evaluation unit, brushstroke evaluation unit, composition evaluation unit and overall evaluation unit. Use the low-rank adaptation method to fine-tune the pre-trained multimodal large model using the calligraphy dataset to obtain the fine-tuned multimodal large model. S2. Obtain the calligraphy image to be evaluated, input the calligraphy image into the single character localization model, and obtain the set of bounding box coordinates of each calligraphy character in the calligraphy image; S3. Cropping the calligraphy image according to the bounding box coordinate set to obtain the single character image corresponding to each calligraphy character; S4. Input the single character image into the fine-tuned multimodal large model. The fine-tuned multimodal large model performs multi-dimensional analysis on the single character image and outputs evaluation text that includes character shape evaluation, stroke evaluation, composition evaluation and comprehensive evaluation. S5. Perform structured parsing on the evaluation text, extract evaluation information and overall evaluation information in three dimensions: character shape, stroke technique, and composition, generate structured evaluation results and output them.

2. The method according to claim 1, characterized in that, In step S1, the structured evaluation annotation information is uniformly constructed using a preset structured evaluation format; the character shape evaluation unit includes fine-grained scoring items for the proportional coordination of radicals and the relationship of stroke interweaving and avoidance; the brushwork evaluation unit includes fine-grained scoring items for the changes in strength and the sharpness of the starting, middle, and ending strokes; the composition evaluation unit includes fine-grained scoring items for the positional deviation of characters in the writing area, the balance of margins and white space, and the overall visual balance; the character shape evaluation unit, brushwork evaluation unit, and composition evaluation unit all include a star rating field, an advantage description field, and a suggestion field for improvement field; the overall evaluation unit weights and merges the results of the character shape evaluation, brushwork evaluation, and composition evaluation according to preset weight coefficients, and outputs a comprehensive score and grade comments.

3. The method according to claim 1, characterized in that, In step S1, a low-rank adaptation method is used to fine-tune the pre-trained multimodal large model, specifically including: adjusting the weight matrix of the pre-trained multimodal large model. W Introducing low-rank incremental parameters ΔW The fine-tuned weight matrix is ​​obtained. ,in , A For the dimension-reduced projection matrix, B To reconstruct the matrix in higher dimensions, and A and B rank r Much smaller than the weight matrix W The dimension where the original weights are frozen during training. W Only for matrices A and B Update the parameters.

4. The method according to claim 1, characterized in that, In step S2, the bounding box coordinate set is represented as follows: ; in, The number of calligraphy characters detected. and Representing the first The coordinates of the top left and bottom right corners of each character area.

5. The method according to claim 1, characterized in that, In step S3, the first The height of a single character image ,width .

6. The method according to claim 1, characterized in that, In step S4, the evaluation text includes a character shape evaluation text segment, a brushstroke evaluation text segment, a composition evaluation text segment, and an overall evaluation text segment; in step S5, the structured parsing of the evaluation text specifically includes: dividing the evaluation text into text subsets corresponding to each evaluation dimension through paragraph segmentation, and extracting star ratings, descriptions of advantages, and suggestions for improvement within the text subsets corresponding to each evaluation dimension.

7. The method according to claim 1, characterized in that, In step S5, the structured evaluation result includes star ratings for three dimensions: character shape, stroke technique, and composition; descriptions of advantages; suggestions for improvement; and an overall evaluation.

8. The method according to claim 1, characterized in that, In step S1, the calligraphy dataset contains no fewer than 5,000 images of individual calligraphy characters.

9. The method according to claim 1, characterized in that, In step S1, the structured evaluation annotation information is generated by multiple reviewers with calligraphy backgrounds who independently score and comment on each calligraphy character, and then integrated through expert review and consistency processing.

10. The method according to claim 1, characterized in that, In step S4, the single-character image and the preset evaluation prompt information are input into the fine-tuned multimodal large model. The preset evaluation prompt information includes a format constraint instruction, which is used to constrain the fine-tuned multimodal large model to output evaluation text with a preset data structure.