Calligraphy work intelligent preliminary screening system and method based on multi-mode large model collaboration and knowledge graph

Through a collaborative system of multimodal large models and knowledge graphs, efficient, transparent and objective initial screening of calligraphy works is achieved, solving the problems of low efficiency and strong subjectivity in manual initial screening in existing technologies. It generates highly interpretable review results and is suitable for high-throughput review scenarios.

CN121938017APending Publication Date: 2026-04-28SHANGHAI UNIV OF ENG SCI +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI UNIV OF ENG SCI
Filing Date
2026-01-19
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing methods for initial screening of calligraphy works are highly subjective, inefficient, and provide vague feedback, making it difficult to meet the needs of high-throughput review. Furthermore, they are costly in terms of manpower and lack unified quantitative criteria and transparency.

Method used

The intelligent initial screening system, which adopts multimodal large model collaboration and knowledge graph, receives image and text data through a multimodal input module, performs quantitative scoring through a multi-model collaborative scoring module, generates interpretability evaluation by combining knowledge graph reasoning, and automatically filters or sends it to manual review through a divergence detection and initial screening decision module.

Benefits of technology

It has achieved automation and intelligence in the initial screening of high-throughput calligraphy works, significantly improving the efficiency and objectivity of the initial screening, generating interpretable evaluations with cultural basis, reducing the burden of manual review, and enhancing the transparency and credibility of the review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121938017A_ABST
    Figure CN121938017A_ABST
Patent Text Reader

Abstract

The invention discloses a calligraphy work intelligent preliminary screening system and method based on multi-mode large model collaboration and a knowledge graph. The system comprises a multi-modal input module used for receiving a work image and a text paraphrase; the multi-model collaborative scoring module is used for calling at least two heterogeneous multi-modal large models and independently outputting quantitative scores from four dimensions of a writing method, a knot body, a chapter method and an ink method; the knowledge graph reasoning module is used for carrying out semantic matching on the score based on the calligraphy field knowledge graph and generating interpretability evaluation; and the divergence detection and preliminary screening decision-making module is used for judging expert re-checking or automatically entering the next round by calculating a scoring standard deviation between models. The corresponding method comprises the following steps: acquiring multi-modal input; performing four-dimensional scoring by cooperating with multiple models; generating interpretability evaluation based on the knowledge graph; and primary screening is completed through bifurcation detection and mean value judgment. Through multi-model cooperation and knowledge reasoning, multi-expert review is effectively simulated, and the objectivity, efficiency and result interpretability of calligraphy review are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and artificial intelligence, and in particular to an intelligent preliminary screening system and method for calligraphy works based on multimodal large model collaboration and knowledge graph. Background Technology

[0002] In current high-throughput judging scenarios such as calligraphy exhibitions, competitions, and grading assessments, the initial screening of calligraphy works mainly relies on human experts for evaluation. These experts typically select works based on traditional aesthetic dimensions such as brushwork, structure, composition, and ink application, combined with their personal experience and aesthetic preferences. However, this human-centric initial screening method has certain limitations.

[0003] First, manual preliminary screening is easily influenced by factors such as expert stylistic preferences, daily condition, and fatigue, leading to inconsistent scoring standards. For example, a judge specializing in cursive script might give a lower evaluation to a regular script work, and differing understandings of elements like "composition density" among experts can also cause inconsistencies in preliminary screening results, thus affecting the fairness of the selection process. Second, with thousands or even tens of thousands of submissions, manual preliminary screening requires a large number of professionals and significant time. For example, a national calligraphy exhibition receives over ten thousand submissions annually, and the preliminary screening stage often requires more than 20 experts to work continuously for one to two weeks, resulting in high labor costs and difficulty in meeting the requirements for a rapid review cycle. Third, while traditional manual preliminary screening has the advantage of cultural understanding, it often lacks unified quantitative criteria, making it difficult to clearly explain to submitters the reasons for not advancing to the next round of selection, such as "overall shortcomings" or "insufficient artistic conception." In addition, for atypical works that blend tradition and innovation, some experts may make misjudgments due to cognitive limitations. In summary, existing manual preliminary screening methods for calligraphy works suffer from problems such as strong subjectivity, low efficiency, and ambiguous feedback. There is an urgent need for an intelligent, high-throughput preliminary screening method that can simulate collaborative evaluation by multiple experts and incorporate professional calligraphy knowledge, so as to significantly improve the efficiency, transparency, and fairness of preliminary screening while ensuring the quality of the review. Summary of the Invention

[0004] This invention aims to overcome the shortcomings of the prior art by providing an intelligent preliminary screening system and method for calligraphy works based on multimodal large-scale model collaboration and knowledge graphs. Through multimodal large-scale models, calligraphy domain knowledge graphs, and human-machine collaborative decision-making schemes, calligraphy works are subjected to multi-dimensional quantitative evaluation and high-throughput preliminary screening. This aims to reduce the workload of manual review and improve the objectivity and interpretability of the preliminary screening process in scenarios such as calligraphy exhibitions, competitions, or grading assessments.

[0005] This invention provides an intelligent preliminary screening system for calligraphy works based on multimodal large-scale model collaboration and knowledge graphs. It comprises: a multimodal input module for receiving image data and corresponding textual explanations of the calligraphy works to be evaluated; a multimodal collaborative scoring module connected to the multimodal input module for calling at least two multimodal large-scale models with differences in model architecture or training objectives to independently analyze the works from four dimensions: brushwork, structure, composition, and ink technique, and output quantitative scores; a knowledge graph reasoning module connected to the multimodal collaborative scoring module for semantic matching of the scoring results based on a pre-constructed knowledge graph in the calligraphy domain, generating interpretable evaluations with cultural basis; and a divergence detection and preliminary screening decision module connected to the knowledge graph reasoning module and the multimodal collaborative scoring module for calculating the standard deviation between the scores of each model. If the difference exceeds a preset divergence threshold, the work is marked as a disputed sample and sent to a queue of human experts for review; otherwise, it automatically determines whether to proceed to the next round of evaluation based on the average score.

[0006] Preferably, the multimodal input module further includes a preprocessing unit for cropping borders, correcting perspective distortion, adjusting white balance, and enhancing contrast of the image; and for using an OCR engine to extract or receive text translations submitted by the user from the image to form a multimodal input pair corresponding to the image and the text.

[0007] Preferably, the multi-model collaborative scoring module calls at least two multimodal large models, including a general visual language model and a domain-specific model fine-tuned on a calligraphy rubbings dataset.

[0008] Preferably, the knowledge graph reasoning module is specifically used to: semantically match the visual and semantic features extracted by the multi-model collaborative scoring module with entities and relationships in the knowledge graph to generate interpretability evaluation. The knowledge graph is constructed using a graph database or a "subject-verb-object" triple format, and its nodes include historical rubbings, calligraphic styles, aesthetic theories, and technical specifications.

[0009] Preferably, the preset divergence threshold is dynamically adjusted according to the exhibition / competition level or type.

[0010] This invention also provides an intelligent preliminary screening method for calligraphy works based on multimodal large-scale model collaboration and knowledge graph, characterized by the following steps: Step S1, acquiring image data and corresponding text interpretations of the calligraphy works to be evaluated as multimodal inputs; Step S2, feeding the multimodal inputs into at least two multimodal large-scale models that differ in model architecture or training objectives, with each model independently analyzing the works from four dimensions: brushwork, structure, composition, and ink technique, and outputting quantitative scores; Step S3, semantically interpreting the scoring results of Step S2 based on a pre-constructed knowledge graph in the calligraphy domain, generating interpretable evaluations with cultural basis; Step S4, calculating the standard deviation between the total scores output by each model, if the standard deviation is greater than a preset disagreement threshold, marking the work as a disputed sample and sending it to the human expert review queue; otherwise, automatically determining whether the work should enter the next round of evaluation based on the mean score and the judgment threshold, thus completing the preliminary screening.

[0011] Preferably, step S1 includes the following steps: step S1-1, acquiring digital images of calligraphy works using high-resolution professional equipment; step S1-2, preprocessing the digital images, and simultaneously extracting corresponding text interpretations from the images using an OCR engine to form multimodal input pairs corresponding to the images and text; step S1-3, packaging the preprocessed images and verified text interpretations into multimodal data units of a unified format, which are used as inputs for subsequent multimodal large models.

[0012] Preferably, step S2 includes the following steps: Step S2-1, calling at least two multimodal large models that differ in model architecture or training objectives as collaborative scoring engines; Step S2-2, inputting the multimodal data units generated in step S1 into each model, with each model independently extracting visual and semantic joint features, and performing structured scoring based on a preset four-dimensional calligraphy evaluation system; Step S2-3, each model outputs quantitative scores in four dimensions, which are then aggregated into a structured scoring vector.

[0013] Preferably, step S3 includes the following steps: Step S3-1, pre-constructing a knowledge graph in the calligraphy field, organizing structured knowledge in the form of "subject-verb-object" triples; Step S3-2, when the semantic similarity between the structured query vector and the knowledge graph entity reaches a threshold, they are considered semantically similar, triggering the subsequent comment generation process; Step S3-3, based on the matching results, calling natural language templates to generate interpretable evaluation statements with cultural context.

[0014] Preferably, step S4 includes the following steps: Step S4-1, normalize the four-dimensional scores output by each model to eliminate dimensional differences, and calculate the total score according to preset weights or equal weighting to form a comprehensive score result for each model; Step S4-2, calculate the standard deviation of the total scores of all models as a quantitative indicator of the degree of divergence between models. If the degree of divergence is greater than a preset divergence threshold, it is marked as a controversial sample and automatically added to the queue of human expert review; Step S4-3, for non-controversial works, calculate the average of their total scores and compare it with a preset threshold. If the average is higher than the threshold, the work is determined to enter the next round of selection.

[0015] Beneficial effects Compared with the prior art, the present invention has the following significant advantages: 1. It has outstanding advantages in high-throughput initial screening efficiency. Traditional manual initial screening requires a large amount of expert resources and takes a long time when faced with a large number of calligraphy submissions, which is difficult to meet the needs of rapid review. However, this invention realizes batch processing and intelligent sorting of works through an automated process, which can complete the first round of screening of thousands or even tens of thousands of works in a short time, and automatically identify disputed samples and send them to the review queue, which greatly shortens the overall review cycle and is suitable for high-concurrency scenarios such as exhibitions, competitions, and grade assessments.

[0016] 2. Significantly enhances the professionalism and objectivity of initial screening results. This invention introduces multiple multimodal large models with differences in architecture or training objectives to work collaboratively, effectively simulating the review process of multiple calligraphy experts and avoiding the bias of subjective preferences of a single model or individual expert on the initial screening results. At the same time, by combining a pre-created knowledge graph in the calligraphy field, the scoring is correlated with visual feature statistics, historical rubbings, calligraphic norms, and aesthetic theories, generating judgment results with cultural basis and significantly enhancing the professionalism of the evaluation.

[0017] 3. It has application value in cost control and review transparency. This invention relies only on conventional image acquisition equipment and software algorithms, requiring no dedicated hardware investment, significantly reducing system deployment and maintenance costs. In particular, the system can generate interpretable comments such as "the center of gravity of the structure is off, not conforming to the European style structural specifications" or "insufficient ink tones, failing to reflect traditional ink aesthetics," providing clear feedback for unselected works, enhancing the credibility and user acceptance of the initial screening process, and effectively alleviating doubts about the selection process. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0019] Figure 1 This is a schematic diagram of the architecture of the intelligent preliminary screening system for calligraphy works based on multimodal large model collaboration and knowledge graph in an embodiment of the present invention. Figure 2 This is a schematic diagram of the scoring discrepancy detection process of the multi-model collaborative scoring module in an embodiment of the present invention; Figure 3 This is a schematic diagram of the semantic matching process for generating interpretability evaluation of knowledge graph reasoning in an embodiment of the present invention. Detailed Implementation

[0020] The following describes a specific embodiment of the present invention in conjunction with a typical application scenario. It should be noted that this embodiment is only for the purpose of facilitating understanding of the technical solution of the present invention and does not constitute a limitation on the scope of protection of the present invention. Those skilled in the art can make adaptive adjustments to the following embodiments without departing from the concept of the present invention, and these modifications all fall within the scope of protection of the present invention.

[0021] Terminology Explanation: 1. Multimodal large model: refers to a deep learning model that can jointly encode and semantically align heterogeneous modal inputs such as images and text, and generate a unified cross-modal representation. A typical example is the Vision-Language Model (VLM).

[0022] 2. Knowledge Graph: A structured knowledge base organized in the form of "entity-relationship-entity" triples, used to represent concepts and their semantic relationships, supporting efficient relational reasoning and semantic retrieval.

[0023] 3. Brushwork, structure, composition, and ink technique: the four core dimensions of traditional calligraphy evaluation. Among them, brushwork refers to the strength, pressure, and central application of the brush; structure refers to the proportion and stability of the character's shape; composition refers to the overall layout and the continuity of the writing; and ink technique focuses on the rhythm and layering of ink density, dryness, and wetness.

[0024] 4. Disagreement Detection: By calculating the statistical differences (such as standard deviations) between the output scores of multiple models, it can be determined whether there are significant disagreements between the models, thereby identifying disputed samples that require human intervention.

[0025] 5. Explainable evaluation: This refers to not only providing scoring results, but also generating natural language feedback with cultural basis and professional context, such as "deviates from the stable characteristics of Ouyang Xun's style" or "does not reflect the ink aesthetics of 'moistening with spring rain and cracking with autumn wind'", in order to enhance the transparency and credibility of the system's decision-making.

[0026] Example 1: Intelligent Preliminary Screening Application in Provincial Youth Calligraphy Exhibition This embodiment uses the initial screening of submissions for a provincial-level youth calligraphy exhibition as an example to detail the system workflow and method steps of the present invention. The exhibition received several thousand calligraphy submissions from across the country, and the first round of screening needed to be completed within a limited timeframe, with a certain percentage of works advancing to the expert review stage. The system is deployed on a local server cluster, employing the method described in this embodiment to achieve high-throughput intelligent initial screening.

[0027] Figure 1 This is a schematic diagram of the architecture of the intelligent initial screening system for calligraphy works based on multimodal large model collaboration and knowledge graph in an embodiment of the present invention.

[0028] like Figure 1 As shown, the multimodal large model collaboration and knowledge graph-based intelligent initial screening system for calligraphy works in this embodiment includes the following modules: Multimodal input module: Receives image data and corresponding text interpretation of the calligraphy work to be evaluated, serving as the basis for subsequent multimodal input for evaluation.

[0029] In this embodiment, firstly, each submitted work is photographed in high definition using professional-grade digital equipment under standard lighting conditions, with the image resolution set to 300 dpi to ensure that the overall composition and brushstroke details are clearly discernible. Then, the acquired images are preprocessed, including automatic white border cropping, perspective correction, white balance adjustment, and contrast enhancement, to eliminate shooting distortion and lighting interference.

[0030] Simultaneously, the PaddleOCR engine is invoked to automatically recognize the text in the image, and low-quality recognition results are filtered out using a confidence threshold (≥0.9). For those below the threshold, staff provide assistance in correction, ultimately forming multimodal data units with a one-to-one correspondence between the image and the text. Each data unit contains a JPEG image (uniformly scaled to 512×512) and a UTF-8 encoded text string, which serve as input for subsequent models.

[0031] Multi-model collaborative scoring module: calls at least two multimodal large models that differ in architecture or training objectives to independently analyze the work from the four core dimensions of calligraphy: brushwork, structure, composition, and ink technique, and output quantitative scores, simulating the independent review process of multiple experts.

[0032] In this embodiment, the system deploys two heterogeneous multimodal large models as collaborative scoring engines: Model A uses the general visual language large model Qwen-VL, which has cross-domain image and text understanding capabilities; Model B uses a domain-specific model finely tuned on a calligraphy rubbings dataset, which enhances the sensitivity to traditional calligraphy features (such as flying white, dry brush, and the flow of energy). This combination is suitable for comprehensive calligraphy exhibitions and competitions, and can cover works of multiple calligraphic styles and genres, taking into account both versatility and professional depth.

[0033] The aforementioned multimodal data units are input into two models respectively. Each model independently extracts joint visual and semantic features and outputs a structured score (0~100 points) based on a pre-defined four-dimensional evaluation system: Brushwork dimension: Based on the analysis of the gradient of the stroke edge and the change of ink width, the rhythm of lifting and pressing is analyzed, and the quality of the use of the central tip is judged in combination with the shape of the beginning and ending strokes; Structural dimension: By detecting key points of the character skeleton, the aspect ratio and center of gravity offset (unit: pixels) of the character are calculated to evaluate the structural stability; Composition dimension: Utilize global attention weight distribution to quantify the standard deviation of line spacing (σ_line_spacing) and the line flow coherence index (values ​​range from 0 to 1). Ink technique dimension: Combining local contrast and grayscale histogram entropy value to measure the richness of tonal range in terms of density, lightness, dryness, and wetness.

[0034] To improve robustness, each model performs inference three times for the same work with different randomly cropped regions, and the average of the four-dimensional scores is taken as the final output of the model, forming a structured score vector: S A = [s A brushwork, s A , structure, s A Composition, s A , Mo method] S B =[s B brushwork, s B , structure, s B Composition, s B , Mo method] Knowledge Graph Reasoning Module: Constructs a knowledge graph in the field of calligraphy that covers rubbings from ancient steles, calligraphic styles, aesthetic theories, and technical norms. Matches the visual and semantic features extracted by the model with entities and relationships in the graph to generate interpretable evaluations with cultural basis, such as "the center of gravity of the structure shifts, deviating from the stable characteristics of the Ouyang Xun style."

[0035] In this embodiment, the pre-constructed knowledge graph of calligraphy is stored in the form of Neo4j graph database, covering classic rubbings such as "Preface to the Orchid Pavilion" and "Draft of a Eulogy for My Nephew", calligraphy styles such as Ouyang Xun, Yan Zhenqing and Zhao Mengfu, aesthetic concepts such as spirit, strength, and flying white, as well as technical terms such as central stroke, lifting, and dry brush. The relationship types include "belong to", "embody", "requirement", and "comparison".

[0036] The intermediate features output by the model (in this embodiment, the composition score σ_line_spacing = 18.7 px and the ink histogram entropy = 3.2) and the scoring deviation (in this embodiment, the composition score < 60) are transformed into structured query vectors, and semantic retrieval is performed in the graph. For example, when the system detects "low composition score and large fluctuation in line spacing", the following path is matched: "Wang Xizhi's Running Script" — embodiment → continuous flow of energy — requirement → close correspondence between lines.

[0037] Based on the matching results, a pre-defined natural language template library is invoked to dynamically generate interpretable comments. For example: if the ink entropy value is too low → "The ink color is dry and monotonous, failing to embody the ink aesthetics of 'moistening like spring rain, cracking like autumn wind'"; if the center of gravity of the structure is shifted to the right by more than 15 px → "The center of gravity of the structure is tilted to the right, which does not conform to the structural norm of 'centrality and stability' in Ouyang Xun's calligraphy style." All comments are accompanied by a graphical path tracing, supporting rapid location of cultural basis during manual review.

[0038] The discrepancy detection and initial screening decision module calculates the standard deviation between the scores of each model. If the difference exceeds a preset threshold, the work is marked as a disputed sample and sent to the human expert review queue. Otherwise, it is automatically determined whether to proceed to the next round based on the average score.

[0039] In this embodiment, firstly, the four-dimensional scores of each model are weighted and summed to obtain the overall total score. This embodiment adopts an equal-weighting strategy: T A =1 / 4(s A , brushstroke + s A , structure + s A Chapter + s A , Mo method) T B =1 / 4(s B , brushstroke + s B , structure + s B Chapter + s B , Mo method) Calculate the standard deviation of the total scores for both models: Set the divergence threshold τ div =8.0 (out of 100). If σ T >τ div If the sample is not found to be in dispute, it is automatically added to the manual review queue; otherwise, the mean value is used. As the final score.

[0040] Then set the initial screening threshold, such as τ. pass =65. If ≥τ pass If the initial screening is successful, the work will proceed to the next round; otherwise, it will be considered as failing the initial screening. The system will record the scoring log, four-dimensional sub-items, and interpretable comments for subsequent analysis, appeal review, or model optimization.

[0041] Through the above process, this embodiment realizes the automation, intelligence and interpretability of high-throughput initial screening of calligraphy works, which significantly reduces the burden of manual review while effectively ensuring the objectivity, fairness and professional depth of the review.

[0042] This embodiment also provides an intelligent initial screening method for calligraphy works based on multimodal large model collaboration and knowledge graph. This method corresponds to the system described above and includes the following steps: Step S1 is executed by the multimodal input module, acquiring the image data and corresponding text transcription of the calligraphy work to be evaluated as multimodal input. The image data comes from high-definition scanning or digital photography, and the text transcription can be automatically extracted through optical character recognition (OCR) technology or manually entered by the user. The specific implementation steps are as follows: Step S1-1: Acquire digital images of calligraphy works using a high-resolution professional-grade camera. The image acquisition resolution should be no less than 300 dpi, and the images should be taken in a uniform light source environment to avoid interference factors such as reflection, shadow, or distortion from affecting the analysis of the works.

[0043] Step S1-2 involves preprocessing the image acquired in step S1-1, including cropping borders, correcting perspective distortion, adjusting white balance, and enhancing contrast. Simultaneously, an OCR engine (such as PaddleOCR or TrOCR, PaddleOCR in this embodiment) is used to extract the corresponding text from the image, or the author's submitted text information is obtained directly from the work information. The text is then manually verified to ensure that the text content is consistent with the calligraphy work content, forming a reliable multimodal input pair, i.e., the corresponding image and text.

[0044] Steps S1-3: Package the preprocessed image and the verified text into a unified format multimodal data unit as input for the subsequent multimodal large model. The image can retain its original size or be scaled proportionally to the resolution required by the model input (e.g., 512×512 or 1024×1024, 512×512 in this embodiment). The text is embedded in the form of a UTF-8 encoded string to ensure that visual information and semantic information are effectively aligned during the model inference process.

[0045] Step S2 is executed by the multi-model collaborative scoring module. The multimodal input is fed into at least two large multimodal models that differ in architecture or training objectives. Each model independently analyzes the work from four dimensions: brushwork, structure, composition, and ink technique, and outputs a quantitative score. Among them, the brushwork dimension focuses on the strength of the brush, the variation of pressure and lifting, and the use of the central tip; the structure dimension evaluates the proportion of the characters and the stability of the center of gravity; the composition dimension analyzes the continuity of the flow and the overall layout; and the ink technique dimension examines the rhythmic expression of the ink density and dryness.

[0046] Figure 2 This is a schematic diagram of the scoring discrepancy detection process of the multi-model collaborative scoring module in an embodiment of the present invention.

[0047] like Figure 2 As shown, the scoring discrepancy detection process of the multi-model collaborative scoring module in this embodiment is as follows: Step S2-1: In this embodiment, two or more multimodal large models with differences in model architecture or training objectives are used as collaborative scoring engines. The model selection must meet the complementary principle of general capabilities and domain adaptability. Model A is a general visual language large model (such as LLaVA or Qwen-VL, Qwen-VL is used in this embodiment), which has cross-domain image and text understanding capabilities. Model B is fine-tuned on a calligraphy rubbing dataset to enhance its sensitivity to traditional calligraphy features, thus forming a complementary evaluation perspective.

[0048] Step S2-2: Input the multimodal data units generated in step S1 into each model. Each model independently extracts the joint visual and semantic features and performs structured scoring based on the preset four-dimensional calligraphy evaluation system. Among them, the brushwork dimension judges the quality of brushwork by analyzing the gradient of the stroke edge, the change of ink width and the shape of the beginning and ending strokes. The structure dimension evaluates the proportion of characters and the center of gravity shift by detecting key points of the character skeleton. The composition dimension measures the inter-line correspondence and spatial rhythm by the global attention weight distribution. The ink technique dimension analyzes the layering of ink density and dryness by combining local contrast and grayscale histogram.

[0049] In steps S2-3, each model outputs a quantitative score in four dimensions (ranging from 0 to 100), which is then aggregated into a structured score vector. To improve the stability of the score, each model can perform multiple inferences on the same work, such as different cropped areas or enhanced views, and take the average value as the final output to ensure that the score results are not affected by local noise or perspective bias.

[0050] Step S3 is executed by the knowledge graph reasoning module. Based on the pre-constructed knowledge graph of the calligraphy field (covering rubbings from ancient steles, calligraphic styles, aesthetic theories, and technical norms), the features extracted in step S2 are semantically matched with the entities in the graph to generate interpretable evaluations with cultural basis, such as "the composition is unbalanced, deviating from the continuous flow of Wang Xizhi's running script" or "the ink color is dry and monotonous, failing to reflect the ink aesthetics of 'moist like spring rain and dry like autumn wind'".

[0051] Figure 3 This is a schematic diagram of the semantic matching process for generating interpretability evaluation of knowledge graph reasoning in an embodiment of the present invention.

[0052] like Figure 3 As shown, the semantic matching process for interpretability evaluation of knowledge graph reasoning in this embodiment is as follows: Step S3-1 involves pre-constructing a knowledge graph for the calligraphy field, organizing structured knowledge in a subject-verb-object tripartite format. Graph nodes include classic calligraphic works from various dynasties such as the *Preface to the Orchid Pavilion* and the *Draft of a Eulogy for My Nephew*, calligraphic styles such as Ouyang Xun, Yan Zhenqing, and Zhao Mengfu, aesthetic concepts like *qiyun* (spirit resonance), *guli* (strength), and *feibai* (flying white stroke), and technical terms like *zhongfeng* (central stroke), *ti'an* (lifting and pressing), and *kubi* (dry brush). Furthermore, edge relationships encompass semantic types such as "belongs to," "influences," "embodies," and "contrasts." Expert verification ensures the accuracy and cultural authority of the knowledge. The expert verification employs a three-tiered review system: first, two calligraphy experts review the accuracy of entities and relationships; second, three calligraphy exhibition and competition judges review the semantic logic; and finally, one researcher in the field of artificial intelligence reviews the adaptability of the knowledge structure. Only graphs with a pass rate greater than 98% can be put into use.

[0053] Step S3-2: Semantic matching between structured query vectors and knowledge graph entities is calculated using cosine similarity, with a matching threshold set to ≥0.8. When the similarity reaches the threshold, they are considered semantically similar, triggering the subsequent comment generation process. The four-dimensional scores and intermediate features output by each model in step S2, such as the composition attention heatmap, ink distribution histogram, and structure centroid coordinates, are converted into structured query vectors, and semantically similar entities and rules are retrieved in the knowledge graph. For example, when the composition score is low and the standard deviation of the line spacing is large, the system matches the graph path "Wang Xizhi's running script → continuous flow of energy → requires close correspondence between lines" as the basis for explanation.

[0054] Step S3-3: Based on the matching results, call natural language templates to generate interpretable evaluation statements with cultural context; the template library has multiple preset expression paradigms (such as "deviation from... characteristics", "failure to reflect... aesthetics", "close to... style but... insufficient"), which are dynamically filled in combination with specific scoring deviations, and output professional feedback such as "the ink color is dry and monotonous, failing to reflect the ink aesthetics of 'moistening with spring rain and cracking with autumn wind'" or "the structure's center of gravity is tilted to the right, which does not conform to the structural norm of 'central and stable' in Ouyang Xun's calligraphy style", for the initial screening decision or subsequent manual review reference.

[0055] Step S4 is executed by the disagreement detection and preliminary screening decision module. It calculates the standard deviation between the total scores output by each model (i.e., the scores after weighted or equally weighted summation of the four dimensions: brushwork, structure, composition, and ink technique). If this deviation is greater than a preset disagreement threshold, the work is marked as a disputed sample and sent to the human expert review queue; otherwise, based on a comparison between the mean score and the preset threshold, it automatically determines whether the work should proceed to the next round of evaluation, completing the preliminary screening. Figure 2 As shown, the specific implementation steps are as follows: Step S4-1: Normalize the four-dimensional scores output by each model to eliminate dimensional differences, and calculate the total score according to preset weights, such as brushwork 30%, structure 30%, composition 20%, and ink technique 20%, or an equal-weighted method, to form the comprehensive score of each model. The weights can be dynamically configured according to different evaluation scenarios. For example, increase the weight of structure in regular script competitions, and increase the weight of composition and ink technique in running script and cursive script exhibitions. If three or more models are deployed, the total score is calculated using a weighted average method, and the weights of each model are dynamically allocated based on its scoring accuracy on the calligraphy validation set. The divergence degree is still calculated using the standard deviation of the total scores of all models, and the divergence threshold remains consistent.

[0056] Step S4-2: Calculate the standard deviation of the total score across all models as a quantitative indicator of the degree of divergence between models. If this degree of divergence is greater than a preset divergence threshold, for example, a standard deviation ≥ 8.0 (based on a percentage score), the work is determined to have evaluation uncertainty, marked as a controversial sample, and automatically added to the human expert review queue to ensure that highly divergent works are not mistakenly screened. This embodiment prioritizes using standard deviation as a quantitative indicator of divergence because it can more stably reflect the dispersion of multi-model scores, adapts to collaborative review scenarios with two or more heterogeneous models, and avoids misjudgments caused by the sensitivity of the range to outliers.

[0057] Step S4-3: For uncontroversial works, calculate the average of their total scores and compare it with a preset threshold (e.g., 60 points on a 100-point scale, 65 points in this embodiment). If the average score is higher than the threshold, the work is determined to enter the next round of selection; otherwise, it fails the screening. The scoring record and interpretable comments are saved for subsequent random checks or appeal verification, thereby completing an efficient and transparent intelligent initial screening process.

[0058] Example 2: Extended Implementation of the Invention It should be noted that Embodiment 1 is merely an example, and the scope of protection of this invention is not limited thereto. In other embodiments of this invention: In the two heterogeneous multimodal large models, Model A can also use the LLaVA-1.6-34B model with a visual Transformer architecture, which extracts stroke trajectory features through the CLIP visual encoder. Model B uses a convolutional-attention hybrid architecture model based on Qwen-VL-Chat, which uses the CogAgent architecture to analyze the structure of individual characters. This combination is suitable for specialized exhibitions and competitions in regular script and running script, enhancing the recognition accuracy of specific calligraphic features.

[0059] Knowledge graphs can be stored not only in Neo4j, but also in other graph databases or standard triplet formats.

[0060] The scoring weights can be configured unequally (e.g., dynamically configured according to the type of exhibition or competition, such as increasing the weight of character structure to 40% in regular script evaluation and increasing the weight of composition to 30% in running script and cursive script exhibitions and competitions) to adapt to different levels of rigor in the judging process. Through the above mechanism, this invention achieves efficient, intelligent, and interpretable initial screening of high-throughput calligraphy works, and is applicable to scenarios such as exhibition and competition submission screening and grading.

[0061] The divergence threshold (τ_div) and pass threshold (τ_pass) can be dynamically adjusted according to different judging scenarios (such as children's competitions and professional competitions).

[0062] The above description is only an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A smart initial screening system for calligraphy works based on multimodal large-scale model collaboration and knowledge graph, characterized in that, include: The multimodal input module is used to receive image data and corresponding text interpretations of the calligraphy works to be evaluated. The multi-model collaborative scoring module, connected to the multi-modal input module, is used to call at least two multi-modal large models that differ in model architecture or training objectives, to independently analyze the work from four dimensions: brushwork, structure, composition, and ink technique, and output quantitative scores. The knowledge graph reasoning module, connected to the multi-model collaborative scoring module, is used to perform semantic matching on the scoring results based on a pre-built knowledge graph in the calligraphy field, and generate interpretable evaluations with cultural basis. The disagreement detection and initial screening decision module is connected to the knowledge graph reasoning module and the multi-model collaborative scoring module. It is used to calculate the standard deviation between the scores of each model. If the difference exceeds the preset disagreement threshold, the work is marked as a disputed sample and sent to the human expert review queue. Otherwise, it is automatically determined whether to enter the next round of selection based on the average score.

2. The intelligent preliminary screening system for calligraphy works based on multimodal large model collaboration and knowledge graph as described in claim 1, characterized in that: in, The multimodal input module also includes a preprocessing unit, which is used to crop the image border, correct perspective distortion, adjust white balance and enhance contrast; and to use the OCR engine to extract or receive text translations submitted by the user from the image to form multimodal data units corresponding to the image and text.

3. The intelligent preliminary screening system for calligraphy works based on multimodal large model collaboration and knowledge graph as described in claim 1, characterized in that: in, The multi-model collaborative scoring module calls at least two multimodal large models, including a general visual language model and a domain-specific model fine-tuned on a calligraphy rubbings dataset.

4. The intelligent preliminary screening system for calligraphy works based on multimodal large model collaboration and knowledge graph as described in claim 1, characterized in that: in, The knowledge graph reasoning module is specifically used to: semantically match the visual and semantic features extracted by the multi-model collaborative scoring module with the entities and relationships in the knowledge graph to generate the interpretability evaluation. The knowledge graph is constructed using a graph database or a "subject-verb-object" triple format, and its nodes include historical rubbings, calligraphic styles, aesthetic theories, and technical specifications.

5. The intelligent preliminary screening system for calligraphy works based on multimodal large model collaboration and knowledge graph as described in claim 1, characterized in that: in, The preset divergence threshold is dynamically adjusted according to the exhibition / competition level or type.

6. A method for intelligent initial screening of calligraphy works based on multimodal large model collaboration and knowledge graph, characterized in that, Includes the following steps: Step S1: Obtain the image data and corresponding text transcription of the calligraphy work to be evaluated as multimodal input; Step S2: The multimodal input is fed into at least two large multimodal models that differ in model architecture or training objectives. Each model independently analyzes the work from four dimensions: brushwork, structure, composition, and ink technique, and outputs a quantitative score. Step S3: Based on the pre-constructed knowledge graph of the calligraphy domain, perform semantic interpretation on the scoring results of step S2 to generate an interpretable evaluation with cultural basis. Step S4: Calculate the standard deviation between the total scores output by each model. If the standard deviation is greater than the preset divergence threshold, the work is marked as a disputed sample and sent to the human expert review queue; otherwise, the work is automatically determined to enter the next round of selection based on the average score and the judgment threshold, thus completing the initial screening.

7. The intelligent initial screening method for calligraphy works based on multimodal large model collaboration and knowledge graph as described in claim 6, Its features are: Step S1 includes the following steps: Step S1-1: Acquire digital images of calligraphy works using high-resolution professional equipment; Step S1-2: Preprocess the digital image and simultaneously use an OCR engine to extract the corresponding text from the image to form a multimodal input pair of image and text. Steps S1-3: Package the preprocessed image and the verified text translation into a unified format multimodal data unit, which will be used as the input for the subsequent multimodal large model.

8. The intelligent initial screening method for calligraphy works based on multimodal large model collaboration and knowledge graph as described in claim 6, Its features are: Step S2 includes the following steps: Step S2-1: Call at least two multimodal large models that differ in model architecture or training objectives as collaborative scoring engines; Step S2-2: Input the multimodal data units generated in step S1 into each model, and each model independently extracts the joint visual and semantic features, and performs structured scoring based on the preset four-dimensional calligraphy evaluation system. In steps S2-3, each model outputs a quantitative score across four dimensions, which is then aggregated into a structured score vector.

9. The intelligent initial screening method for calligraphy works based on multimodal large model collaboration and knowledge graph as described in claim 6, Its features are: Step S3 includes the following steps: Step S3-1: Pre-construct a knowledge graph for the calligraphy domain, organizing structured knowledge in the form of "subject-verb-object" triplets; Step S3-2: When the semantic similarity between the structured query vector and the knowledge graph entity reaches the threshold, they are considered to be semantically similar, triggering the subsequent comment generation process. Step S3-3: Based on the matching results, call the natural language template to generate interpretable evaluation statements with cultural context.

10. The intelligent initial screening method for calligraphy works based on multimodal large model collaboration and knowledge graph as described in claim 6. Its features are: Step S4 includes the following steps: Step S4-1: Normalize the four-dimensional scores output by each model to eliminate differences in dimensions, and calculate the total score according to the preset weight or equal weight method to form the comprehensive score of each model. Step S4-2: Calculate the standard deviation of the total score of all models as a quantitative indicator of the degree of divergence between models. If the degree of divergence is greater than the preset divergence threshold, it is marked as a controversial sample and automatically added to the human expert review queue. Step S4-3: For uncontroversial works, calculate the average of their total scores and compare it with a preset threshold; if the average score is higher than the threshold, the work is determined to enter the next round of evaluation.