A Cultural Image Recognition Method and System Based on Prior Knowledge Space and Ontology Gating

By using a method based on prior knowledge space and ontology gating, cultural images are identified step by step, solving the problems of cultural illusion and untraceability in cultural image recognition. This enables the generation of auditable cultural image recognition reports, improving the credibility and transparency of the output.

CN122493462APending Publication Date: 2026-07-31ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies for cultural image recognition suffer from cultural illusion failure modes and reasoning paradigm defects, failing to effectively trace back to actual visual evidence and lacking a systematic ontological representation of the core categories of cultural images, resulting in untraceable interpretation conclusions.

Method used

We employ a method based on prior knowledge space and ontology gating. Through a multi-stage ontology gating mechanism, combined with domain ontology, knowledge graph and multimodal large language model, we gradually identify cultural images and generate auditable cultural image recognition reports to ensure that each output claim can be traced back to the underlying evidence.

Benefits of technology

It effectively suppresses cultural illusions, enables end-to-end auditable reasoning, ensures the credibility and transparency of outputs, and possesses good scalability and uncertainty management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493462A_ABST
    Figure CN122493462A_ABST
Patent Text Reader

Abstract

This invention discloses a cultural image recognition method and system based on prior knowledge space and ontology gating, belonging to the field of image recognition technology. The method includes: acquiring a cultural image set and constructing a prior knowledge space; obtaining a cultural image recognition report based on an ontology gating model using the target cultural image, work metadata, and the prior knowledge space, including: generating content candidate sets and form candidate sets for the target cultural image based on a domain ontology and a controlled vocabulary; extracting a focus set based on the work metadata and content candidate set and obtaining the output of the description stage; obtaining the output of the analysis stage based on the focus set and form candidate set; obtaining the output of the interpretation stage based on the prior knowledge space, focus set, description stage, and analysis stage outputs; obtaining the output of the evaluation stage based on the description stage, analysis stage, and interpretation stage outputs; and summarizing the outputs of each stage to generate a cultural image recognition report. This method effectively suppresses cultural illusions, and the evidence is traceable and highly credible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image recognition technology, specifically relating to a cultural image recognition method and system based on prior knowledge space and ontology gating. Background Technology

[0002] Cultural images, especially traditional figure paintings, belong to a complex visual medium. The complete identification of cultural images is not a simple object recognition task, but a progressive chain of cognitive reasoning: from the objective description of observable visual content, through the analysis of formal language and brushwork structure, to the interpretation of iconographic motifs and cultural significance, and finally to comparative evaluation placed within a historical context. Any semantic leap between levels, or inferences made before sufficient lower-level visual evidence is established, will undermine the traceability upon which credible identification depends.

[0003] Existing technologies typically employ the following technical solutions:

[0004] 1) Multimodal Large Language Model (MLLM) is used for cultural image recognition. However, a cultural illusion failure mode exists, meaning that it is impossible to trace back to actual visual evidence or historical context. Specifically, this manifests as: a structural disconnect between the ability to recognize visual content and the ability to infer cultural meaning; the end-to-end black-box generation paradigm lacks stage constraints, leading to the premature activation of high-level cultural conclusions before the underlying evidence is established; and the output claims lack visual anchoring and auditability.

[0005] 2) Cultural image recognition based on the Retrieval Enhanced Generation (RAG) method. Although the introduction of an art domain knowledge base partially alleviates the illusion problem, the following shortcomings still exist: 1) Reasoning paradigm deficiency: Existing RAG methods operate in a one-time generation mode, inputting the retrieved knowledge as a whole into a multimodal large language model, lacking structured constraints on cognitive sequences, leading to premature activation of high-level cultural interpretations; 2) Knowledge base deficiency: Existing art domain knowledge bases lack systematic ontological representations of core categories of cultural images (such as penmanship, visual style, cultural semantics, etc.), and cannot provide accurate semantic constraints for high-level cultural interpretations.

[0006] Furthermore, existing technologies cannot provide explicit evidence for each output claim, resulting in untraceable and unverifiable interpretations. Therefore, this paper proposes a cultural image recognition method and system based on prior knowledge space and ontology gating. Summary of the Invention

[0007] The purpose of this invention is to address the above-mentioned problems by proposing a cultural image recognition method and system based on prior knowledge space and ontology gating. By combining prior knowledge space and multi-stage ontology gating mechanism, it achieves progressive and auditable recognition from visual content to cultural meaning, effectively suppresses cultural illusion, and ensures that each output claim can be traced back to underlying evidence.

[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0009] The cultural image recognition method based on prior knowledge space and ontology gating proposed in this invention includes the following steps:

[0010] S1. Obtain the cultural image set and construct the corresponding prior knowledge space. ,in, For the domain ontology, It is a knowledge graph, constructed from a controlled vocabulary. For symbol mapping rule base;

[0011] S2. Based on the target cultural image and corresponding work metadata, as well as the prior knowledge space, an ontology-gated model is used to obtain the corresponding cultural image recognition report, as detailed below:

[0012] S21. Based on the domain ontology and controlled vocabulary, a multimodal large language model is used to generate content candidate sets and form candidate sets for target cultural images;

[0013] S22. Based on the work's metadata and content candidate set, use a multimodal large language model to extract the focus set of the target cultural image, and obtain the output of the description stage, including several description claims corresponding to the focus set, and each description claim carries an L1 evidence ID and a corrected confidence level for each D-stage claim.

[0014] S23. Based on the focus set and the formal candidate set, use a multimodal large language model to obtain the output of the analysis phase, including several analysis claims, and each analysis claim carries a corresponding L2 evidence ID and a modified confidence level for each A-stage claim;

[0015] S24. Based on the outputs of the prior knowledge space, focus set, description stage, and analysis stage, the output of the interpretation stage is obtained using a multimodal large language model.

[0016] S25. Based on the outputs of the description, analysis, and interpretation phases, use a multimodal large language model to obtain the output of the evaluation phase, including an evaluation claim along with the corresponding output of the interpretation phase and the evaluation confidence level.

[0017] S26. Summarize the outputs of each stage to generate a cultural image recognition report.

[0018] Preferably, the domain ontology is used to divide the semantic elements of all cultural images in the cultural image set into a first semantic layer, a second semantic layer, and a third semantic layer, wherein:

[0019] The first semantic layer includes the entities of the cultural image, entity attributes, spatial location and interaction relationships between entities;

[0020] The second semantic layer includes composition, brushwork techniques, and ink coloring.

[0021] The third semantic layer includes iconographic motifs, narrative themes, visual styles, and cultural semantics;

[0022] The controlled vocabulary includes a set of legal fields, a set of enumerable tags, and inter-layer semantic relationship types. The set of legal fields includes several predefined fields corresponding to the first semantic layer, the second semantic layer, and the third semantic layer. The set of enumerable tags includes several predefined tags allowed for each field in each semantic layer. The inter-layer semantic relationship types include the indicative relationship between the first semantic layer and the second semantic layer, the evocation relationship between the second semantic layer and the third semantic layer, and the mapping relationship between the first semantic layer and the third semantic layer.

[0023] The symbol mapping rule base includes several cross-layer triggering rules extracted from the knowledge graph. Each cross-layer triggering rule is represented as follows:

[0024]

[0025] In the formula, , , As a combination of content clues, Formal clues are combined. This is a candidate set of content, belonging to the first semantic layer of the controlled vocabulary. This is a formal candidate set, belonging to the second semantic layer of the controlled vocabulary. It is an L3 semantic candidate, belonging to the third semantic layer of the controlled vocabulary. For conjunction operation, This indicates that it has been triggered.

[0026] Preferably, the types of nodes in the knowledge graph include work nodes, work metadata nodes, image region nodes, L1 content entity nodes, L2 formal feature nodes, L3 meaning nodes, text evidence nodes, and rule nodes;

[0027] The types of edges in a knowledge graph include:

[0028] Entity Relationship Edges: connect work nodes to work metadata nodes, connect work nodes to L1 content entity nodes, connect work nodes to L2 formal feature nodes, and connect work nodes to L3 meaning nodes;

[0029] First spatial relation edge: connects the image region node and the L1 content entity node, carrying the corresponding initial confidence and image region coordinates;

[0030] Second spatial relation edge: connects image region nodes with L2 form feature nodes, carrying the corresponding initial confidence and image region coordinates;

[0031] The first co-occurrence support edge: connects the L3 meaning node with the L1 content entity node that supports the L3 meaning node, and connects the L2 form feature node with the L3 meaning node that supports the L3 meaning node, carrying corresponding weights;

[0032] Second co-occurrence support edge: connects L3 meaning nodes and text evidence nodes, carrying corresponding weights;

[0033] Rule-triggered edges: connect rule nodes to L3 significance nodes that trigger the rule node;

[0034] Contrast-based edge: connects L3 meaning nodes with other L3 meaning nodes, connects work metadata nodes with L3 meaning nodes, and connects text evidence nodes with L3 meaning nodes, carrying corresponding weights;

[0035] Cross-work comparison edge: connects image region nodes with similarity scores higher than a preset similarity threshold, and connects comparable nodes across works with similarity scores higher than a preset similarity threshold, carrying the corresponding similarity scores.

[0036] Preferably, the focus set is extracted as follows:

[0037] Based on the work's metadata and content candidate set, a multimodal large language model is used to identify the target entities of the target cultural image. The target entities, as well as entities that have spatial location and interaction relationships with the target entities, are included from the content candidate set into the initial focus set. The work's metadata includes the work's name, author, dynasty, work type, material, and collection information.

[0038] The initial focus set is matched and corrected with the entity patterns of comparable works in the knowledge graph using the correction operators of the knowledge graph to obtain the focus set. The correction operators of the knowledge graph are used to perform bidirectional correction, as follows:

[0039] 1) When the co-occurrence frequency of entities in the initial focus set in the entity patterns of the comparable works set is less than the false detection threshold, the corresponding entity is removed from the initial focus set or the initial confidence of the spatial relation edge of the first semantic layer associated with the corresponding entity is reduced. The co-occurrence frequency is the ratio of the number of cultural images in the comparable works set that contain the corresponding entity to the total number of cultural images in the comparable works set. The comparable works set consists of a preset number of cultural images in the knowledge graph that are adjacent to or the same as the target cultural image in terms of dynasty, or have the same category of narrative theme, or have a pre-defined number of cultural images that have a pre-defined number of cross-work comparison edges that are related. The entity pattern includes all entity combinations of the corresponding entity in the comparable works set and their corresponding co-occurrence frequencies.

[0040] 2) If an entity in the content candidate set is not included in the initial focus set because its initial confidence level is lower than the first preset threshold, but the co-occurrence frequency of the entity in the entity mode of the comparable works set is greater than or equal to the false detection threshold, then the corresponding entity will be re-included from the content candidate set into the initial focus set, and the initial focus set after bidirectional correction will be regarded as the focus set.

[0041] Preferably, the multimodal large language model adopts the GPT-5.4 model;

[0042] The descriptive claims are represented by “people-objects-environment-interaction relationship”, where people, objects, environment, and interaction relationship are the four D-stage claims in the descriptive claims, and L1 evidence ID is the ID of the corresponding element in the content candidate set;

[0043] The analytical claims are represented by “composition-brushwork-coloring”, where composition, brushwork, and coloring are the three A-stage claims in the analytical claims, and L2 evidence ID is the ID of the corresponding element in the formal candidate set.

[0044] Preferably, the corrected confidence level of the D-stage claim is a weighted sum of the initial confidence level of the spatial relation edge associated with the first semantic layer in the current D-stage claim and the co-occurrence frequency of the current D-stage claim;

[0045] The revised confidence level of the A-stage claim is the weighted sum of the initial confidence level of the spatial relation edge of the current A-stage claim associated with the second semantic layer and the historical form convention conformity of the current A-stage claim.

[0046] Preferably, the explanatory phase includes:

[0047] (1) Theme Explanation Sub-stage:

[0048] Based on the corresponding cross-layer triggering rules in the focus set activation symbol mapping rule base, all generated L3 semantic candidates form the L3 candidate set;

[0049] The output of the topic explanation sub-stage is obtained by using a multimodal large language model, including L3 semantic candidates in the L3 candidate set and carrying the corresponding rule node ID, L1 evidence ID, L2 evidence ID, L3 meaning node ID, text evidence node ID associated with cross-work comparison edge, and first confidence score. The L3 semantic candidates in the L3 candidate set are either topic type or specific motif.

[0050] (2) Atmosphere Explanation Sub-stage:

[0051] Based on the output of the focus set and analysis phase, the output of the atmosphere interpretation sub-phase is obtained using a multimodal large language model, including L3 atmosphere slots and corresponding L1 evidence IDs, L2 evidence IDs and second confidence scores. The L3 atmosphere slots include global mood, character emotions and environmental atmosphere.

[0052] (3) Value synthesis sub-stage:

[0053] Based on the outputs of the description stage, analysis stage, theme interpretation sub-stage, and atmosphere interpretation sub-stage, the output of the value synthesis sub-stage is obtained using a multimodal large language model, including the cultural context and carrying the evidence IDs and synthesis confidence scores of the theme interpretation sub-stage and atmosphere interpretation sub-stage;

[0054] L3 semantic candidates with a comprehensive confidence level higher than the second preset threshold are used as the main explanations, while L3 semantic candidates with a comprehensive confidence level lower than the second preset threshold are retained with uncertainty labels, which indicate the reason for insufficient evidence.

[0055] Preferably, the overall confidence level of the Phase I claim output by the value synthesis sub-stage is the weighted sum of the observed confidence level of the current Phase I claim, the rule activation strength, the knowledge graph verification consistency score, and the evidence chain integrity score, where:

[0056] The observed confidence level is the average of the corrected confidence levels of all L1 evidence IDs that trigger the current Phase I claim and the adjusted confidence levels of all L2 evidence IDs.

[0057] The rule activation strength is the product of the initial weight of the rule triggering edge that triggers the current I-stage claim in the symbol mapping rule base and the minimum evidence set satisfaction. The minimum evidence set satisfaction is the ratio of the total number of elements in the first semantic layer and the second semantic layer that trigger the current I-stage claim to the total number of all elements in the minimum evidence set. The minimum evidence set is the minimum combination of clues formed by elements in the content layer L1 and the form layer L2 that trigger the current I-stage claim.

[0058] Knowledge graph consistency score The formula is as follows:

[0059]

[0060] in, This indicates taking the minimum value. This refers to the co-occurrence frequency claimed in the current Phase I. For co-occurrence reward coefficient, For proof by contradiction, the penalty coefficient, The evidence suppression coefficient;

[0061] The evidence chain integrity score is the sum of the number of L1 evidence IDs, L2 evidence IDs, and text evidence node IDs associated with cross-work comparison edges that the current phase I claim is traced, divided by the preset number of evidence. The preset number of evidence includes the sum of the number of rule node IDs, L1 evidence IDs, L2 evidence IDs, L3 meaning node IDs, and text evidence node IDs associated with cross-work comparison edges that the current phase I claim is traced.

[0062] Preferably, the confidence level of the E-stage claim is a weighted sum of the overall confidence level, cross-work support, textual source support, and uncertainty penalty of the current E-stage claim, where:

[0063] Cross-work support is the weighted sum of cross-work similarity and co-occurrence support scores of the current E-stage claim. Cross-work similarity is the average of the similarity scores of all cross-work comparison edges associated with the current E-stage claim. Co-occurrence support score is the average of the weights of all co-occurrence support edges associated with the current E-stage claim.

[0064] The minimum value among the text source support score of 1 and the weighted sum of the text evidence normalization score, text evidence reliability score, and text evidence connection strength claimed in the current E stage, and the text evidence normalization score is the average of the preset source coefficients of the text evidence corresponding to the text evidence node claimed in the current E stage, and the text evidence connection strength is the average of the weights of all co-occurring support edges of the associated text evidence claimed in the current E stage.

[0065] The uncertainty penalty term is the minimum value among the weighted sum of the average weight of all counter-evidence edges associated with the current E-stage claim and the uncertainty label score. When the uncertainty label does not exist, the uncertainty label score is 0, and when the uncertainty label exists, the uncertainty label score is 0.20~0.80.

[0066] A cultural image recognition system based on prior knowledge space and ontology gating includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, it implements the cultural image recognition method based on prior knowledge space and ontology gating as described above.

[0067] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0068] 1) Effectively suppressing cultural illusion: The ontology gating mechanism, through hard constraints on staged field access permissions, can prevent the premature activation of high-level cultural meanings before the underlying evidence is established, fundamentally solving the problem of cultural illusion in multimodal large language models for cultural image understanding; 2) Achieving end-to-end auditable reasoning: Each output claim carries explicit evidence tracing links, confidence levels, and uncertainty annotations, making the entire process from cultural image input to cultural image recognition report output traceable and questionable, meeting the requirements for transparency of explanation, and achieving auditable output; 3) Decoupling knowledge and reasoning: By separating the prior knowledge space from the ontology gating mechanism, the prior knowledge space can be independently updated and expanded without affecting the reasoning architecture of the ontology gating mechanism, exhibiting good scalability; 4) Explicit management of uncertainty: When evidence is insufficient, competing L3 semantic candidates are retained instead of forcibly generating a single conclusion, achieving explicit management of reasoning uncertainty and improving the credibility of the output. Attached Figure Description

[0069] Figure 1 This is a flowchart of the cultural image recognition method based on prior knowledge space and ontology gating of the present invention;

[0070] Figure 2 This is a flowchart illustrating the construction process of the ontology in the field of this invention.

[0071] Figure 3 This is a schematic diagram illustrating the extraction of the content candidate set and form candidate set of the target cultural image in an embodiment of the present invention;

[0072] Figure 4 This is a schematic diagram of the DAIE inference stage of the target cultural image in an embodiment of the present invention;

[0073] Figure 5 This is a schematic diagram of a cultural image recognition report for a target cultural image according to an embodiment of the present invention. Detailed Implementation

[0074] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0075] It should be noted that, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application.

[0076] Example 1:

[0077] like Figures 1-5 As shown, the cultural image recognition method based on prior knowledge space and ontology gating includes the following steps:

[0078] S1. Obtain several cultural images to form a cultural image set and construct the prior knowledge space of the cultural image set. ,in, For the domain ontology, For knowledge graphs, This is a symbol mapping rule base.

[0079] Prior knowledge space The structure is as follows:

[0080] (1) Domain Ontology , means as follows:

[0081] The semantic elements of all cultural images in the cultural image collection are divided into three progressive semantic layers:

[0082] The first semantic layer (content layer L1) includes entities of cultural images, entity attributes, and spatial positions and interactions between entities, providing field constraints for the description stage (D). Specifically, entities include people, objects, and scenery; entity attributes include identity, clothing, posture, and expression; spatial positions and interactions between entities refer to spatial relationships, interaction relationships, and environmental attribution relationships between entities; spatial relationships include front and back, distance, left and right, up and down, and orientation; interaction relationships include adjacency, surrounding, holding, and sitting; and environmental attribution relationships include indoors, in a pavilion, by the water, objects placed on a table, and objects placed in the hands of people.

[0083] The second semantic layer (formal layer L2) includes composition, brushwork techniques, and ink coloring, providing field constraints for the analysis phase (A). Specifically, composition refers to the overall organization of cultural images, such as centralizing the subject, left-right symmetry, density arrangement, white space, and handscroll-style scattered unfolding; brushwork techniques refer to brushwork methods such as line drawing and texturing, such as iron wire drawing, gossamer drawing, hemp fiber texturing, and axe-cut texturing; ink coloring refers to the density, dryness, wetness, layering, and rendering relationship of ink, such as ink wash, dark ink outlining, light ink shading, and dry brush rubbing.

[0084] The third semantic layer (meaning layer L3) includes iconographic motifs, narrative themes, visual styles, and cultural semantics, providing field constraints for the interpretation stage (I) and the evaluation stage (E). Specifically, iconographic motifs refer to cultural image units triggered by a combination of clues, such as the zither, pine, and mountain dwelling, which can all point to the theme of reclusion. Narrative themes refer to the type of event or plot expressed by the cultural image as a whole, such as visiting friends, banquets, and listening to the zither. Visual styles (atmosphere) refer to the overall atmosphere formed by formal factors such as composition, brushwork, and ink coloring, such as coldness, desolation, leisure, and solemnity. Cultural semantics (artistic conception) refers to the spiritual interest or cultural stance presented by the cultural image, such as the ideal of reclusion, the consciousness of the remnants of the Ming dynasty, and the lofty aspirations of living in forests and springs.

[0085] Each semantic layer is configured with a corresponding sub-controlled vocabulary. The sub-controlled vocabularys of all semantic layers constitute the controlled vocabulary. The controlled vocabulary includes a set of legal fields, a set of enumerable tags, and inter-layer semantic relationship types. The inter-layer semantic relationship types include indexes between the content layer and the form layer, evokes between the form layer and the meaning layer, and maps onto between the content layer and the meaning layer. For example, the content layer "the fisherman is located at the corner of the picture" indicates the form layer "corner composition", the form layer "iron wire drawing" evokes the meaning layer "orderly and solemn", and the content layer "fisherman + fishing boat + river" maps the meaning layer "fisherman in seclusion, fisherman sleeping soundly, returning boat in wind and rain", etc. The indexes between the content layer and the form layer are used for retrieving indexes from content layer L1 to form layer L2; the evokes between the form layer and the meaning layer are used for retrieving evokes from form layer L2 to meaning layer L3; and the maps onto between the content layer and the meaning layer are used for retrieving maps onto from content layer L1 to meaning layer L3. These three relationships can be used for knowledge graph construction to achieve L3 semantic candidate triggering in subsequent interpretation stage I and evidence backtracking in DAIE structured report generation, etc.

[0086] In this embodiment, the controlled vocabulary contains 1,171 terms, divided into 28 tertiary categories under 8 secondary categories (L1: 639 terms; L2: 121 terms; L3: 411 terms). Consistency in annotation and support for specific retrieval in subsequent stages are ensured by defining the allowed set of legal fields, the enumerable tag set, and the types of semantic relationships between layers. For example, in the controlled vocabulary, the set of legal fields (secondary categories) constrains the fields allowed in each layer of the content layer L1, form layer L2, and meaning layer L3, including 8 secondary categories: character image, natural landscape, man-made environment, form and layout, style and technique, theme and narrative, atmosphere and tone, and cultural context. Character image, natural landscape, and man-made environment correspond to the content layer L1; form and layout and style and technique correspond to the form layer L2; and theme and narrative, atmosphere and tone, and cultural context correspond to the meaning layer L3. The enumerable tag set (tertiary categories) includes several element tags used for... The constraints correspond to the predefined tags allowed for each field in the semantic layer. For example, character image includes identity, clothing, posture, and expression; natural landscape includes landform, plants, animals, and seasonal weather; man-made environment includes architecture, landscape, furniture, and utensils; form and layout include painting form, composition mode, visual structure, and spatial atmosphere; style and technique include style type, line drawing techniques, coloring techniques, and ink painting techniques; thematic narrative includes theme type and specific motif; atmosphere and tone include overall artistic conception, character emotions, and environmental atmosphere; and cultural context includes style lineage, semantic foundation, and creative motivation. It should be noted that the names of the dividing fields in the controlled vocabulary and the domain ontology may not be exactly the same, but objects at the same level should belong to the same level. For example, character image is located in the content layer L1 of the controlled vocabulary, entities are located in the content layer L1 of the domain ontology, and character image is a type of entity, both of which are located in the content layer L1. It is easy to understand that the objects in the secondary and tertiary categories of the controlled vocabulary are conventionally labeled objects that are well known to those skilled in the art. They may be labeled manually by experts or according to actual needs. The labeling can be based on the cultural heritage concept framework defined by the international standard CIDOC-CRM (ISO 21127:2023) and the subject classification principles of the internationally accepted iconographic classification system ICONCLASS. The key difference of the method of this invention is that each object is classified into different semantic layers and the semantic relationship types between layers are established.

[0087] Domain Ontology Its core function is for ontology gating: during implementation, it strictly limits the accessible semantic layer fields according to the current reasoning stage, preventing any high-level fields from being activated before the underlying evidence is fully established. Specifically, it adopts a unidirectional hierarchical structure: the description stage (D) is limited to fields at the content layer (L1), the analysis stage (A) is limited to fields at the formal layer (L2), and only in the interpretation stage (I) and the evaluation stage (E) are fields at the meaning layer (L3) accessed. By implementing corresponding hard constraints in stages, it ensures that no high-level cultural meaning is activated before the underlying evidence is fully established.

[0088] (2) Knowledge Graph , means as follows:

[0089] Constructing a knowledge graph based on a controlled vocabulary Knowledge graph The types of nodes include Artwork, Dynasty, ImageRegion, L1 Item, L2 Item, L3 Item, TextSource, and Rule. A cultural image is considered a work node. The work metadata includes the work name, author, dynasty, work type, material, and collection information. One work metadata node corresponds to one element in the work metadata. One image region node corresponds to at least one entity. That is, the image region corresponding to one image region node may include one or more entities. One element in the content layer L1 corresponds to one L1 content entity node. One element in the form layer L2 corresponds to one L2 form feature node. One element in the meaning layer L3 corresponds to one L3 meaning node. Textual evidence includes three types of inscriptions, colophons, and commentary texts of cultural images in historical documents. One element in the textual evidence corresponds to one textual evidence node. Historical documents can be adjusted according to actual needs. This embodiment includes 36 painting works, museum archives, and 1,323 figure paintings in encyclopedias. One rule node corresponds to one minimum evidence set. The minimum evidence set is the minimum combination of clues formed by elements in the content layer L1 and the form layer L2 that is required to trigger an element in the meaning layer L3.

[0090] As shown in Table 1 below, the types of edges in a knowledge graph include entity relationship edges, spatial relationship edges, co-occurrence support edges, proof by contradiction edges, rule triggering edges, and cross-work comparison edges, among which:

[0091] ENTITY (Entity Relationship Edge): Connects work nodes to work metadata nodes, connects work nodes to L1 content entity nodes, connects work nodes to L2 formal feature nodes, and connects work nodes to L3 meaning nodes;

[0092] HAS_L1 (First Spatial Relation Edge): Connects the image region node and the L1 content entity node, carrying the initial confidence score and image region coordinates;

[0093] HAS_L2 (Second Spatial Relation Edge): Connects the image region node to the L2 form feature node, carrying the initial confidence score and image region coordinates;

[0094] SUPPORTED_BY (First Co-occurrence Support Edge): Connects the L3 meaning node to the L1 content entity node that supports (i.e., there is a mapping relationship between the content layer and the meaning layer) the L3 meaning node, and connects the L3 meaning node to the L2 formal feature node that supports (i.e., there is an evoking relationship between the formal layer and the meaning layer), carrying weights.

[0095] GROUNDED_IN (Second Co-occurrence Support Edge): Connects L3 meaning nodes and text evidence nodes, carrying weights;

[0096] TRIGGERS (rule-triggered edges): connects the rule node to the L3 significance node that triggered the rule node;

[0097] REFUTES (proof of contradiction edge): connects L3 meaning nodes with other L3 meaning nodes, connects work metadata nodes with L3 meaning nodes, and connects text evidence nodes with L3 meaning nodes, carrying weights;

[0098] SIMILAR_TO (Cross-work comparison edge): Connects nodes of different image regions with similarity scores higher than the preset similarity threshold, connects nodes of different cross-work comparable nodes with similarity scores higher than the preset similarity threshold, and carries the similarity score. The preset similarity threshold is 0.75.

[0099] Table 1

[0100]

[0101] Among them, nodes and nodes Similarity score The formula is as follows:

[0102]

[0103] in,

[0104]

[0105]

[0106]

[0107] in, For nodes and nodes The cosine similarity of the semantic embedding vectors. , This refers to the node index in the knowledge graph. For nodes and nodes Tag overlap For nodes and nodes Metadata similarity of works The vector similarity weight (preferably 0.50) The label overlap weight is 0.30 (preferably). Ea represents the similarity weight of the work's metadata (preferably 0.20), and Ea represents the node's similarity weight. The semantic embedding vector, where Eb is the node. semantic embedding vector, Using the Euclidean norm, a semantic embedding vector refers to a fixed-length vector obtained by encoding the content layer L1, formal layer L2, or semantic layer L3 of the corresponding node in the knowledge graph G through a text embedding model (such as the text encoder of the CLIP model, or the BERT model, or the RoBERTa model, etc.). For nodes A collection of tags, For nodes A collection of tags, This represents the number of elements in the corresponding set. For author similarity, For dynasty similarity, For similarity of work types, For the similarity of collection information, when , , , Middle node and nodes If the information is the same, a value of 1 is used; if it is different, a value of 0 is used. In particular, if the dynasties are adjacent, a value of 0.5 is used. For author similarity weighting, Weighting based on dynasty similarity. Weighting based on the similarity of work types. Weighting based on the similarity of collection information, and selecting the best option. =0.25, =0.35, =0.25, =0.15. It should be noted that the calculation of the similarity of work metadata is related to the content in the work metadata and can be adjusted according to actual needs. For example, when the work metadata does not contain collection information, the calculation of the similarity of work metadata only covers the author, dynasty, and work type. The corresponding weights can be readjusted according to the actual situation, such as 0.30, 0.40, and 0.30 respectively.

[0108] For example, nodes The tag set is {fisherman, fishing boat, riverbank, blank space}, node If the tag set is {fisherman, fishing boat, reeds, blank space}, then =3 / 5=0.60. If =0.86, =0.80, then:

[0109]

[0110] With a preset similarity threshold of 0.75, then 0.77 ≥ 0.75. Therefore, a SIMILAR_TO (cross-work comparison edge) is established between node a and node b, and similarity=0.77 is stored as the edge attribute (similarity score) of this cross-work comparison edge.

[0111] In Table 1 above, the edges include six types of semantic relations: entity relation edges, spatial relation edges, co-occurrence support edges, counter-evidence relation edges, rule triggering edges, and cross-work comparison edges. ImageRegion(g) and ImageRegion(h) represent image region nodes of different cultural images, Artwork(c) and Artwork(d) represent different work nodes, L3Item(e) and L3Item(f) represent different L3 meaning nodes, → indicates a one-way connection, ↔ indicates a two-way connection, and the arrow indicates the direction.

[0112] knowledge graph It supports retrieving stage-specific subgraphs by DAIE reasoning stage (including description stage D, analysis stage A, interpretation stage I, and evaluation stage E). (Description stage D is mainly represented by L1 subgraphs, analysis stage A by L2 subgraphs, and interpretation stage I and evaluation stage E by L3 subgraphs and cross-work subgraphs). The stage-specific subgraphs are obtained from the knowledge graph based on the current stage, the set of legal fields limited by ontology gating, and the current work (or scene fragments segmented from the current work). The work is a cultural image. Among them, the L1 subgraph is the knowledge graph. The L1 content layer is a subgraph associated with the current work (i.e., there are edges), and the L2 subgraph is a knowledge graph. The L2 subgraph in the middle formal layer is associated with the current work, and the L3 subgraph is the knowledge graph. The L3 subgraph in the middle meaning layer is associated with the current work, and the cross-work subgraph is a knowledge graph. A subgraph relating the current work to comparable portfolios.

[0113] knowledge graph It is used to support evidence retrieval, co-occurrence verification, and rebuttal penalty calculation in multi-stage reasoning (DAIE reasoning stage). Entity relation edges are used to extract stage-specific subgraphs and initial focus sets. With knowledge graph Entity pattern matching correction in comparable portfolios; spatial relationship edges used to extract content candidate sets. and formal candidate set The L3 semantic candidates are subjected to evidence chain verification based on co-occurrence supporting edges, penalties based on rule-triggered edges and counter-evidence edges, and cross-work verification based on cross-work comparison edges and co-occurrence frequencies. In other words, the L3 semantic candidates generated in the interpretation stage I are verified using a knowledge graph. When the entity pattern of the L3 semantic candidate is consistent with the clue combination in the comparable works set, a reward is given. When there is counter-evidence relationship, dynastic conflict, textual evidence conflict or cross-work comparison inconsistency, a penalty is imposed on it, thereby obtaining the confidence result of the L3 semantic candidate.

[0114] This embodiment utilizes multi-source data, including 36 painting books, museum archives, and encyclopedias, as historical documents to annotate 1,323 portrait paintings (cultural images) created by 488 artists. Annotation can be done manually, semi-automatically, or automatically, generating a structured JSON file, and ultimately presenting it as a graph-structured knowledge graph. Presented.

[0115] (3) Symbol mapping rule base , means as follows:

[0116] From knowledge graph Extract the cross-layer triggering rules. Each cross-layer triggering rule is represented as follows:

[0117]

[0118] In the formula, , , As a combination of content clues, Formal clues are combined. This is the content candidate set, belonging to the content layer L1 of the controlled vocabulary. This is a formal candidate set, belonging to the formal layer L2 of the controlled vocabulary. These are elements in the L3 semantic layer of the controlled vocabulary, namely L3 semantic candidates (thematic narrative, atmosphere, cultural context). For conjunction operation, This indicates an exit, or trigger. (This refers to the knowledge graph.) It contains 370 cross-layer triggering rules. The key lies in the combination effect: the combined appearance of a scholar, a zither, a pine tree, and a mountain dwelling is more reliable in triggering the reclusive candidate state than any single clue existing alone.

[0119] Each cross-layer triggering rule is configured with a minimum evidence set (MES) and a corresponding initial weight. The minimum evidence set (MES) specifies the minimum combination of clues required to trigger the corresponding L3 semantic candidate. The cross-layer triggering rule is activated only when the input clues satisfy the Minimum Evidence Set (MES), at which point the activation state is "activated"; otherwise, the activation state is "inactive". The initial weight is a value preset based on experience, representing the graphic authority (graphic evidence strength) of the cross-layer triggering rule.

[0120] S2. Obtain the corresponding cultural image recognition report based on the target cultural image, the corresponding work metadata, and the prior knowledge space using an ontology-gated model.

[0121] Among them, acquiring target cultural images The input features of the ontology gating model are the target cultural image, the corresponding work metadata, and prior knowledge space. The metadata includes the work's name, author, dynasty, type, material, and collection information. Work types include hanging scrolls, handscrolls, and album leaves. The output features constitute the corresponding cultural image recognition report. The target cultural image and the corresponding work's metadata form the bimodal input of the ontology gating model. The cultural image recognition report carries evidence links, confidence levels, and uncertainty annotations for the corresponding claims.

[0122] Specifically, the ontology gating model is constructed as follows:

[0123] S21, Based on Domain Ontology Controlled vocabulary for target cultural images Candidate feature extraction (initial evidence extraction) is performed using a multimodal large language model to generate target culture images. Content candidate set and formal candidate set The details are as follows:

[0124] (1) Content candidate set Extracting the domain ontology The entities, attributes, spatial locations and interactions between entities in the content layer L1 are used as candidate elements of the content layer L1 and normalized to the predefined labels of the content layer L1 in the controlled vocabulary. Each candidate element of the normalized content layer L1 is bound to a corresponding image region label (image region coordinates) and an initial confidence level.

[0125] (2) Formal candidate set Extracting the domain ontology The composition, brushwork, and ink color of the L2 form layer are used as candidate elements of the L2 form layer and normalized to the predefined labels of the L2 form layer in the controlled vocabulary. Each candidate element of the normalized L2 form layer is then bound to a corresponding image region label (image region coordinates) and an initial confidence level.

[0126] The multimodal large language model uses a first prompt word template for candidate feature extraction. In the first prompt word template, the system prompt words are used to constrain the output format of the multimodal large language model, and the user prompt words are used to execute the current task.

[0127]

[0128] Domain Ontology The controlled vocabulary and a dual-intervention approach are applied to candidate feature extraction in multimodal large language models: Field constraints—locking the output of the multimodal large language model to predefined fields and predefined labels, i.e., passing a whitelist of fields in the domain ontology and the element labels in the controlled vocabulary corresponding to each field into the multimodal large language model through structured hints (structured JSON files), requiring that each output of the multimodal large language model must correspond to the corresponding semantic layer in the domain ontology and the element label in the controlled vocabulary; Semantic enhancement—mapping open vocabulary predictions to standard terms, preserving... To ensure consistency of annotation across stages and improve retrieval efficiency, when the multimodal large language model describes any observable feature in a cultural image using open vocabulary, a second prompt word template is used to normalize the open vocabulary to standard terms in a controlled vocabulary. For example, the open vocabulary "delicate and smooth brushstrokes" is normalized to the element label "silk lines" in the controlled vocabulary, and the open vocabulary "light ink color" is normalized to the element label "light ink" in the controlled vocabulary. Through this dual intervention, the annotation terms of each inference stage are kept consistent, and the retrieval hit rate of candidate elements in the knowledge graph is improved.

[0129]

[0130] In this embodiment, the multimodal large language model adopts the GPT-5.4 model. However, existing multimodal large language models can be selected according to actual needs.

[0131] S22, Description Phase (D) – L1 Gated Content Description:

[0132] During the description phase (D), when applying ontology gating constraints, access is only allowed to the content layer L1, while access to the formal layer L2 and the semantic layer L3 is prohibited, and content candidate sets are used. Extracting target culture images focus set Get the focus set The corresponding pre-defined format of the description claims, each of the D-stage claims in the description claims carries the L1 evidence ID and the adjusted confidence level.

[0133] Specifically, target cultural images focus set The extraction process is as follows:

[0134] (1) Anchor point identification: Based on the work's metadata and content candidate set, the multimodal big language model is used to identify the target entity of the target cultural image as the anchor point. When perceiving the cultural image, the multimodal big language model works with the clues provided by the work's metadata and content candidate set to guide the multimodal big language model to locate the target subject as the anchor point, such as a person.

[0135] Relationship-centric expansion: This involves removing the target entity, and entities with spatial and interactive relationships with it, from the content candidate set. Included in the initial focus set The multimodal large language model uses a third-party cue word template to obtain the initial focus set. The details are as follows.

[0136]

[0137] (2) Initial focus set correction: using the correction operator of the knowledge graph The initial focus set With knowledge graph Match and correct the entity patterns of comparable portfolios to obtain a focus set. Correction operators for knowledge graphs Perform bidirectional corrections, specifically as follows: 1) Suppress the perceived incidental entities in the comparable works set, i.e., when the initial focus set... When the co-occurrence frequency of an entity in a comparable portfolio's entity patterns is less than the false detection threshold (preferably 0.2), it indicates that the corresponding entity is a false detection by the multimodal large language model, starting from the initial focus set. The corresponding entity will be removed or its weight reduced (reducing weight means removing the corresponding entity-related content layer output by the multimodal large language model). The initial confidence of the spatial relationship edge is multiplied by a preset coefficient less than 1 (e.g., 0.70) to update the initial confidence; the co-occurrence frequency is the ratio of the number of cultural images in the comparable works set containing entity combinations of the corresponding entity to the total number of cultural images in the comparable works set; 2) Recover the candidate set of content that is iconographically important but was missed due to the initial confidence being lower than the first preset threshold. The entities in the content candidate set Some entities did not enter the initial focus set because their initial confidence level was lower than the first preset threshold. However, if the co-occurrence frequency of this entity in the entity patterns of comparable portfolios is greater than or equal to the false detection threshold, it indicates that it has image saliency and is therefore removed from the content candidate set. Reintegration into the initial focus set The initial focus set after bidirectional correction is regarded as the focus set. .

[0138] Specifically, the comparable works set is a subset of all works (cultural images) in the knowledge graph, selected according to the following criteria: 1) Cultural images that are adjacent to or the same dynasty as the target cultural image; 2) Cultural images that are consistent with the narrative theme category of the target cultural image; 3) The top-K cultural images (e.g., Top-K=8) of cross-work comparison edges associated with the target cultural image retrieved based on the KNN algorithm in the knowledge graph. Meeting any one of these criteria qualifies a work as a comparable works set. Entity patterns are the statistical regularities of entity co-occurrence combinations in the comparable works set, including: 1) High-frequency co-occurrence combinations, i.e., entity combinations that appear repeatedly in the same theme, such as the combination of literati, pine trees, guqin (a seven-stringed zither), and mountain dwelling in the reclusive theme; 2) Co-occurrence frequency, i.e., the narrative importance of an entity combination in the same scene, such as pine trees having a high co-occurrence frequency in the reclusive theme, while background rocks have a low co-occurrence frequency in the reclusive theme. Matching is a category comparison of element labels in a controlled vocabulary, which does not require consistent image region coordinates or initial confidence levels; it only requires matching the initial focus set. The co-occurrence frequency of each entity combination in the dataset is statistically analyzed against that of entity combinations in comparable works. The probability of each entity appearing in similar themes is then determined to evaluate iconographic saliency. Only semantic category matching is required; visual pixel-level consistency is not necessary. For example, for L3 semantic candidates... Statistics on its role in knowledge graphs Combination of triggering content clues The ratio of the number of culturally related images to the total number of cultural images in comparable collections is the corresponding co-occurrence frequency.

[0139] Based on the focal set of the target culture image Determine the corresponding descriptive claims. Each descriptive claim (DBullet) output in the description phase is presented in a first preset format, such as "person—object—environment—interaction relationship". The specific format can be adjusted according to actual needs. Each D-stage claim in each descriptive claim carries a corresponding L1 evidence ID and corrected confidence level, strictly gated to the L1 content layer fields of the controlled vocabulary, forming the auditable factual basis for subsequent stages. Person, object, environment, and interaction relationship correspond to the four D-stage claims in the descriptive claim. The L1 evidence ID is the content candidate set. The corresponding L1 content entity node (L1Item) is assigned a unique number (i.e., a knowledge graph number). The L1 evidence ID (corresponding node number) is formatted as L1-{segId}-{serial number}, for example, L1-001-03, where segId represents the index of the descriptive claim. Each stage D claim must reference at least one L1 evidence ID during its generation, which is equivalent to tagging each descriptive claim with its source, making each descriptive claim traceable.

[0140] The corrected confidence level of the D-stage claim in the description phase output is given by the following formula:

[0141]

[0142] in, The corrected confidence level for the current Phase D claim; The current D phase advocates for the related content layer. The initial confidence level of the spatial relation edges; This refers to the co-occurrence frequency claimed in the current D phase; The preset description weight parameter is preferably 0.6 to 0.8, and in this embodiment, it is 0.7.

[0143] S23, Analysis Phase (A) – L2 Gated Formal Language Analysis:

[0144] During the analysis phase, ontology gating constraints are applied, allowing access only to the formal layer L2 of the controlled vocabulary, focusing on the set of focal points of the target cultural image. Using anchor points, for the formal candidate set Formal language analysis is performed using a multimodal large language model to obtain corresponding analytical claims, along with corresponding L2 evidence IDs and revised confidence levels. The formal language analysis includes:

[0145] 1) Micro-formal analysis: Extracting the focal set applied to the target cultural image. The brushwork techniques and ink coloring of each entity are used as technical characteristics;

[0146] 2) Macro-composition analysis: Extract the compositional layout of the target culture image as compositional features;

[0147] 3) Based on the knowledge graph G, calculate the historical form convention conformity of technique features, composition features, or combinations of technique features and composition features to determine whether they conform to the recurring formal expression rules of cultural images in historical documents, and obtain the modified confidence level based on the historical form convention conformity.

[0148] Each analytical claim output during the analysis phase is presented in the second preset format, "Composition—Brushwork—Coloring," which can be adjusted according to actual needs. Each A-stage claim within the analytical claim carries a corresponding L2 evidence ID and a revised confidence level, and is passed into the interpretation phase as high-priority formal evidence. Composition, brushwork, and coloring are the three A-stage claims within the analytical claim. The L2 evidence ID is a formal candidate set. The corresponding L2-form feature node (L2Item) is assigned a unique number (i.e., knowledge graph). The L2 evidence ID (corresponding node number) is formatted as L2-{anaId}-{serial number}, for example, L2-001-03, where anaId represents the index of the analytical claim. The multimodal large language model uses the fourth cue word template to obtain the output of the analysis phase, as follows:

[0149]

[0150] The revised confidence level of the Phase A claim output from the analysis phase is given by the following formula:

[0151]

[0152]

[0153] in, The revised confidence level for the current Phase A claim; The initial confidence level of the spatial relation edges in the current A-stage assertion of the associated form layer L2; The degree of conformity of the historical forms and conventions advocated in the current D phase; The preset analysis weight parameter is preferably 0.6 to 0.8, and in this embodiment, it is 0.7; It can be a technical feature, a compositional feature, or a combination of technical and compositional features; For the target cultural images and comparable works collection No. Similarity scores for cross-work comparisons of cultural images; The first in the comparable works collection Cultural Images The formal layer, M, represents the number of cultural images in comparable collections; This is an indicator function; it takes the value 1 if the condition within the parentheses is true, and 0 otherwise.

[0154] S24, Interpretive Stage (I) – Explanation of the Significance of L3 Gated Culture:

[0155] The interpretation phase (I), based on the established content layer (L1) and formal layer (L2) of the target cultural image, enters the semantic space of the meaning layer (L3) for the first time. The interpretation phase includes three sub-phases: thematic interpretation, atmosphere interpretation, and value synthesis. The thematic interpretation and atmosphere interpretation sub-phases are completed in parallel before the value synthesis sub-phase begins. The value synthesis sub-phase is further divided into the description phase (D), the analysis phase (A), and the thematic interpretation sub-phase (L3). ), Atmosphere Explanation Sub-phase ( The output of the explanatory stage is obtained using a multimodal large language model. The output of the explanatory stage includes the outputs of the topic explanation sub-stage, the atmosphere explanation sub-stage, and the value synthesis sub-stage. Specifically:

[0156] (1) Thematic Explanation Sub-stage ( )

[0157] Based on the focal set of the target culture image Activate symbol mapping rule base The cross-layer triggering rules that satisfy the minimum evidence set are activated, and the generated L3 semantic candidates corresponding to the activated cross-layer triggering rules form the L3 candidate set.

[0158] The output of the topic interpretation sub-stage includes L3 semantic candidates in the L3 candidate set, along with the corresponding rule node ID, L1 evidence ID, L2 evidence ID, L3 meaning node ID, text evidence node ID associated with the cross-work comparison edge, and the first confidence level. The L3 semantic candidates in the L3 candidate set are either topic type or specific motif, and the L3 meaning node ID is the ID of the L3 semantic candidate itself.

[0159] (2) Atmosphere Explanation Sub-stage ( )

[0160] Based on the focal set of the target culture image The output of the analysis phase yields L3 atmosphere slots, which include global mood, character emotions, and environmental atmosphere. The cross-sensory consistency test of the knowledge graph G on the L3 semantic candidates is then completed. The output of the atmosphere interpretation sub-phase is the L3 semantic candidates corresponding to the three atmosphere slots, each carrying the corresponding L1 evidence ID, L2 evidence ID, and second confidence level.

[0161] (3) Value synthesis sub-stage ( )

[0162] The overall description phase (D), the analysis phase (A), and the theme interpretation sub-phase (D) ), Atmosphere Explanation Sub-phase ( The output of the cultural context (the I-stage claim output of the value synthesis sub-stage) is obtained and carries the evidence ID and synthesis confidence of the theme interpretation sub-stage and the atmosphere interpretation sub-stage.

[0163] The overall confidence level of the Phase I claims output by the value synthesis sub-phase The formula is as follows:

[0164]

[0165]

[0166] in, The observed confidence level is the average of the corrected confidence levels of all L1 evidence IDs and the revised confidence levels of all L2 evidence IDs that trigger the current Phase I claim. The activation strength of the rule, i.e., the symbol mapping rule base. The initial weight of the triggering edge that triggers the current I-stage claim is multiplied by the minimum evidence set satisfaction. The minimum evidence set satisfaction is the ratio of the total number of elements in the content layer L1 and the form layer L2 that trigger the current I-stage claim to the total number of all elements in the minimum evidence set. To verify the consistency score of the knowledge graph; The evidence chain integrity score is calculated by dividing the sum of the number of L1 evidence IDs, L2 evidence IDs, and text evidence node IDs associated with cross-work comparison edges claimed in the current Phase I by a preset number of evidence nodes. The preset number of evidence nodes includes the sum of the number of rule node IDs, L1 evidence IDs, L2 evidence IDs, L3 meaning node IDs, and text evidence node IDs associated with cross-work comparison edges claimed in the current Phase I, as preferred in this embodiment. The values ​​are 0.30, 0.25, 0.30, and 0.15, respectively.

[0167]

[0168] in, This refers to the co-occurrence frequency claimed in the current Phase I. For the co-occurrence reward coefficient, when When the frequency is greater than or equal to a preset frequency threshold (e.g., 0.30), It should be 1.2, otherwise It is 1.0; The coefficient for the proof-of-contrast penalty is used when, in the current phase I, a proof-of-contrast edge is claimed to exist. Take a value less than 1, such as 0.7, otherwise... Set the value to 1.0; The evidence suppression coefficient is used when the current stage I claim lacks supporting textual evidence nodes or does not trigger L3 significance nodes. Take a value less than 1, such as 0.8, otherwise... Take 1.0.

[0169] The multimodal large language model uses the fifth cue word template to obtain the output of the interpretation stage, as follows:

[0170]

[0171] L3 semantic candidates with a comprehensive confidence level higher than the second preset threshold (0.75) are the primary explanations (at least one). L3 semantic candidates with a comprehensive confidence level lower than the second preset threshold are retained with uncertainty labels instead of being discarded, in order to ensure the auditability of the reasoning. That is, when the comprehensive confidence level is lower than the second preset threshold, the reason for insufficient evidence (auxiliary clues) is labeled, such as "The dynasty context is questionable: This social function is mainly seen in the Northern Song court, which is inconsistent with the context of the late Ming and early Qing dynasties in the input work; at the same time, there is a lack of textual evidence nodes to support it."

[0172] When evidence is insufficient, competing candidates (i.e., controversial L3 semantic candidates) are retained with uncertainty annotations (caution notes), rather than being forced to generate a single definitive conclusion. Each claim can be fully traced back to the underlying visual observation through the evidence ID, achieving end-to-end transparency and questionability of reasoning. Retaining competing candidates instead of discarding them has the following effects: 1) Supporting the auditability of reasoning: The core commitment of the DAIE report is that every conclusion can be traced and questioned. If candidates are discarded directly when there is insufficient evidence, and only one conclusion is output, the reviewer cannot know which possibilities were considered, nor can they judge whether the selection is reasonable. Retaining competing candidates with uncertainty annotations makes the entire reasoning process transparent; 2) Supporting human-machine collaborative review: Domain experts can see all low-confidence competing candidates and their uncertainty annotations, and manually confirm, refute, or revise their judgments. If candidates are directly discarded, experts lose the opportunity to intervene and correct errors; 3) Preventing cultural interpretation illusion: Forcing the output of a single definitive conclusion is one of the main mechanisms that generate cultural interpretation illusion, that is, it will still generate semantically fluent but unfounded interpretations when there is insufficient evidence. Retaining competing candidates with uncertainty annotations is a hard constraint to avoid forcibly drawing conclusions.

[0173] S25, Evaluation Phase (E) – Cross-contextualized assessment:

[0174] The evaluation phase prohibits the generation of new internal semantics within the cultural image, strictly limiting it to placing the target cultural image within a broader historical and cultural context. The confidence levels from each stage—Description Phase D, Analysis Phase A, and Interpretation Phase I—are passed to the evaluation phase. The claims in the evaluation phase are generated by a multimodal large language model using a sixth cue word template, carrying the output of the interpretation phase and the evaluation confidence level (Conf_E). This academically comparable cross-work argumentation prevents impressionistic understandings arising from unconstrained generation.

[0175] The output of the evaluation phase is hard-constrained into the following three dimensions:

[0176] Stylistic lineage positioning: Assess the position of cultural images in the stylistic lineage of painting history, such as whether they represent the Southern Song court style or embody the "Ma Yuan's corner" composition tradition;

[0177] Painter's Creative Positioning: Evaluate how the painter organizes and expresses his creative intentions through composition, blank space, brushwork, and specific motifs, such as using a fisherman in a slumber, riverbank ripples, and large areas of blank space to create a reclusive lifestyle;

[0178] Philosophical foundation positioning: assess the ideological basis of cultural images, such as the visual transformation of "principle," "mind," and inner cultivation in Song Dynasty Neo-Confucianism and Zen Buddhism.

[0179]

[0180] In this embodiment, the evaluation confidence level Conf_E claimed in the current E stage is calculated as follows:

[0181]

[0182] in, The overall confidence level of the current Phase E claim; To support the cross-works argument of the current E phase, The textual support for the current Phase E claims; The uncertainty penalty term claimed in the current E phase, as preferred in this embodiment. The values ​​are 0.35, 0.30, 0.25, and 0.10, respectively.

[0183] Among them, the cross-work support level advocated in the current E phase ( The calculation is as follows:

[0184]

[0185] in, This represents the average similarity score of all cross-work comparison edges related to the current E-stage claim. This is the co-occurrence support score claimed in the current E phase, which is the average weight of all co-occurrence support edges associated with the current E phase claim.

[0186] The textual support for the current E-phase claims ( The calculation is as follows:

[0187]

[0188] in, The normalized score of the textual evidence claimed in the current Phase E argument. , This indicates taking the minimum value. The number of textual evidence nodes connected to the current E-phase claim. The preset number of text evidence nodes is set, for example, 3; The reliability score of the text evidence claimed in the current E stage is determined based on the average of the preset source coefficients of the text evidence corresponding to the text evidence nodes claimed in the current E stage. For example, the preset source coefficient for the source of 36 painting works is 0.90, the preset source coefficient for the source of museum archives is 1.00, and the preset source coefficient for the source of encyclopedias is 0.80. The text evidence connection strength is the average weight of all co-occurring support edges of the associated text evidence.

[0189] For example, the current E phase advocates connecting two text evidence nodes. If we take 3, then =2 / 3=0.67; the textual evidence includes one museum archival document and one encyclopedia, then... =(0.90+0.80) / 2=0.85; The average weight of all co-occurring supporting edges of the associated textual evidence is 0.80, therefore:

[0190] = 0.40×0.67 + 0.40×0.85 + 0.20×0.80 = 0.768.

[0191] The uncertainty penalty term currently advocated in Phase E ( The calculation is as follows:

[0192]

[0193] in, This is the average weight of all the counter-evidence edges that are currently asserted in stage E; The uncertainty labeling score is the one advocated in the current E stage. When there is no uncertainty label, it is 0; when there is uncertainty label, it is 0.20~0.80, preferably 0.60.

[0194] For example, if the current E-stage asserts that there is no contradictory edge, then... =0; if uncertainty is marked as "none", then =0, calculated to obtain = 0.60×0 + 0.40×0 = 0. If there exists a proof-by-contrast edge with a corresponding weight of 0.70 and with uncertain labeling, then the calculation yields: = 0.60×0.70 + 0.40×0.60 =0.66.

[0195] Step 26: Summarize the outputs of the four stages—Description Stage D, Analysis Stage A, Interpretation Stage I, and Evaluation Stage E—to generate the Cultural Image Recognition Report Y_DAIE, where:

[0196] The output of the description phase D includes several description claims. Each description claim includes four D-phase claims (people, objects, environment, and interaction relationships). Each D-phase claim carries an L1 evidence ID and a corrected confidence level.

[0197] The output of analysis phase A includes several analytical claims, each of which includes three phase A claims (composition, brushwork, and coloring). Each phase A claim carries an L2 evidence ID and a revised confidence level.

[0198] The output of Explanation Phase I includes several explanatory claims, each comprising four Phase I claims: two Phase I claims from the Theme Explanation sub-phase (the main explanation and competing candidates), one Phase I claim from the Atmosphere Explanation sub-phase, and one Phase I claim from the Value Synthesis sub-phase. Specifically, the two Phase I claims from the Theme Explanation sub-phase (L3 semantic candidates in the L3 candidate set) carry the corresponding rule node ID, L1 evidence ID, L2 evidence ID, L3 meaning node ID, textual evidence node ID associated with the cross-work comparison edge, and a first confidence level. The L3 semantic candidates in the L3 candidate set are either theme type or specific motif; the L3 meaning node ID is the ID of the L3 semantic candidate itself. The Phase I claims from the Atmosphere Explanation sub-phase (L3 semantic candidates corresponding to the three atmosphere slots) carry the corresponding L1 evidence ID, L2 evidence ID, and a second confidence level. The Phase I claim from the Value Synthesis sub-phase (cultural context) carries the evidence IDs from the Theme Explanation sub-phase and the Atmosphere Explanation sub-phase, along with a comprehensive confidence level.

[0199] The output of evaluation phase E includes an evaluation claim, and each evaluation claim includes an E-phase claim. Each E-phase claim (stylistic lineage positioning, painter's creative positioning, and philosophical foundation positioning) carries the evidence ID and evaluation confidence level Conf_E from the output of interpretation phase I. When the evaluation confidence level of the E-phase claim... If the cultural image recognition report is deemed reliable, it is considered reliable; otherwise, it is considered unreliable.

[0200] For ease of understanding, the present invention will be described in detail below with reference to the embodiments.

[0201] Input: Prior knowledge space K, target cultural image X and corresponding work metadata. The work metadata of target cultural image X is "Autumn River Fishing Retreat" painted by Song Dynasty painter Ma Yuan. It is 37cm in height and 29cm in width. The painting is set against the background of fishing boats and reeds on an autumn river, depicting a scene of a fisherman sleeping soundly with his oars tucked in at the bow of the boat.

[0202] (1) Prior knowledge space Pre-built knowledge graph It includes 1,323 figure paintings from the Tang to the Qing dynasties, covering six major themes: elegant gatherings, ladies, religious myths, historical narratives, scholars, and folk scenes. The knowledge graph includes entity relationship edges, spatial relationship edges, co-occurrence support edges, counter-evidence relationship edges, rule triggering edges, and cross-work comparison edges.

[0203] (2) Receive the target cultural image X and the corresponding work metadata, confirm that the dynasty is "Song Dynasty" and the work type is "book format", i.e., book.

[0204] (3) Based on the first prompt word template, candidate features are extracted using the GPT-5.4 model to generate a content candidate set. (Including: bamboo, inscriptions, fishing rods, headscarves, fishing rods and straw hats, reeds, fishermen, resting, simple clothes, fishing boats, oars, water ripples, riverbanks) and a collection of candidate forms. (Including: corner composition, combination of meticulous and freehand brushwork, light ink rendering, simplified brushwork, and blank space), and retaining the work's metadata (including album format and paper).

[0205] (4) Description phase: Ontology gating locks the fields of content layer L1. Anchor point identification confirms the target entity as "fisherman"; through relation center expansion, fishing boat, oar, fishing rod, headscarf, cloth, bamboo, reeds, water ripples, and riverbank are included in the initial focus set. Correction operators for knowledge graphs By searching for similar works on the theme of fishermen and comparable landscape paintings from the Southern Song Dynasty, it was confirmed that the combination of "fisherman-fishing boat-riverbank-bamboo / reeds" has iconographic salience in the context of fishermen's seclusion and is therefore retained; the focal set of the target cultural image is output. ={Fisherman, headscarf, cloth clothes, fishing rod, fishing boat, oar, bamboo, reeds, water ripples, riverbank, inscription}, with corresponding L1 evidence ID and confidence level. Figure 4 In the target cultural image, the background, including bamboo, reeds, and oars, is rendered in light ink, creating an atmosphere that embodies elegance and refinement. The reeds and fishing boats symbolize wandering, and the fisherman lies prone on the fishing boat on the water ripples, holding the oar, showing the fisherman taking a short nap. The quiet and serene atmosphere reflects the dialogue between the literati and the fisherman, symbolizing the virtuous conduct of a gentleman and the spirit of reclusion.

[0206] (5) Analysis phase: Ontology gating locks the fields of form layer L2. Micro-form analysis identifies the brushwork technique as "using simplified brushstrokes and combining meticulous brushwork with freehand brushwork, with simple lines that contain control", and the ink color as "light ink rendering, creating a hazy feeling with ink layers"; macro-composition analysis identifies the composition layout as "corner composition, the main subject is compressed in one corner of the picture, and a large area of ​​white space forms a spacious space", and at the same time confirms that the work type is "album format" and the material is "paper"; based on the knowledge graph It confirms that this type of "corner composition, blank space and light ink rendering" has historical formal conventions to support it in the Southern Song court style and Ma Yuan style; outputs the corresponding analytical claims, with corresponding L2 evidence IDs and confidence levels.

[0207] (6) Interpretation stage: Fields that enter the meaning layer L3 for the first time.

[0208] Theme Explanation Sub-phase ( ): Activation symbol mapping rule base of the focus set of the target cultural image. The minimum evidence set (MES) corresponding to rule-023 is {fisherman∧fishing boat∧riverbank∧blank space}, which is activated when the condition is met, triggering L3 semantic candidates: ① "literati's reclusive lifestyle of fishing", with an initial weight of 0.71; ② "wandering and fisherman's slumber", with an initial weight of 0.58; Knowledge graph G performs Song Dynasty context verification: "literati's reclusive lifestyle of fishing" has a significantly higher co-occurrence frequency in comparable works on the theme of fishermen and small landscape paintings in the Southern Song Dynasty than other L3 semantic candidates, and has a stable association with "corner composition + large blank space + fisherman theme"; the final comprehensive confidence score ( L3 semantic candidate ① =0.86, L3 semantic candidate ② =0.74; Output L3 semantic candidate ① as the main explanation, and L3 semantic candidate ② as an auxiliary clue (with uncertainty labeling) retained.

[0209] Atmosphere Explanation Sub-phase ( ): Candidate set in L2 form The primary inputs are "light ink rendering", "white space" and "corner composition", which trigger L3 semantic candidates "tranquil and lonely" (second confidence 0.77) and "sparse and elegant" (second confidence 0.76). Further, the descriptive information such as "fishermen resting", "fishing boat", "riverbank", "water ripples" and "bamboo and reeds" are combined to strengthen the atmosphere of "quiet and peaceful, a moment of rest, and elegant and refined".

[0210] Value synthesis sub-stage ( ): Comprehensive description phase (D), analysis phase (A), topic interpretation sub-phase ( The output of the "literati's reclusive lifestyle" and the atmosphere explanation sub-stage ( The output of “a tranquil and serene state of mind, a fisherman’s solitude and detachment, and a vast and desolate environment with misty waters, a chilly autumn river, and a quiet and secluded atmosphere” confirms through knowledge graph G that “fisherman’s seclusion” has long-term textual support in comparable works of Song Dynasty painting and small landscape painting; the output of cultural context: the pursuit of seclusion and inner peace.

[0211] (7) Evaluation stage: Based on the knowledge graph G, the nodes of the Southern Song court paintings, Ma Yuan's "Ma Yijiao" corner composition system, and the related cross-work relationship edges of "Fisherman's Picture" and "Jiangzhu Xiaojing" from the same period were retrieved. It was found that the combination of "fishing boat, fisherman, blank space, and corner composition" has a high co-occurrence relationship with the reclusive atmosphere of the Southern Song Dynasty; the comprehensive confidence level (Conf_I) of the interpretation stage is 0.86 for Iv-01, the cross-work support level (Comp) is 0.99, and the text source support level (... The value is 1.00, and the uncertainty penalty term is ( The confidence level was calculated to be 0.84, with a value of 0.05. The output assessment assertions are: Stylistic lineage positioning = representing the style of the Southern Song Dynasty painting academy; Ma Yuan established his signature "Ma Jiao" (Ma Yijiao) style during this period; Artist creative positioning = Ma Yuan came from a family of court painters, yet he yearned for mountains and streams, longing for a secluded life. The emotions in his paintings shifted from tranquility to a yearning for reclusive culture, a fusion of the Taoist hermit spirit and the literati self-projection tradition of the fisherman archetype; Philosophical foundation positioning = a synthesis of Neo-Confucianism (Neo-Confucianism) and Zen Buddhism in the Song Dynasty: advocating the inner cultivation of "Li" (moral principles) and "Xin" (the mind of reflection).

[0212] (8) Generate a DAIE structured report (cultural image recognition report), such as Figure 5 As shown. All 12 claims carry evidence IDs and confidence levels. The outputs for each stage are shown in Table 2 below:

[0213] Table 2

[0214]

[0215] According to Table 2, all 12 claims include 4 Phase D claims, 3 Phase A claims, 4 Phase I claims, and 1 Phase E claim, each accompanied by a corresponding evidence ID, confidence level, and uncertainty label. Low-confidence L3 semantic candidates with uncertainty labels are retained, and no claims lack source links.

[0216] like Figure 5 As shown, the DAIE structured report generated in this embodiment is as follows:

[0217] D. Description (Description Phase):

[0218] A fisherman wearing a headscarf and simple clothes lies on his side on a fishing boat moored along the river with his eyes slightly closed. The boat is moored on the riverbank in autumn, where bamboo branches and reeds dot the shore.

[0219] A. Analysis (Analysis Phase):

[0220] The work is presented on the scale of a small album leaf, employing a corner composition strategy with a sparse and singular focal point, leaving large areas of blank space: deeply embodying the principles of "treating white as black" and "a corner of a horse." The brushwork combines meticulous detail with freehand strokes, using simplified outlines to create concise yet rigorous lines. The colors are applied purely with ink, using the varying shades and tones of ink to replace all other colors.

[0221] I. Explanation (Explanation Stage):

[0222] The resting fisherman in the painting is elevated to the theme of "fishing in seclusion," transforming a daily rest into an ideal state of "living in seclusion among the rivers and lakes." The overall artistic conception presents a serene and distant style; the fisherman is fishing alone, his expression serene and detached, leisurely and content; the surrounding waters are misty and the autumn river is clear and cold: the environment is spacious and quiet. Its deeper value system integrates the Taoist spirit of reclusion and the tradition of literati using the image of the fisherman to express their personal aspirations.

[0223] E. Evaluation (Evaluation Phase):

[0224] This work belongs to the Southern Song court style, encoding the mindset of a dynasty in its concise composition of a small landscape. Ma Yuan's signature corner framing and simple brushwork create a lonely autumnal atmosphere, reflecting the painter's yearning for seclusion in the mountains. Philosophically, it advocates the inner cultivation of "reason" (moral principles) and "heart" (reflective mind); the "fishing retreat" motif continues the long-standing tradition of literati self-projection, inviting viewers into a state of "hesitation."

[0225] Figure 5 For visualization purposes, different colors represent different types of evidence. For example, tan represents L1 evidence, light green represents L2 evidence, light orange represents L3 evidence, and light red represents contextual evidence. L1 evidence corresponds to elements in the content layer (L1) of the controlled vocabulary; L2 evidence corresponds to elements in the form layer (L2) of the controlled vocabulary; L3 evidence corresponds to elements in the thematic narrative and atmosphere of the meaning layer (L3) of the controlled vocabulary; and contextual evidence corresponds to elements in the work's metadata and the cultural context of the meaning layer (L3) of the controlled vocabulary. For each color-coded evidence type, hovering the mouse over the text in the corresponding color area will display the evidence ID and confidence level, as shown in Table 2, facilitating traceability and reference for operators.

[0226] To verify the effectiveness of this invention, the experimental procedure is as follows:

[0227] 1) Experimental setup:

[0228] Benchmarking. This experiment constructed a benchmark set containing 100 cultural images (figure paintings), covering five major dynasties (Tang, Song, Yuan, Ming, and Qing) and six thematic categories: elegant gatherings, court ladies, religious mythology, historical narratives, recluses, and folk scenes. The works are evenly distributed across scrolls, handscrolls, and album leaves. Each cultural image is accompanied by metadata (artwork title, artist, dynasty, art type, material, and collection information) and references from museum catalogs and art history research. To test performance both within and outside the knowledge graph's coverage area, the benchmark set (all) is divided into 50 cultural images within the knowledge graph (in-KG subset) and 50 cultural images outside the knowledge graph (out-of-KG subset). The internal knowledge graph subset represents images included in the knowledge graph. Of the 1,323 cultural images, the subset of external knowledge graphs represents those not included in the knowledge graph. Of the 1,323 cultural images.

[0229] Among them, 35 cultural images were annotated to form a gold subset consisting of several fine-grained gold reports. Each gold report in the gold subset contains a DAIE structured report annotated by experts. In addition, 8 handscroll cultural images from the 100 cultural images (each divided into 5-8 scene fragments, totaling approximately 48 scene fragments) constitute an independent narrative subset for evaluating the narrative coherence of multi-scene handscrolls, that is, assessing whether narrative consistency can be maintained across paragraphs when dealing with handscrolls containing multiple scene fragments.

[0230] 2) Baseline model:

[0231] The GPT-5.4 model is used as the multimodal large language model, with a temperature parameter T=0.3 (used to control the randomness of the multimodal large language model output; the smaller the temperature parameter, the more deterministic the output). The maximum output length is 4,096. During each stage's query to the knowledge graph, the top 8 nodes semantically closest to the current stage are retrieved as evidence input. The proposed method (TCFP-Sight method) is compared with four baseline models, adopting the principle of "same backbone, different mechanisms," i.e., maintaining consistency in the multimodal large language model. B1 (MLLM-Direct) outputs directly without using a corpus (knowledge graph). B2 (MLLM-Structured) also does not use a corpus but performs the DAIE inference stage. B3 (Text-RAG) retrieves and accesses the same corpus as the TCFP-Sight method and outputs directly after retrieval. B4 (ArtRAG) retains the ArtRAG multi-granularity subgraph process but replaces the original knowledge graph with the same corpus. The experimental results are shown in Tables 3 and 4. An upward arrow (↑) indicates that the corresponding value is as large as possible, and a downward arrow (↓) indicates that the corresponding value is as small as possible.

[0232] Table 3

[0233]

[0234] Table 4

[0235]

[0236] The evaluation system is organized into three complementary levels, specifically:

[0237] Level 1: The GPT-5.4 model is used as the judging model (all 100 cultural images), equipped with dimension-specific scoring criteria. Each output is scored independently, and a 5-point Likert scale is used to evaluate four dimensions (each dimension is scored from 1 to 5 by the GPT-5.4 model for a single cultural image, and the average score for all 100 images is taken to obtain the corresponding dimension score): Factual basis (FG) is used to measure the accuracy of elements in the content layer L1 and the formal layer L2 relative to the cultural image; Depth of formal analysis (FA) is used to measure the accuracy of composition and brushwork techniques; Cultural iconography interpretation (CI) is used to measure the accuracy of the outputs of the three sub-stages of the interpretation stage; Argument transparency (AT) is used to measure whether the cultural semantics cites L1 and L2 evidence.

[0238] Level Two: Auditability and Illusion Indicators (Golden Subset). Illusion indicators include: Evidence Support Rate (ESR), which is the ratio between the number of claims in Stage I and Stage E that simultaneously carry L1 and L2 evidence, and the total number of claims in Stage I and Stage E; Gated Compliance Rate (GCR), which is the proportion of all stages in the DAIE reasoning stage that output elements belonging to the corresponding stage's domain body; and Hall. Total. Hall. Total is decomposed into four types: OH (Entity Illusion), FM (Formal Misreading), IH (Imaginary Illusion), and CM (Context Mismatch), as shown in Table 4. OH is the number of entities not present in the cultural image, FM is the number of misjudgments of elements in the formal L2 layer, IH is the number of elements in the semantic L3 layer lacking support from L1 and L2 evidence, and CM is the number of misjudgments of dynasty. Hall. Total = OH + FM + IH + CM.

[0239] The third level: expert evaluation (30 cultural images). Five evaluators with art history backgrounds conducted a blind evaluation. The evaluators scored the four dimensions using a 5-point Likert scale. The results showed a high degree of convergence between the TCFP-Sight method and the expert evaluation, further enhancing the robustness of the overall evaluation conclusions.

[0240] 3) Main results:

[0241] Overall Performance. The TCFP-Sight method achieved best results in all five critical metrics (CI, AT, ESR, GCR, and Hall. Total). Compared to the strongest baseline model (CI of 3.52 in B4), its CI score increased by 0.39 points (+11.1%), a difference approximately 2.8 times the baseline CI range, indicating a qualitative leap in cultural semantic reasoning that the baseline model could not match. The GCR score was 17.3 percentage points higher than the strongest baseline model (GCR of 0.743 in B2), confirming that the gating mechanism achieved a breakthrough in stage-level lexical compliance that structured prompts (structured JSON files) could not match.

[0242] Structured cues and the cost of hallucinations. B2 achieved the highest FG and FA, confirming that structured cues improve the surface quality of the text. However, B2 also produced the highest total number of hallucinations (428) – a 22.3% increase compared to the uncued B1 (350). Table 4 shows that although B2 reduced OH (247 compared to B1's 273), its IH (70) and CM (69) soared to 3.2 times and 3.0 times that of B1, respectively. The TCFP-Sight method performed worse than B2 in FG and FA metrics because its controlled vocabulary system trades narrative richness for higher reliability – a trade-off that yields significant returns in AT, ESR, and hallucination suppression. Importantly, this reduction is not due to a decrease in the amount of content generated: on the golden subset, the TCFP-Sight method generated 459 stage I and stage E claims (compared to 403 for B2), but its advanced illusion rate (the sum of IH and CM divided by the number of generated claims) was 0.039—nearly an order of magnitude lower than B2's 0.345. This finding validates the irreplaceable nature of the hierarchical field constraints imposed on the DAIE inference stages.

[0243] Hallucination type analysis. Among all four hallucination types, the TCFP-Sight method had the lowest frequency of occurrence. The reduction in the CM index was the most significant—from 23 cases (B1) to 5 cases (a decrease of 78.3%), which directly validates the correction operator of the knowledge graph. The effectiveness of the IH and CM indicators in verifying dynastic background. As the two most typical failure modes in cultural interpretation, the frequency of occurrence of the IH and CM indicators decreased from 45 cases (the combined value of B1) to 18 cases (a decrease of 60.0%), providing the main empirical support for the ability to resist illusions in high-level cultural reasoning.

[0244] A comparison of the generalization capabilities of internal knowledge graph subsets (in-KG) and external knowledge graph subsets (out-of-KG). The internal and external knowledge graph subsets show negligible differences in AT (3.40 vs. 3.39), ESR (0.912 vs. 0.917), and GCR (0.921 vs. 0.911). "vs." indicates a control, suggesting that the main advantages of the evidence association mechanism (evidence tracing) and ontology gating mechanism stem from architectural constraints (DAIE reasoning stage) rather than knowledge graph retrieval results. For example, 3.40 vs. 3.39 represents the internal knowledge graph subset's AT index of 3.40 versus the external knowledge graph subset's AT index of 3.39; the rest are analogous. The internal knowledge graph subset shows a moderate advantage in the FA (3.28 vs. 3.04) and CI (3.98 vs. 3.84) metrics. The illusion suppression results for the external knowledge graph subset indicate that even for cultural images not included in the knowledge graph, the total number of illusions achieved by the proposed method (71 vs. 51) is significantly lower than the baseline model. Specifically, the total number of illusions for the internal knowledge graph subset (in-KG) is 51, and the total number of illusions for the external knowledge graph subset (out-of-KG) is 71, demonstrating good generalization ability. The core advantage of the TCFP-Sight method lies in its high auditability and low illusion rate, maintaining its advantage even outside the knowledge graph coverage area. This proves that its performance improvement does not stem from simple memorization of instances within the knowledge graph.

[0245] 4) Ablation test

[0246] Four ablation variants were evaluated on a golden subset containing 35 cultural images. Each ablation variant removed one core mechanism (including removal of the DAIE inference stage, removal of ontology gating (i.e., domain-free ontology and authorization vocabulary), removal of the knowledge graph, and removal of the symbolic mapping rule base). As shown in Table 5, the inter-stage violation rate (CVR) is the ratio of the number of elements in the L3 meaning layer output by stages D and A to the total number of elements output by stages D and A. HL is the high-level illusion rate, and Inst is the ratio of the average number of illusions in the cultural image to the total number of illusions.

[0247] Table 5

[0248]

[0249] The DAIE inference phase employs a multi-stage ontology gating mechanism. Directly using single-generation (without a DAIE inference phase and corresponding layer field constraints) leads to a severe performance degradation: the Evidence Support Rate (ESR) drops to 0.314 (a decrease of 0.598), the Inter-Stage Violation Rate (CVR) climbs to its peak (0.160), the Gated Compliance Rate (GCR) drops to 0.658, and both the IH and CM metrics reach their peaks with an advanced illusion rate as high as 0.658. Notably, the CI metric only decreases by 0.11, indicating that even without an evidence anchoring mechanism (corresponding layer field constraints), seemingly reasonable explanatory text can still be generated—confirming that relying solely on prior knowledge space without phased inference is insufficient.

[0250] Ontology gating ensures auditability. Removing ontology gating caused the AT (Accuracy Time) metric to drop from 3.40 to 2.00 (a decrease of 1.40), the largest single-metric decrease among all variants; the Gated Compliance Rate (GCR) decreased from 0.916 to 0.723, confirming that the stage-level ontology gating mechanism is a direct cause of compliance; the CVR (Conformity Rate) increased by 75.0%; and the Higher Illusion Rate (HL) nearly doubled to 0.422. Ontology gating effectively prevents premature activation of high-level decisions by strictly limiting semantic access at each stage.

[0251] Knowledge graphs are primarily used for fact verification; remove the correction operator from the knowledge graph. This only resulted in a slight decrease in the ESR index and a relatively unchanged GCR index. However, the total number of hallucinations increased by 32.4%, and the CM index increased by 1.9 times, confirming that the core role of knowledge graphs in suppressing dynastic background misattribution is for fact verification.

[0252] Symbol mapping rule base Activate controlled interpretation functionality. Removing the symbolic mapping rule base (R) results in the lowest FG metric and reduces the ESR metric to 0.755, but also records the lowest CM metric. This is because the lack of a rule-triggered activation mechanism reduces the generation of high-level claims, thereby lowering the violation rate and interpretability. The core function of rules is to provide stable and controllable candidate activation paths for high-level interpretations. Furthermore, through complementarity, it achieves an optimal balance between interpretation quality, auditability, and illusion control.

[0253] In summary, the TCFP-Sight method establishes a four-stage ontology gating mechanism on the prior knowledge space. Building upon this foundation, the ontology gating mechanism is used to delineate the boundaries of corresponding stages, reconstructing traditional figure painting recognition into a structured, empirically-based cognitive reasoning process. This approach suppresses cultural illusions while ensuring the traceability of complete claims. Ablation experiments confirm that the ontology gating mechanism is the most critical core component, and its staged reasoning cannot be replaced solely by prior knowledge space.

[0254] It should be noted that this invention is also applicable to the recognition of multi-scene narrative cultural images of handscrolls (handcrolls containing multiple consecutive scene segments). Specifically: for handscrolls containing multiple consecutive scene segments, they are divided into several scene segments, each scene segment corresponding to a scene segment node (Segment) in the knowledge graph (containing the semantic embedding vector of the corresponding scene segment). Similar to the single-scene segment recognition described above, an ontology gating mechanism is independently executed for each scene segment to perform multi-stage reasoning and generate a DAIE structured report for the corresponding scene segment. The difference lies in that cross-scene segment narrative coherence verification is additionally introduced in the evaluation stage (E): that is, using the semantic embedding vector of each scene segment as the query vector in the knowledge graph... Based on the KNN algorithm, the top-K nodes connected to the cross-work comparison edge of the corresponding scene fragment are retrieved as the corresponding scene fragment with the highest similarity score. The system identifies the evolution of narrative clues and thematic repetition patterns across scenes. The evolution of narrative clues is to compare the L3 semantic candidates of adjacent scene fragments to identify the direction of the evolution of narrative clues across scenes (such as from "travel" → "banquet" → "seclusion", or from "listening to music" → "watching dance" → "resting"). Thematic repetition patterns are identified by detecting the recurring motif combinations in multiple scene fragments to identify global thematic repetition patterns. For example, the combination of "recluse + zither + pine" appears repeatedly in multiple scene fragments of the handscroll, which strengthens the theme of seclusion.

[0255]

[0256] The multimodal large language model uses the seventh prompt word template to integrate cross-scene analysis results to generate a global narrative evaluation report, ensuring the coherence of cross-scene segments in the long scroll multi-scene narrative recognition.

[0257] Example 2:

[0258] A cultural image recognition system based on prior knowledge space and ontology gating includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements the cultural image recognition method based on prior knowledge space and ontology gating as described in Example 1.

[0259] It should be understood that specific limitations regarding the cultural image recognition system based on prior knowledge space and ontology gating can be found in the limitations of the cultural image recognition method based on prior knowledge space and ontology gating in Example 1, and will not be repeated here. Each module in the aforementioned cultural image recognition system based on prior knowledge space and ontology gating can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0260] The memory and processor are electrically connected directly or indirectly to enable data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines. The memory stores a computer program that can run on the processor. By running the computer program stored in the memory, the processor implements the cultural image recognition method based on prior knowledge space and ontology gating in Embodiment 1 of the present invention.

[0261] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), and Electrically Erasable Programmable Read-Only Memory (EEPROM). The memory stores computer programs, and the processor executes these programs after receiving execution instructions.

[0262] The processor may be an integrated circuit chip with data processing capabilities. The aforementioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc. It can implement or execute the methods, steps, and logic block diagrams disclosed in Embodiment 1 of this invention. The general-purpose processor can be a microprocessor or any conventional processor.

[0263] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0264] The embodiments described above are merely specific and detailed examples of the embodiments described in this application, and should not be construed as limiting the scope of the application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and improvements all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.

Claims

1. A cultural image recognition method based on prior knowledge space and ontology gating, characterized in that: Includes the following steps: S1. Obtain the cultural image set and construct the corresponding prior knowledge space. ,in, For the domain ontology, It is a knowledge graph, constructed from a controlled vocabulary. For symbol mapping rule base; S2. Based on the target cultural image and corresponding work metadata, as well as the prior knowledge space, an ontology-gated model is used to obtain the corresponding cultural image recognition report, as detailed below: S21. Based on the domain ontology and controlled vocabulary, a multimodal large language model is used to generate content candidate sets and form candidate sets for target cultural images; S22. Based on the work's metadata and content candidate set, use a multimodal large language model to extract the focus set of the target cultural image, and obtain the output of the description stage, including several description claims corresponding to the focus set, and each description claim carries an L1 evidence ID and a corrected confidence level for each D-stage claim. S23. Based on the focus set and the formal candidate set, use a multimodal large language model to obtain the output of the analysis phase, including several analysis claims, and each analysis claim carries a corresponding L2 evidence ID and a modified confidence level for each A-stage claim; S24. Based on the outputs of the prior knowledge space, focus set, description stage, and analysis stage, the output of the interpretation stage is obtained using a multimodal large language model. S25. Based on the outputs of the description, analysis, and interpretation phases, use a multimodal large language model to obtain the output of the evaluation phase, including an evaluation claim along with the corresponding output of the interpretation phase and the evaluation confidence level. S26. Summarize the outputs of each stage to generate a cultural image recognition report.

2. The cultural image recognition method based on prior knowledge space and ontology gating as described in claim 1, characterized in that: The domain ontology is used to divide the semantic elements of all cultural images in the cultural image set into a first semantic layer, a second semantic layer, and a third semantic layer, wherein: The first semantic layer includes the entities of the cultural image, entity attributes, spatial location and interaction relationships between entities; The second semantic layer includes composition, brushwork techniques, and ink coloring. The third semantic layer includes iconographic motifs, narrative themes, visual styles, and cultural semantics; The controlled vocabulary includes a set of legal fields, a set of enumerable tags, and inter-layer semantic relationship types. The set of legal fields includes several predefined fields corresponding to the first semantic layer, the second semantic layer, and the third semantic layer. The set of enumerable tags includes several predefined tags allowed for each field in each semantic layer. The inter-layer semantic relationship types include the indicative relationship between the first semantic layer and the second semantic layer, the evocation relationship between the second semantic layer and the third semantic layer, and the mapping relationship between the first semantic layer and the third semantic layer. The symbol mapping rule base includes several cross-layer triggering rules extracted from the knowledge graph. Each cross-layer triggering rule is represented as follows: In the formula, , , As a combination of content clues, Formal clues are combined. This is a candidate set of content, belonging to the first semantic layer of the controlled vocabulary. This is a formal candidate set, belonging to the second semantic layer of the controlled vocabulary. It is an L3 semantic candidate, belonging to the third semantic layer of the controlled vocabulary. For conjunction operation, This indicates that it has been triggered.

3. The cultural image recognition method based on prior knowledge space and ontology gating as described in claim 2, characterized in that: The types of nodes in the knowledge graph include work nodes, work metadata nodes, image region nodes, L1 content entity nodes, L2 formal feature nodes, L3 meaning nodes, text evidence nodes, and rule nodes. The types of edges in a knowledge graph include: Entity Relationship Edges: connect work nodes to work metadata nodes, connect work nodes to L1 content entity nodes, connect work nodes to L2 formal feature nodes, and connect work nodes to L3 meaning nodes; First spatial relation edge: connects the image region node and the L1 content entity node, carrying the corresponding initial confidence and image region coordinates; Second spatial relation edge: connects image region nodes with L2 form feature nodes, carrying the corresponding initial confidence and image region coordinates; The first co-occurrence support edge consists of an L1 content entity node that connects the L3 meaning node to the L1 content entity node that supports the L3 meaning node, and an L2 form feature node that connects the L3 meaning node to the L3 meaning node that supports the L3 meaning node, carrying corresponding weights. Second co-occurrence support edge: connects L3 meaning nodes and text evidence nodes, carrying corresponding weights; Rule-triggered edges: connect rule nodes to L3 significance nodes that trigger the rule node; Contrast-based edge: connects L3 meaning nodes with other L3 meaning nodes, connects work metadata nodes with L3 meaning nodes, and connects text evidence nodes with L3 meaning nodes, carrying corresponding weights; Cross-work comparison edge: connects image region nodes with similarity scores higher than a preset similarity threshold, and connects comparable nodes across works with similarity scores higher than a preset similarity threshold, carrying the corresponding similarity scores.

4. The cultural image recognition method based on prior knowledge space and ontology gating as described in claim 3, characterized in that: The extraction process for the focus set is as follows: Based on the work's metadata and content candidate set, a multimodal large language model is used to identify the target entities of the target cultural image, and the target entities, as well as entities with spatial location and interaction relationships with the target entities, are included from the content candidate set into the initial focus set. The work's metadata includes the work's name, author, dynasty, work type, material, and collection information. The initial focus set is matched and corrected with the entity patterns of comparable works in the knowledge graph using the knowledge graph's correction operator to obtain the focus set. The knowledge graph's correction operator is used to perform bidirectional correction, as follows: 1) When the co-occurrence frequency of entities in the initial focus set in the entity patterns of the comparable works set is less than the false detection threshold, the corresponding entity is removed from the initial focus set or the initial confidence of the spatial relation edge of the first semantic layer associated with the corresponding entity is reduced. The co-occurrence frequency is the ratio of the number of cultural images in the comparable works set that contain the corresponding entity to the total number of cultural images in the comparable works set. The comparable works set consists of a preset number of cultural images in the knowledge graph that are adjacent to or the same as the target cultural image in terms of dynasty, or have the same category of narrative theme, or have a pre-defined number of cultural images that have a pre-defined number of cross-work comparison edges that are related. The entity pattern includes all entity combinations of the corresponding entity in the comparable works set and their corresponding co-occurrence frequencies. 2) If an entity in the content candidate set is not included in the initial focus set because its initial confidence level is lower than the first preset threshold, but the co-occurrence frequency of the entity in the entity mode of the comparable works set is greater than or equal to the false detection threshold, then the corresponding entity will be re-included from the content candidate set into the initial focus set, and the initial focus set after bidirectional correction will be regarded as the focus set.

5. The cultural image recognition method based on prior knowledge space and ontology gating as described in claim 1, characterized in that: The multimodal large language model adopts the GPT-5.4 model; The description claim is represented by "person-object-environment-interaction relationship", where person, object, environment and interaction relationship are the four D-stage claims in the description claim, and L1 evidence ID is the ID of the corresponding element in the content candidate set; The analytical claims are represented by "composition-brushwork-coloring", where composition, brushwork, and coloring are the three A-stage claims in the analytical claims, and L2 evidence ID is the ID of the corresponding element in the formal candidate set.

6. The cultural image recognition method based on prior knowledge space and ontology gating as described in claim 3, characterized in that: The corrected confidence level of the D-stage claim is a weighted sum of the initial confidence level of the spatial relation edge associated with the first semantic layer in the current D-stage claim and the co-occurrence frequency of the current D-stage claim. The revised confidence level of the A-stage claim is a weighted sum of the initial confidence level of the spatial relation edge of the current A-stage claim associated with the second semantic layer and the historical form convention conformity of the current A-stage claim.

7. The cultural image recognition method based on prior knowledge space and ontology gating as described in claim 3, characterized in that: The explanatory phase includes: (1) Theme Explanation Sub-stage: Based on the corresponding cross-layer triggering rules in the focus set activation symbol mapping rule base, all generated L3 semantic candidates form the L3 candidate set; The output of the topic explanation sub-stage is obtained by using a multimodal large language model, including L3 semantic candidates in the L3 candidate set and carrying the corresponding rule node ID, L1 evidence ID, L2 evidence ID, L3 meaning node ID, text evidence node ID associated with cross-work comparison edge, and first confidence score. The L3 semantic candidates in the L3 candidate set are either topic type or specific motif. (2) Atmosphere Explanation Sub-stage: Based on the output of the focus set and analysis phase, the output of the atmosphere interpretation sub-phase is obtained using a multimodal large language model, including L3 atmosphere slots and carrying corresponding L1 evidence IDs, L2 evidence IDs and second confidence scores. The L3 atmosphere slots include global mood, character emotions and environmental atmosphere. (3) Value synthesis sub-stage: Based on the outputs of the description stage, analysis stage, theme interpretation sub-stage, and atmosphere interpretation sub-stage, the output of the value synthesis sub-stage is obtained using a multimodal large language model, including the cultural context and carrying the evidence IDs and synthesis confidence scores of the theme interpretation sub-stage and the atmosphere interpretation sub-stage. L3 semantic candidates with a comprehensive confidence level higher than the second preset threshold are used as the main interpretations, while L3 semantic candidates with a comprehensive confidence level lower than the second preset threshold are retained with uncertainty labels, which are the reasons for insufficient evidence.

8. The cultural image recognition method based on prior knowledge space and ontology gating as described in claim 7, characterized in that: The overall confidence level of the Phase I claim output by the value synthesis sub-stage is the weighted sum of the observed confidence level of the current Phase I claim, the rule activation strength, the knowledge graph verification consistency score, and the evidence chain integrity score, where: The observed confidence level is the average of the corrected confidence levels of all L1 evidence IDs that trigger the current Phase I claim and the revised confidence levels of all L2 evidence IDs. The rule activation strength is the product of the initial weight of the rule triggering edge that triggers the current I-stage claim in the symbol mapping rule base and the minimum evidence set satisfaction. The minimum evidence set satisfaction is the ratio of the total number of elements in the first semantic layer and the second semantic layer that trigger the current I-stage claim to the total number of all elements in the minimum evidence set. The minimum evidence set is the minimum cue combination formed by elements in the content layer L1 and the form layer L2 that trigger the current I-stage claim. The consistency score of the knowledge graph verification The formula is as follows: in, This indicates taking the minimum value. This refers to the co-occurrence frequency advocated in the current Phase I stage. For co-occurrence reward coefficient, For proof by contradiction, the penalty coefficient, The evidence suppression coefficient; The evidence chain integrity score is the sum of the number of L1 evidence IDs, L2 evidence IDs, and text evidence node IDs associated with cross-work comparison edges that are claimed to be traced in the current I stage, divided by a preset number of evidence. The preset number of evidence includes the sum of the number of rule node IDs, L1 evidence IDs, L2 evidence IDs, L3 meaning node IDs, and text evidence node IDs associated with cross-work comparison edges that are claimed to be traced in the current I stage.

9. The cultural image recognition method based on prior knowledge space and ontology gating as described in claim 8, characterized in that: The confidence level of the E-stage claim is a weighted sum of the overall confidence level, cross-work support, textual source support, and uncertainty penalty of the current E-stage claim, where: The cross-work support score is the weighted sum of the cross-work similarity and co-occurrence support scores of the current E-stage claim. The cross-work similarity score is the average of the similarity scores of all cross-work comparison edges associated with the current E-stage claim, and the co-occurrence support score is the average of the weights of all co-occurrence support edges associated with the current E-stage claim. The minimum value among the text source support degree of 1 and the weighted sum of the text evidence normalization score, text evidence reliability score and text evidence connection strength claimed in the current E stage, is given. The text evidence normalization score is the average of the preset source coefficients of the text evidence corresponding to the text evidence node claimed in the current E stage, and the text evidence connection strength is the average of the weights of all co-occurrence support edges of the associated text evidence claimed in the current E stage. The uncertainty penalty term is the minimum value among the weighted sum of the average weight of all counter-evidence edges associated with the current E-stage claim and the uncertainty label score. When the uncertainty label does not exist, the uncertainty label score is 0, and when the uncertainty label exists, the uncertainty label score is 0.20~0.

80.

10. A cultural image recognition system based on prior knowledge space and ontology gating, characterized in that: It includes a memory and a processor, the memory being used to store a computer program, which, when executed by the processor, implements the cultural image recognition method based on prior knowledge space and ontology gating as described in any one of claims 1 to 9.