A multimodal knowledge graph automatic construction and interactive narration method and system
By aligning and classifying multimodal data, a global cultural relic knowledge graph is constructed, and topological constraint retrieval is enhanced. This solves the problems of cross-modal fusion and limited interaction in multimodal knowledge graphs, and achieves an efficient, reliable knowledge dissemination and a user-friendly system solution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI
- Filing Date
- 2026-05-20
- Publication Date
- 2026-06-16
AI Technical Summary
Existing technologies for constructing multimodal knowledge graphs suffer from difficulties in cross-modal data fusion, redundancy in static graph structures, and limited interaction methods, resulting in insufficient graph accuracy and poor user experience.
By acquiring multimodal data, performing cross-modal alignment and classification, extracting knowledge and resolving conflicts, constructing a global cultural relic knowledge graph, and enhancing retrieval based on graph topological constraints, the visualization and narrative interaction of the knowledge graph are realized.
It improves the accuracy and completeness of knowledge graphs, supports complex queries with multi-hop reasoning, provides an immersive and story-based knowledge exploration experience, and solves the problems of integration and monotonous interaction forms in existing technologies.
Smart Images

Figure CN122221982A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal knowledge graph construction, and more specifically, to a method and system for automatic construction and interactive narrative of multimodal knowledge graphs. Background Technology
[0002] With the growing demand for the digitization of cultural heritage, the need to integrate multimodal data of cultural relics, including text and images, is becoming increasingly urgent. However, existing technologies face core bottlenecks in the construction and application of multimodal knowledge graphs: at the knowledge integration level, they rely heavily on a single text modality, making it difficult to effectively integrate the semantics of non-textual information such as images, and lack the ability to resolve descriptive conflicts between cross-modal data, resulting in insufficient graph accuracy; at the knowledge organization level, static graph structures are redundant, making it difficult to clearly distinguish attributes and relationships, while retrieval techniques based on simple text slicing easily lose deep logical connections between knowledge, restricting complex queries and reasoning; at the interaction level, the display format is monotonous, lacking dynamic narrative generation and multimodal intelligent interaction capabilities based on user intent and graph logic, limiting knowledge dissemination efficiency and user experience. Summary of the Invention
[0003] The purpose of this invention is to provide a method and system for automatic construction and interactive narrative of multimodal knowledge graphs, so as to improve the above-mentioned problems.
[0004] To achieve the above objectives, the embodiments of this application provide the following technical solutions:
[0005] On the one hand, embodiments of this application provide a method for automatic construction and interactive narrative of multimodal knowledge graphs, the method comprising:
[0006] Acquire multimodal data, which includes textual data of the cultural relics and image data from different angles;
[0007] Cross-modal alignment and classification are performed on the multimodal data to obtain a standardized and associated multi-source dataset;
[0008] Knowledge extraction and conflict resolution are performed on the standardized multi-source dataset to obtain triple information;
[0009] Based on the triplet information, data structuring and multi-source knowledge fusion processing are performed to obtain a global cultural relic knowledge graph;
[0010] Based on the global cultural relic knowledge graph, a retrieval enhancement process based on graph topology constraints is performed to obtain the retrieval system;
[0011] Based on the global cultural relic knowledge graph and the retrieval system, the knowledge graph is visualized and narrated interactively to obtain the narrative storyline.
[0012] Secondly, embodiments of this application provide a multimodal knowledge graph automatic construction and interactive narrative system, the system comprising:
[0013] The acquisition module is used to acquire multimodal data, which includes text data of cultural relics and image data from different angles;
[0014] The first processing module is used to perform cross-modal alignment and classification on the multimodal data to obtain a standardized and associated multi-source dataset;
[0015] The second processing module is used to perform knowledge extraction and conflict resolution on the standardized associated multi-source dataset to obtain triple information;
[0016] The third processing module is used to perform data structuring and multi-source knowledge fusion processing based on the triplet information to obtain a global cultural relic knowledge graph.
[0017] The fourth processing module is used to perform retrieval enhancement processing based on graph topology constraints according to the global cultural relic knowledge graph to obtain the retrieval system;
[0018] The fifth processing module is used to perform visualization and narrative interaction processing of the knowledge graph based on the global cultural relic knowledge graph and the retrieval system to obtain the narrative storyline.
[0019] Thirdly, embodiments of this application provide an automatic construction and interactive narrative device for multimodal knowledge graphs, the device including a memory and a processor. The memory is used to store a computer program; the processor is used to execute the computer program to implement the steps of the above-described automatic construction and interactive narrative method for multimodal knowledge graphs.
[0020] Fourthly, embodiments of this application provide a readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described method for automatic construction and interactive narrative of multimodal knowledge graphs.
[0021] The beneficial effects of this invention are as follows:
[0022] This invention achieves cross-modal association between text and images through automated alignment and classification, solving the problems of relying on a single modality and difficulty in integrating multi-source semantics, thus improving the accuracy and completeness of knowledge graphs. By extracting labeled triples and performing structured mapping, it achieves precise separation of attributes and associations, overcoming the shortcomings of traditional graph structures such as redundancy and difficulty in discerning core logic. Furthermore, it enhances retrieval by extracting local context based on graph topological hop count constraints, ensuring semantic coherence and overcoming the fragmentation problem caused by simple slicing, supporting complex queries with multi-hop reasoning. Finally, based on the global graph and retrieval results, combined with visual interaction and subgraph narrative generation, it provides an immersive, story-based knowledge exploration experience, solving the problems of single interaction forms and low dissemination efficiency in existing technologies. This invention achieves full-process automation from multimodal fusion and knowledge construction to intelligent retrieval and narrative interaction, providing an efficient, reliable, and user-friendly systematic solution for the digital protection and intelligent dissemination of cultural relics.
[0023] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings. Attached Figure Description
[0024] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a schematic diagram of the process of automatic construction and interactive narrative of multimodal knowledge graphs as described in an embodiment of the present invention.
[0026] Figure 2 This is a schematic diagram of the structure of the multimodal knowledge graph automatic construction and interactive narrative device described in this embodiment of the invention.
[0027] The diagram is labeled as follows: 800, Automatic construction and interactive narrative device for multimodal knowledge graphs; 801, Processor; 802, Memory; 803, Multimedia component; 804, I / O interface; 805, Communication component. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0029] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0030] Example 1:
[0031] This embodiment provides a method for automatically constructing and interactively narrating multimodal knowledge graphs. It can be understood that this embodiment can present a scenario, such as: during the digitization process of museums, high-resolution images of their artifacts and scattered, multi-source, unstructured textual materials form data silos that are difficult to automatically connect, and the information contains inconsistent or even contradictory terminology. This makes it difficult for researchers to efficiently construct clear and reliable knowledge systems for in-depth research, and curators cannot easily plan vivid and coherent narrative exhibitions that reveal the deep connections between artifacts, forming the core problems of difficult knowledge mining and low dissemination efficiency.
[0032] See Figure 1 The figure shows that the method includes steps S1-S6.
[0033] Step S1: Acquire multimodal data, which includes text data of the cultural relic and image data from different angles;
[0034] In this step, we first acquire multi-source heterogeneous data related to the guqin using a scanner and web crawler. Text data includes professional literature and official museum descriptions related to the artifact; OCR character recognition technology is used to convert the scanned documents into editable text. Image data includes high-resolution photographs of the artifact from different angles; for example, for the guqin, this includes the soundboard, back, sound hole, inscriptions, and detailed ornamentation.
[0035] Step S2: Perform cross-modal alignment and classification on the multimodal data to obtain a standardized and associated multi-source dataset;
[0036] Step S2 further includes steps S21-S24, which specifically include:
[0037] Step S21: Process the multimodal data using the CLIP dual-tower architecture to obtain image feature information and text feature information;
[0038] In the CLIP dual-tower architecture used in this step, images are processed by a visual encoder, while pre-defined semantic libraries of guqin components, such as the soundboard bottom and inscriptions, are processed by a text encoder. Both are mapped to a unified high-dimensional feature space. To ensure the comparability of features in the space, the system performs norm normalization on the original feature vectors output by the encoders for subsequent similarity calculations. It should be noted that this application, while keeping the visual encoder parameters frozen, incrementally fine-tunes the text encoder using a professional guqin dataset. The fine-tuning process employs a cross-entropy loss function for optimization, with the formula:
[0039]
[0040] In the above formula, This represents the contrastive loss value during the training process; the smaller the value, the more accurate the model's recognition of the guqin's parts. This represents the total number of samples in the current training batch; This represents the total number of preset categories of guqin components; Indicates the first Similarity score between an image and its corresponding correct text description (positive sample pair); Indicates the first The image and the first The similarity score (negative sample pair) between unrelated text descriptions is used. This formula enhances the model's discriminative power by bringing positive samples closer together and pushing negative samples further apart.
[0041] Step S22: Calculate the similarity matrix based on the image feature information and the text feature information;
[0042] In this step, the similarity matrix is obtained by calculating the dot product of image feature information and text feature information. The specific process is as follows:
[0043]
[0044] In the above formula, A score is given for matching images and text. and These are the normalized feature vectors for the image and the text, respectively; This is a temperature hyperparameter used to adjust the smoothness of the score distribution and prevent the prediction results from being too extreme.
[0045] Step S23: Use the Softmax operator to transform the similarity matrix to generate the posterior probability corresponding to each category, and obtain the predicted probability distribution;
[0046] In this step, the specific calculation process for the posterior probability is as follows:
[0047]
[0048] In the above formula, The posterior probability represents the likelihood that an image belongs to a certain category of a guqin component. This is the set of scores for the image and all candidate categories.
[0049] Step S24: Label the multimodal data according to the predicted probability distribution to obtain the standardized associated multi-source dataset.
[0050] In this step, the label corresponding to the maximum probability value is selected as the classification attribute, and an automated file renaming logic is triggered accordingly to modify the original image name to the corresponding semantic label. To prevent naming conflicts between images of similar parts (such as multiple close-ups of Longchi), the system introduces a counter-based conflict detection mechanism: if a file with the same name is detected, an incrementing sequence number (such as longchi_1.jpg, longchi_2.jpg) is automatically appended to the file name, thereby completing the standardized association annotation of the guqin multimodal data.
[0051] Step S3: Perform knowledge extraction and conflict resolution on the standardized multi-source dataset to obtain triple information;
[0052] Step S3 further includes steps S31-S34, which specifically include:
[0053] Step S31: Perform semantic enhancement and rewriting processing on the text data obtained from the standardized association multi-source dataset to obtain semantically enhanced text data;
[0054] Traditional information extraction techniques typically process the raw text directly. However, a common phenomenon in guqin (a seven-stringed zither) literature is the "zero-subject" phenomenon, where the core instrument name is often omitted when describing the instrument's surface, dimensions, or inscriptions. Direct extraction using existing techniques leads to missing subjects in triples, resulting in numerous misaligned and isolated attribute nodes in the graph, highlighting the problem of subjectless triples. Therefore, this step involves semantic enhancement rewriting, using a large language model to pre-identify core artifact entities in the text, and reconstructing fragmented descriptions lacking subjects into complete sentences centered on these entities. This process follows a mapping function. in, This represents the raw, unstructured text obtained through OCR recognition. The unique identifier for the identified core cultural relic entity (such as "the '11_Xiang' Qin"); The enhanced text with explicit logical associations after rewriting, ensuring that all attributes extracted subsequently can accurately point to the cultural relic entity. By performing a unique retrieval of the cultural relics, it solves the problem of entity reference conflicts caused by multiple names or unclear references for the same qin in qin literature.
[0055] Step S32: Perform predicate semantic clustering processing on the semantically enhanced text data to obtain a standardized predicate system;
[0056] Since the descriptive words for the same attribute from different literature sources are often different (e.g., "length", "total length", "overall length" all refer to the same physical attribute), in this step, a semantic similarity algorithm is used to execute an aggregation function where represents the set of heterogeneous relationship predicates initially extracted from the full text (such as "1. The total length is", "The overall length is approximately"); represents the aggregated standard predicate (such as uniformly mapped to "The length is"). Automatically induct, merge, and standardize these large amounts of heterogeneous predicates initially extracted from the text, and finally form a unified and standardized predicate system. This step eliminates semantic redundancy, constructs a standardized qin knowledge system, and provides a unified corpus basis for subsequent complex reasoning.
[0057] Step S33: Extract labeled triples according to the semantically enhanced text data and the standardized predicate system to obtain initial triples;
[0058] The traditional triple structure is single and cannot distinguish the relationships between entities (such as the production process of Qin A is similar to that of Qin B) and the inherent attributes of entities (such as the material of Qin A is paulownia). In visualization and RAG retrieval, if specific values are regarded as independent nodes, it will cause tens of thousands of meaningless isolated leaf nodes in the graph. In addition, homonymous entities (such as different "Xiang" qins are recorded in different literatures) will lead to incorrect node merging. Therefore, in this step, a labeled triple extraction algorithm is executed, and the structure is used for storage, where represents the main entity node with a unique file identifier prefix (such as "8_Xiang", where 8 is the document number), effectively solving the problem of name conflicts; is the predicate after standardized processing; is the object (including associated entities or specific descriptive values); The classification labels are used to distinguish whether the knowledge belongs to the relationship between entities or describes the attribute of an entity. This step achieves precise separation of graph node and edge types: attribute information is stored as the metadata of the node, while relationship information is constructed as the edge of the graph, thus ensuring the conciseness of the graph structure.
[0059] Step S34: Perform knowledge conflict resolution and structure verification on the initial triples to obtain the triple information.
[0060] Because automated extraction can produce formatting errors or "knowledge illusions" that logically contradict domain common sense (e.g., mistakenly extracting material as length), existing technologies typically lack feedback correction mechanisms, resulting in a significant amount of noise and logical errors in the knowledge entering the database. This step introduces automated verification and conflict resolution logic. The process uses a verification function... Implementation, in which The triplet to be verified; This is a predefined set of constraint rules for the guqin domain (including predicate whitelists, numerical validity, two-hop path logic rules, etc.). Processing logic: If the rules are satisfied (denoted as...)... () was determined to be valid data If not satisfied (denoted as...) If ), then the correction function will be activated. The correction function utilizes structured text. Abnormal triples are subjected to secondary comparison and logical reconstruction to resolve conflicts arising from model inference. For example, if "length is 'Paulownia wood'" is extracted, the correction logic will identify the predicate error through context and correct it to "material is", ensuring the rigor of the graph data.
[0061] Step S4: Perform data structuring and multi-source knowledge fusion processing based on the triplet information to obtain a global cultural relic knowledge graph;
[0062] Step S4 further includes steps S41-S44, which specifically include:
[0063] Step S41: Perform knowledge element structured mapping processing based on the triplet information to obtain a structured knowledge representation object. The structured knowledge representation object includes attribute nodes and associated edges.
[0064] In existing technologies, all triples are treated as equivalent and processed through direct mapping. However, this method cannot handle the complex data of artifacts like the guqin, which involves both intricate physical attributes (such as size and material) and social relationships (such as provenance and similar instruments). Failure to distinguish between attributes and relationships would result in a large number of invalid entity nodes representing specific numerical values in the graph. Therefore, in this step, a knowledge element structured mapping logic is executed, using classification labels to guide data flow through a mapping function. Implementation, in which Represents the set of labeled triples of the input; This is the generated structured JSON object. The mapping logic is based on... The value of the branch is processed: if The data is categorized as relationships between entities. ;like The data is then mapped to the attribute nodes of the entity. This improvement, through tag-guided branch mapping, decouples attributes and relationships before data is stored, avoiding graph redundancy and ensuring that unstructured text can be transformed into semantically clear and structurally rigorous JSON knowledge representation.
[0065] Step S42: Extract the main identifier of the cultural relic from the text data in the multimodal data to obtain the standard name identifier of the cultural relic;
[0066] The names for the guqin in historical documents often vary (e.g., sometimes called "Xiangqin," sometimes "this guqin"), making it difficult for current technology to accurately identify the official, standard names of the artifact from the messy, long texts recognized by OCR, resulting in blurred central nodes in the image. Therefore, this step executes an automatic ontology name recognition algorithm to extract core identifiers from the original OCR text. This process uses a regular expression matching function. Implementation, in which This is the original OCR text data; For pre-defined semantic pattern operators for the guqin (such as specific matching rules for "qin name" or "this qin has no name"); This step extracts the standard name of the guqin (e.g., "Xiang" guqin). Through semantic pattern operators, this extraction method filters out redundant information from highly colloquial or non-standard descriptions, accurately capturing the unique identifier of the artifact and laying the foundation for global knowledge alignment.
[0067] Step S43: Perform entity attribute enhancement annotation processing based on the structured knowledge representation object and the standard name identifier of the cultural relic to obtain a local knowledge table for each cultural relic. The local knowledge table of the cultural relic includes a set of entity nodes with the main marker of the cultural relic.
[0068] In this step, all generated entity nodes (corresponding to each object in the knowledge graph, such as a musical instrument, a person, or a location) are compared with the extracted standard names of cultural relics. An indicator function is used to assign attribute values to each node, defined by the following formula: in Indicates the current entity node Assigning values to cultural relics attributes; The name string of the entity; This is the standard subject identifier extracted in the previous step. When... If the name matches exactly, a value of 1 is assigned, indicating that the node is the main cultural relic; otherwise, a value of 0 is assigned. This step automates the assignment through indicator functions, giving each entity record a topic tag, enabling automated anchoring of the core guqin target even in large-scale node sets.
[0069] Step S44: Perform fusion processing based on the local knowledge table of each cultural relic to obtain the global cultural relic knowledge graph.
[0070] When processing multiple sets of guqin data from different museums and books, existing techniques often merge them by simply appending text. This leads to a large number of duplicate entities, such as multiple documents mentioning "Paulownia wood," making the datasheets extremely bloated and logically conflicting. Therefore, this step utilizes multi-source data fusion and deduplication logic to merge local tables into a global database. The merging follows the union deduplication operator. in This is the merged global cultural relics knowledge graph dataset; This represents the total number of guqin samples collected. In order to target the A set of local entities and relation tables generated by a guqin; This represents the deduplication operator. During the merging process, the entity tables are checked for overlap based on the name consistency principle, and deduplication is performed on the relation tables based on the "triple uniqueness" principle (i.e., the head entity, predicate, and tail entity must be completely identical). This ultimately constructs a highly consistent global cultural relic knowledge graph that eliminates data redundancy. This step effectively resolves logical redundancy between data from different sources, constructing a highly rigorous, conflict-free, large-scale guqin knowledge network, significantly improving the graph's query efficiency and accuracy.
[0071] Step S5: Perform retrieval enhancement processing based on graph topology constraints according to the global cultural relic knowledge graph to obtain the retrieval system;
[0072] Step S5 further includes steps S51-S53, which specifically include:
[0073] Step S51: Extract local context based on the graph topology of the global cultural relic knowledge graph to obtain semantic background;
[0074] Traditional RAG systems usually slice the original long text directly, resulting in semantic fragmentation and losing the logical relationships between the attributes of cultural relics. In knowledge graph retrieval, simple keyword queries can only obtain isolated nodes with direct associations and cannot capture the complex implicit connections between cultural relics (for example, although two guqins have no direct contact, they are associated through common manufacturing techniques or dynasty backgrounds). In this step, through the knowledge graph topology feature extraction logic, the semantic background is reconstructed using node marking and hop count constraints. First, a set of nodes that meet the condition is screened out from the global cultural relic knowledge graph (where represents any node in the graph, is the cultural relic attribute marker, is the set of cultural relic main body nodes). For each core node , its local knowledge context document is extracted through the breadth-first traversal algorithm, and this process follows the following hop count constraint criteria: where is the enhanced semantic description document of node ; represents the shortest path hop count between node and node in the graph; is the set of attribute triples with a path distance not exceeding 2 hops, defined as the inherent physical attributes of this cultural relic; is the association path with a path distance of 3 or 4 hops and the end node belonging to the cultural relic main body set , and this path represents the similarity mapping of the current cultural relic with other cultural relics in terms of craftsmanship style or historical background. In this step, the hierarchical extraction of physical attributes and logical associations is achieved through hop count constraints. The attributes within 2 hops ensure the accuracy of question answering, while the cross-body associations of 3 to 4 hops挖掘 the potential similarities between guqins, enabling the system to answer deep questions such as "Which guqins in the museum have similar craftsmanship to this guqin?".
[0075] Step S52: Perform vector representation processing on heterogeneous knowledge blocks according to the semantic background to obtain feature vectors;
[0076] Existing technologies often directly vectorize unstructured text. In the field of guqins, the text contains a large number of professional terms (such as "dragon pond", "phoenix pool", "frosting"), and direct vectorization is likely to cause semantic deviation, and the retrieval efficiency decreases significantly as the document volume increases. Therefore, in this step, vector storage construction and hierarchical indexing logic are executed. The extracted topological paths are converted into standardized natural language descriptions (such as converting the triple "Xiang - material - paulownia wood" into "The material of the 'Xiang' guqin is paulownia wood"), and then mapped to a high-dimensional space using a deep embedding model. The formula is This represents a pre-trained deep text embedding operator (such as the OpenAI embedding model). For the generated enhanced semantic document; This step generates a feature vector for the artifact. It replaces the unstructured text with a standardized description based on graph reconstruction, eliminating noise in the original OCR text and resulting in a cleaner vector representation. The system stores the vectors in vector databases such as FAISS, significantly improving the retrieval speed for massive amounts of artifact data.
[0077] Step S53: Perform retrieval enhancement processing based on the feature vector to obtain the retrieval system.
[0078] In traditional retrieval and question answering systems, complex questions often lead to incomplete or conflicting background information, causing the large model to "illusion" or provide overly broad answers. Therefore, this step executes semantic matching and retrieval enhancement generation logic. When a user query is received... First, extract the query vector. And calculate its matching degree with the vectors in the library. in The cosine similarity score is used to measure the relevance between user questions and knowledge of cultural relics. The feature vector representing the knowledge background of the v-th cultural relic in the global cultural relic knowledge graph; and These are the magnitudes of the corresponding vectors. This step selects the vectors with the highest similarity scores. The knowledge context serves as a retrieval enhancement background. Finally, a large language model is used to generate the answer. in, Indicates the answer result; Representing a large language model; This represents the enhanced background knowledge provided to generate the current answer. This step, through a graph topology-driven RAG mechanism, ensures that the background information input to the large model has undergone pre-screening by graph logic. Even if the user's question is vague, the system can retrieve relevant cultural relic background information through pre-defined topological relationships in the graph, enabling the generation of highly professional and logically rigorous answers to cultural relic questions without human intervention.
[0079] Step S6: Perform visualization and narrative interaction processing of the knowledge graph based on the global cultural relic knowledge graph and the retrieval system to obtain the narrative storyline.
[0080] Step S6 further includes steps S61-S66, which specifically include:
[0081] Step S61: Construct a visual interactive system based on the global cultural relic knowledge graph and the retrieval system to obtain the interactive system;
[0082] In this step, the existing global knowledge graph and retrieval system are utilized to provide users with an immersive interactive loop, from active exploration to active acquisition, and from data association to story generation. This step constructs a dynamic, force-oriented layout-based visual interface using front-end technology. This interface allows users to perform intuitive operations such as dragging, scaling, clicking, and selecting primitives (nodes and edges) representing entities like artifacts, figures, and locations. It is supported by a global knowledge graph data model and retrieval system, enabling each interactive operation to drive real-time querying and calculation of underlying data. This step provides an intuitive platform and entry point for all subsequent interactive functions. It should be noted that this application's interactive system introduces visual feature encoding logic, directly mapping the unstructured features of artifacts onto visual symbols by deploying barcodes or color wrapping mechanisms around nodes. Each type of artifact attribute corresponds to a visual encoding function. in For the entity's first One attribute, For a specific color value to be assigned, This corresponds to a texture or shape pattern. Through this visual encoding method, users can quickly determine the age range, technological sophistication, and other core dimensions of an artifact simply by observing the color distribution and pattern features of the map, without needing to consult a detailed list, thus improving the efficiency of information retrieval under large-scale maps.
[0083] Step S62: Obtain the user's operation instructions in the interactive system, and determine the target knowledge subgraph based on the user's operation instructions;
[0084] In traditional knowledge graph visualization, when researchers need to compare the commonalities of multiple guqin (such as the "Xiang" guqin and the "Zhonghe" guqin), they often have to manually search for common nodes in a complex network structure, which is extremely inefficient and prone to missing key connections. However, this application constructs a visualization interface based on force-directed layout, allowing users to activate specific entity node sets through single or multiple selection operations. It will also calculate and highlight the set of public adjacency nodes of these entities. in This represents the calculated set of common nodes (such as the "Paulownia wood" material node or the "Tang Dynasty" era node shared by two zithers). For multiple guqin entity nodes selected by the user on the interface; Represents a node The set of one-hop adjacent nodes in the graph (i.e., all attributes directly connected to the instrument). Furthermore, this application supports logical deletion of redundant nodes and manual addition of new nodes to complete the knowledge chain. For complex artifact attributes, the system supports node splitting, i.e., when a node contains multiple composite semantics, it is decomposed into multiple atomic nodes using a splitting operator. For a specific subgraph selected by the user... The system maps this data to a temporary storage area at the bottom of the interface, using it as context material for subsequent storyline generation. This process is achieved through a context mapping function. Implementation, in which This allows users to lock a local topology using box selection or multi-selection. This describes the generated structured storyline.
[0085] Step S63: Perform semantic structure extraction processing on the target knowledge subgraph to obtain a structured semantic context;
[0086] In existing technologies, narrative generation typically relies on user-input keywords or the entire text, lacking an awareness of the local logic within the graph. When users focus on specific relationships within a complex guqin (a seven-stringed zither) graph (such as the relationship between a guqin's unique patterns and its manufacturing process), existing technologies cannot accurately extract this local topological structure, resulting in generic narratives that fail to focus on the subset of knowledge that interests the user. Therefore, in this step, a visual selection capture logic is used, allowing users to select specific guqin nodes and their associated attributes within the interface, and then execute a subgraph semantic compression operator. in Represents the local topology structure that the user has locked by selecting or multi-selecting; This represents a serialization function, responsible for converting the topological path between nodes (such as "'Xiang' Qin—Material—Paulownia Wood") into a natural language logical chain that can be understood by the large language model; This is the final structured context corpus.
[0087] Step S64: Perform deep knowledge background retrieval processing based on the structured semantic context to obtain retrieval results for enhancing the narrative;
[0088] To enhance the depth, detail, and professionalism of the final generated story, this step treats the structured semantic context as a new and more complex query. The retrieval system then uses this context as a guide to extract more relevant, logically connected, and in-depth knowledge background from the global graph, such as the craft features related to Dynasty C. This forms an expanded retrieval result to enhance the narrative, ensuring that the generated story possesses rich details that go beyond superficial connections.
[0089] Step S65: Perform narrative storyline generation processing based on the retrieval results used to enhance the narrative to obtain a preliminary narrative storyline;
[0090] Existing generative AI often produces content with a single style, unable to automatically switch narrative perspectives based on different dimensions of cultural relics (such as craftsmanship vs. historical transmission). Relying solely on users manually inputting complex commands is too difficult for non-professional users. Therefore, this step implements a storyline-driven prompt customization logic, dynamically adjusting the generation strategy based on the narrative type selected by the user in the interaction panel (such as "cultural relic craftsmanship storyline" or "transmission and evolution storyline"). For different types, the system introduces a type weighting function. To intervene in the generation of guidance: in The final instruction set (Prompt) for inputting the large language model; This is the final structured context corpus; The generated deep knowledge supplements are used to enhance retrieval and enrich the details of the story; For users' personalized generation intent submitted through the input box (such as requesting a story style that is "science popularization" or "literary"); To ensure that the generated storylines conform to the preset knowledge logic, the constraints parameters are designed for specific narrative dimensions such as "craftsmanship," "history," and "heritage."
[0091] Meanwhile, a loose narrative logic exists when generating the initial storyline. This step also introduces a linear serialization algorithm based on cross-entity topological associations, which is implemented by executing a narrative logic weaving function. Preprocess the retrieved subgraph paths. The specific formula is as follows: ,in This is the final structured context corpus; It weaves functions for narrative logic, responsible for mapping non-linear mesh topological subgraphs into linear logic flows; Provides a local topology that users can lock by selecting a box or multiple selections; This is the set of nodes anchored to the core logic, representing the topological intersection points between different entities; These are constraint parameters for the narrative dimension, used to adjust the weight distribution of different logical paths (such as craft paths and transmission paths) during generation. In this step, for a specific cultural relic entity, the algorithm traverses the knowledge graph... By skipping attribute paths, subordinate attribute nodes, including material mechanisms, process parameters, and inscription features, are extracted and transformed from discrete topological points into narrative primitives with causal logical chains, thus realizing in-depth mining of high-order attribute paths. At the same time, through the topological structure of the global knowledge graph, common logical subgraphs or implicit related paths between different cultural relics entities are automatically identified and defined as core logical anchor nodes, thereby realizing cross-temporal and spatial cultural relic association modeling.
[0092] Step S66: Perform knowledge consistency verification and correction on the preliminary narrative storyline to obtain the narrative storyline.
[0093] Large language models are prone to generating "factual illusions" when generating long narratives, especially in highly specialized fields like the guqin (a traditional Chinese stringed instrument), where issues such as time travel and misuse of craft terminology frequently occur. During the generation phase, a consistency score is calculated to ensure content authenticity, specifically:
[0094]
[0095] In the above formula, This represents the initial narrative text generated by the model; It is the deep historical background knowledge related to the current subgraph retrieved through the RAG mechanism; Represents the semantic cosine similarity calculation function; A preliminary semantic consistency score is then given. If... If the value falls below a preset threshold, the system will automatically activate a self-correction mechanism. Certain facts in The process involves reconstruction, ultimately outputting a logically rigorous narrative storyline about the cultural relics. It should be noted that the generated narrative storyline will be dynamically rendered on the front end and displayed in a card-based format in the temporary storage area at the bottom of the visualization interface. Each story card retains its corresponding partial image. The mapping relationship allows users to perform secondary editing, deletion, or node splitting operations. This process is achieved through mapping storage functions. Implementation, in which For displaying story card objects, This serves as a unique identifier for the subgraph that triggers the generation of the story. This design achieves a closed-loop process across the entire chain of "graph exploration - semantic extraction - story generation - interactive feedback," greatly enhancing the efficiency of digital display of cultural relic knowledge.
[0096] Furthermore, to address potential logical paradoxes in long-range narratives, this application introduces a graph topological consistency check operator based on the aforementioned semantic cosine similarity. By constructing a comprehensive consistency scoring function Perform hard constraint validation on narrative logic:
[0097]
[0098] in, A comprehensive consistency score that integrates semantic measures and topological constraints; As a weighting factor, used to adjust the proportion of semantic similarity and logical consistency, its value should be dynamically weighted based on the knowledge density of the narrative dimension: in high-precision logical scenarios such as the craft of the guqin, it is recommended to take [value missing]. To strengthen the hard constraints of the graph topology; while in emotional narrative scenarios such as cultural background, it is recommended to take... To enhance the richness of semantic expression. For preliminary semantic consistency scoring; This is a graph topology consistency check operator used to quantify the degree of fit between textual logic and graph facts; For the logical predicate chain extracted from the generated text; This represents a local knowledge graph topological subgraph that serves as a logical reference benchmark. In this step, the operator ensures the authenticity of the generated content through hard logical constraints:
[0099] Topological path mapping alignment: Operators extract preliminary narrative text The logical reasoning chain in the graph is compared with the original knowledge graph. The multi-hop fact path in the process is used for isomorphic mapping verification.
[0100] Logical validity hard constraint: Determines whether causal inferences in the generated text fall within the valid topological space defined by the knowledge graph. If the narrative logic violates the preset attribute constraints or spatiotemporal partial order relationships in the graph (such as temporal contradictions in technological evolution), this operator will determine that the logic is invalid and trigger a reconstruction process based on deterministic facts.
[0101] Example 2:
[0102] This embodiment provides a multimodal knowledge graph automatic construction and interactive narrative system. The system includes an acquisition module, a first processing module, a second processing module, a third processing module, a fourth processing module, and a fifth processing module, specifically including:
[0103] The acquisition module is used to acquire multimodal data, which includes text data of cultural relics and image data from different angles;
[0104] The first processing module is used to perform cross-modal alignment and classification on the multimodal data to obtain a standardized and associated multi-source dataset;
[0105] The second processing module is used to perform knowledge extraction and conflict resolution on the standardized associated multi-source dataset to obtain triple information;
[0106] The third processing module is used to perform data structuring and multi-source knowledge fusion processing based on the triplet information to obtain a global cultural relic knowledge graph.
[0107] The fourth processing module is used to perform retrieval enhancement processing based on graph topology constraints according to the global cultural relic knowledge graph to obtain the retrieval system;
[0108] The fifth processing module is used to perform visualization and narrative interaction processing of the knowledge graph based on the global cultural relic knowledge graph and the retrieval system to obtain the narrative storyline.
[0109] In one specific embodiment of this disclosure, the first processing module includes a first processing unit, a second processing unit, a third processing unit, and a fourth processing unit, specifically including:
[0110] The first processing unit is used to process the multimodal data using the CLIP dual-tower architecture to obtain image feature information and text feature information;
[0111] The second processing unit is used to calculate a similarity matrix based on the image feature information and the text feature information.
[0112] The third processing unit is used to transform the similarity matrix using the Softmax operator to generate the posterior probability corresponding to each category, thereby obtaining the predicted probability distribution.
[0113] The fourth processing unit is used to label the multimodal data according to the predicted probability distribution to obtain the standardized associated multi-source dataset.
[0114] In one specific embodiment of this disclosure, the second processing module includes a fifth processing unit, a sixth processing unit, a seventh processing unit, and an eighth processing unit, specifically including:
[0115] The fifth processing unit is used to perform semantic enhancement and rewriting processing on the text data obtained from the standardized and associated multi-source dataset to obtain semantically enhanced text data.
[0116] The sixth processing unit is used to perform predicate semantic clustering on the semantically enhanced text data to obtain a standardized predicate system;
[0117] The seventh processing unit is used to extract labeled triples based on the semantically enhanced text data and the standardized predicate system to obtain initial triples;
[0118] The eighth processing unit is used to perform knowledge conflict resolution and structure verification on the initial triples to obtain the triple information.
[0119] In one specific embodiment of this disclosure, the third processing module includes a ninth processing unit, a tenth processing unit, an eleventh processing unit, and a twelfth processing unit, specifically comprising:
[0120] The ninth processing unit is used to perform knowledge element structured mapping processing based on the triplet information to obtain a structured knowledge representation object, which includes attribute nodes and associated edges.
[0121] The tenth processing unit is used to extract the main identifier of cultural relics from the text data in the multimodal data to obtain the standard name identifier of the cultural relics.
[0122] The eleventh processing unit is used to perform entity attribute enhancement annotation processing based on the structured knowledge representation object and the standard name identifier of the cultural relic, to obtain a local knowledge table for each cultural relic, wherein the local knowledge table of the cultural relic includes a set of entity nodes with the main marker of the cultural relic.
[0123] The twelfth processing unit is used to perform fusion processing based on the local knowledge table of each cultural relic to obtain the global cultural relic knowledge graph.
[0124] In one specific embodiment of this disclosure, the fourth processing module includes a thirteenth processing unit, a fourteenth processing unit, and a fifteenth processing unit, specifically comprising:
[0125] The thirteenth processing unit is used to perform local context extraction based on the graph topology structure of the global cultural relics knowledge graph to obtain semantic background.
[0126] The fourteenth processing unit is used to perform vector representation processing of heterogeneous knowledge blocks based on the semantic background to obtain feature vectors;
[0127] The fifteenth processing unit is used to perform retrieval enhancement processing based on the feature vector to obtain the retrieval system.
[0128] In one specific embodiment of this disclosure, the fifth processing module includes a sixteenth processing unit, a seventeenth processing unit, an eighteenth processing unit, a nineteenth processing unit, a twentieth processing unit, and a twenty-first processing unit, specifically comprising:
[0129] The sixteenth processing unit is used to construct a visual interactive system based on the global cultural relic knowledge graph and the retrieval system, thereby obtaining the interactive system.
[0130] The seventeenth processing unit is used to obtain the user's operation instructions in the interactive system and determine the target knowledge subgraph based on the user's operation instructions;
[0131] The eighteenth processing unit is used to perform subgraph semantic structure extraction processing based on the target knowledge subgraph to obtain a structured semantic context;
[0132] The nineteenth processing unit is used to perform deep knowledge background retrieval processing based on the structured semantic context to obtain retrieval results for enhancing the narrative;
[0133] The twentieth processing unit is used to perform narrative storyline generation processing based on the retrieval results for enhancing the narrative, and to obtain a preliminary narrative storyline.
[0134] The twenty-first processing unit is used to perform knowledge consistency verification and correction processing on the preliminary narrative storyline to obtain the narrative storyline.
[0135] It should be noted that the specific methods by which each module performs operations in the system described in the above embodiments have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0136] Example 3:
[0137] Corresponding to the above method embodiments, this embodiment also provides a device for automatic construction and interactive narrative of multimodal knowledge graphs. The device for automatic construction and interactive narrative of multimodal knowledge graphs described below and the method for automatic construction and interactive narrative of multimodal knowledge graphs described above can be referred to in correspondence.
[0138] Figure 2 This is a block diagram illustrating an automatic construction and interactive narrative device 800 for multimodal knowledge graphs, according to an exemplary embodiment. Figure 2 As shown, the multimodal knowledge graph automatic construction and interactive narrative device 800 may include: a processor 801 and a memory 802. The multimodal knowledge graph automatic construction and interactive narrative device 800 may also include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.
[0139] The processor 801 controls the overall operation of the multimodal knowledge graph automatic construction and interactive narrative device 800 to complete all or part of the steps in the aforementioned multimodal knowledge graph automatic construction and interactive narrative method. The memory 802 stores various types of data to support the operation of the multimodal knowledge graph automatic construction and interactive narrative device 800. This data may include, for example, instructions for any application or method operating on the multimodal knowledge graph automatic construction and interactive narrative device 800, as well as application-related data, such as contact data, sent and received messages, images, audio, video, etc. The memory 802 can be implemented using any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. I / O interface 804 provides an interface between processor 801 and other interface modules, such as keyboards, mice, and buttons. These buttons can be virtual or physical. Communication component 805 is used for wired or wireless communication between the multimodal knowledge graph automatic construction and interactive narrative device 800 and other devices. Wireless communication includes Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof. Therefore, the corresponding communication component 805 may include a Wi-Fi module, a Bluetooth module, and an NFC module.
[0140] In an exemplary embodiment, the multimodal knowledge graph automatic construction and interactive narrative device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the above-described multimodal knowledge graph automatic construction and interactive narrative method.
[0141] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided. When executed by a processor, these program instructions implement the steps of the above-described method for automatic construction and interactive narration of multimodal knowledge graphs. For example, the computer-readable storage medium may be the memory 802 including the program instructions, which may be executed by the processor 801 of the device 800 for automatic construction and interactive narration of multimodal knowledge graphs to complete the above-described method for automatic construction and interactive narration of multimodal knowledge graphs.
[0142] Example 4:
[0143] Corresponding to the above method embodiments, this embodiment also provides a readable storage medium. The readable storage medium described below and the multimodal knowledge graph automatic construction and interactive narrative method described above can be referred to in relation to each other.
[0144] A readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the multimodal knowledge graph automatic construction and interactive narrative method described in the above method embodiments.
[0145] Specifically, the readable storage medium can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or any other readable storage medium capable of storing program code.
[0146] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0147] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for automatic construction and interactive narrative of multimodal knowledge graphs, characterized in that, include: Acquire multimodal data, which includes textual data of the cultural relics and image data from different angles; Cross-modal alignment and classification are performed on the multimodal data to obtain a standardized and associated multi-source dataset; Knowledge extraction and conflict resolution are performed on the standardized multi-source dataset to obtain triple information; Based on the triplet information, data structuring and multi-source knowledge fusion processing are performed to obtain a global cultural relic knowledge graph; Based on the global cultural relic knowledge graph, a retrieval enhancement process based on graph topology constraints is performed to obtain the retrieval system; Based on the global cultural relic knowledge graph and the retrieval system, the knowledge graph is visualized and narrated interactively to obtain the narrative storyline.
2. The method for automatic construction and interactive narrative of multimodal knowledge graphs according to claim 1, characterized in that, Cross-modal alignment and classification of the multimodal data includes: The multimodal data is processed using the CLIP dual-tower architecture to obtain image feature information and text feature information; A similarity matrix is obtained by calculating based on the image feature information and the text feature information; The similarity matrix is transformed using the Softmax operator to generate the posterior probabilities corresponding to each category, thus obtaining the predicted probability distribution; The multimodal data is labeled according to the predicted probability distribution to obtain the standardized associated multi-source dataset.
3. The method for automatic construction and interactive narrative of multimodal knowledge graphs according to claim 1, characterized in that, Knowledge extraction and conflict resolution are performed on the standardized, multi-source datasets, including: The text data obtained from the standardized multi-source dataset is semantically enhanced and rewritten to obtain semantically enhanced text data. The semantically enhanced text data is subjected to predicate semantic clustering to obtain a standardized predicate system; Based on the semantically enhanced text data and the standardized predicate system, labeled triples are extracted to obtain initial triples; The initial triples are subjected to knowledge conflict resolution and structure verification to obtain the triple information.
4. The method for automatic construction and interactive narrative of multimodal knowledge graphs according to claim 1, characterized in that, Based on the triplet information, data structuring and multi-source knowledge fusion processing are performed, including: Based on the triplet information, knowledge element structure mapping is performed to obtain a structured knowledge representation object, which includes attribute nodes and associated edges. The main identifier of cultural relics is extracted from the text data in the multimodal data to obtain the standard name identifier of the cultural relics; Based on the structured knowledge representation object and the standard name identifier of the cultural relic, entity attribute enhancement annotation processing is performed to obtain a local knowledge table for each cultural relic. The local knowledge table of the cultural relic includes a set of entity nodes with the main marker of the cultural relic. The global cultural relic knowledge graph is obtained by fusing the local knowledge tables of each cultural relic.
5. The method for automatic construction and interactive narrative of multimodal knowledge graphs according to claim 1, characterized in that, The retrieval enhancement process based on graph topology constraints is performed according to the global cultural relic knowledge graph, including: Based on the global cultural relic knowledge graph, local context extraction based on the graph topology is performed to obtain the semantic background; Based on the semantic background, the heterogeneous knowledge blocks are processed into vector representations to obtain feature vectors; The retrieval system is obtained by performing retrieval enhancement processing based on the feature vector.
6. A multimodal knowledge graph automatic construction and interactive narrative system, characterized in that, include: The acquisition module is used to acquire multimodal data, which includes text data of cultural relics and image data from different angles; The first processing module is used to perform cross-modal alignment and classification on the multimodal data to obtain a standardized and associated multi-source dataset; The second processing module is used to perform knowledge extraction and conflict resolution on the standardized associated multi-source dataset to obtain triple information; The third processing module is used to perform data structuring and multi-source knowledge fusion processing based on the triplet information to obtain a global cultural relic knowledge graph. The fourth processing module is used to perform retrieval enhancement processing based on graph topology constraints according to the global cultural relic knowledge graph to obtain the retrieval system; The fifth processing module is used to perform visualization and narrative interaction processing of the knowledge graph based on the global cultural relic knowledge graph and the retrieval system to obtain the narrative storyline.
7. The multimodal knowledge graph automatic construction and interactive narrative system according to claim 6, characterized in that, The first processing module includes: The first processing unit is used to process the multimodal data using the CLIP dual-tower architecture to obtain image feature information and text feature information; The second processing unit is used to calculate a similarity matrix based on the image feature information and the text feature information. The third processing unit is used to transform the similarity matrix using the Softmax operator to generate the posterior probability corresponding to each category, thereby obtaining the predicted probability distribution. The fourth processing unit is used to label the multimodal data according to the predicted probability distribution to obtain the standardized associated multi-source dataset.
8. The multimodal knowledge graph automatic construction and interactive narrative system according to claim 6, characterized in that, The second processing module includes: The fifth processing unit is used to perform semantic enhancement and rewriting processing on the text data obtained from the standardized and associated multi-source dataset to obtain semantically enhanced text data. The sixth processing unit is used to perform predicate semantic clustering on the semantically enhanced text data to obtain a standardized predicate system; The seventh processing unit is used to extract labeled triples based on the semantically enhanced text data and the standardized predicate system to obtain initial triples; The eighth processing unit is used to perform knowledge conflict resolution and structure verification on the initial triples to obtain the triple information.
9. The multimodal knowledge graph automatic construction and interactive narrative system according to claim 6, characterized in that, The third processing module includes: The ninth processing unit is used to perform knowledge element structured mapping processing based on the triplet information to obtain a structured knowledge representation object, which includes attribute nodes and associated edges. The tenth processing unit is used to extract the main identifier of cultural relics from the text data in the multimodal data to obtain the standard name identifier of the cultural relics. The eleventh processing unit is used to perform entity attribute enhancement annotation processing based on the structured knowledge representation object and the standard name identifier of the cultural relic, to obtain a local knowledge table for each cultural relic, wherein the local knowledge table of the cultural relic includes a set of entity nodes with the main marker of the cultural relic; The twelfth processing unit is used to perform fusion processing based on the local knowledge table of each cultural relic to obtain the global cultural relic knowledge graph.
10. The multimodal knowledge graph automatic construction and interactive narrative system according to claim 6, characterized in that, The fourth processing module includes: The thirteenth processing unit is used to perform local context extraction based on the graph topology structure of the global cultural relics knowledge graph to obtain semantic background. The fourteenth processing unit is used to perform vector representation processing of heterogeneous knowledge blocks based on the semantic background to obtain feature vectors; The fifteenth processing unit is used to perform retrieval enhancement processing based on the feature vector to obtain the retrieval system.