Museum exhibit deep research system fused with multi-modal large model

By constructing a multimodal large-scale model for in-depth research on museum exhibits, the system solves the problems of automatic cross-modal information association and deep semantic understanding in museum exhibit research, realizes automated research path planning and structured argument construction, and improves research efficiency and quality.

CN121962628APending Publication Date: 2026-05-01CHINA ACAD OF ART +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA ACAD OF ART
Filing Date
2025-12-17
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Current technologies for studying museum exhibits rely on manual cross-referencing, lack the ability to automatically associate cross-modal information and achieve deep semantic understanding, and have a low degree of automation in research path planning and deep connection discovery, resulting in low research efficiency.

Method used

A system for in-depth research on museum exhibits integrating a multimodal big model is constructed, including a multimodal data aggregation module, a cross-modal semantic alignment module, a research path generation module, a dynamic argumentation construction module, and a research conclusion generation module. This system enables unified semantic representation and deep association of image and text data, automatically plans research paths, and constructs structured argumentation structures.

Benefits of technology

It has enabled the intelligent and automated study of museum exhibits, enhanced the systematicness and rigor of the research process, and improved the efficiency and quality of research output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962628A_ABST
    Figure CN121962628A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning, and particularly discloses a museum exhibit deep research system fused with a multi-modal large model. The system comprises a multi-modal data aggregation module, a cross-modal semantic alignment module, a research path generation module, a dynamic argumentation construction module and a research conclusion generation module, through multi-modal data fusion, semantic alignment and intelligent path planning, an argumentation structure is automatically constructed, and a standard research conclusion is generated. The whole-process intelligent support of exhibit research is realized, and the research efficiency and systematicness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

In-depth research system for museum exhibits integrating multimodal large models Technical Field

[0001] This invention belongs to the field of deep learning technology, specifically relating to a deep research system for museum exhibits that integrates multimodal large models. Background Technology

[0002] Artificial intelligence technology is being applied more and more widely in the field of cultural heritage, aiming to improve the efficiency and depth of digital preservation, display, and research of cultural relics. Multimodal data processing and intelligent analysis, as a key technological direction, are committed to achieving a comprehensive understanding and value discovery of cultural heritage by integrating multi-source information such as visual and textual data.

[0003] In-depth academic research on museum exhibits constitutes an important specific application scenario. The goal of this technological direction is to integrate scattered exhibit information, academic literature, and image data, and to assist researchers in exploring the historical background, artistic features, and cultural connections of exhibits through intelligent analysis, thereby enhancing the systematicness and innovation of research work.

[0004] Existing technologies primarily rely on independent digital document databases, collection information management systems, and traditional search engines. These tools generally lack the ability to deeply understand the semantics of unstructured documents, making it impossible to automatically associate and fuse cross-modal information. When conducting multi-source information verification, researchers must manually perform cross-referencing and comparative analysis, resulting in low research efficiency and a high dependence on personal experience.

[0005] The existing system architecture fails to integrate advanced large language models and intelligent research algorithms, making it difficult to automatically plan research paths or systematically uncover deep connections in massive amounts of data, which seriously restricts the intelligent development level of in-depth research on museum exhibits. Summary of the Invention

[0006] This invention aims to provide a system for in-depth research on museum exhibits that integrates multimodal large models, in order to solve the problems in existing technologies such as reliance on manual cross-referencing, lack of automatic cross-modal information association and deep semantic understanding capabilities, and low degree of automation in research path planning and deep connection discovery.

[0007] The technical solution of this invention is a museum exhibit in-depth research system that integrates multimodal large models. The system includes a multimodal data aggregation module, a cross-modal semantic alignment module, a research path generation module, a dynamic argumentation construction module, and a research conclusion generation module.

[0008] The multimodal data aggregation module is used to collect and structure the storage of exhibit image data, exhibit basic attribute data, full-text data of relevant academic literature, and historical archive text data.

[0009] The cross-modal semantic alignment module is used to perform joint embedding representation learning on image data and text data output by the multimodal data aggregation module, mapping data from different modalities to a unified semantic vector space, and realizing a deep association between visual features and text concepts.

[0010] The research path generation module is based on the unified semantic representation output by the cross-modal semantic alignment module. It automatically identifies the key elements of the research topic through a preset research goal parsing algorithm and dynamically generates multiple candidate research paths based on the semantic correlation strength between the elements.

[0011] The dynamic argument construction module is used to retrieve relevant evidence fragments from the multimodal data aggregation module in real time based on the research path selected by the research path generation module, and automatically construct an argument structure that supports the research hypothesis based on the logical relationship and credibility weight between the evidence.

[0012] The research conclusion generation module receives the complete argument structure output by the dynamic argument construction module, and uses a preset conclusion synthesis algorithm to summarize and refine the argument chain to generate a final research conclusion text with a standardized format and evidence citations.

[0013] Furthermore, the multimodal data aggregation module specifically includes an image acquisition unit, a text acquisition unit, and a data preprocessing unit.

[0014] The image acquisition unit acquires multi-angle image data of the exhibits through high-resolution digital scanning equipment and extracts semantic embedding feature vectors of the images using a pre-trained multimodal embedding model. The text acquisition unit retrieves academic papers, monograph chapters, historical records, and curatorial descriptions related to the exhibits from digital libraries, academic databases, and the museum's internal archives. The data preprocessing unit performs word segmentation, stop word removal, and named entity recognition on the raw text acquired by the text acquisition unit, and normalizes and reduces the dimensionality of the image features output by the image acquisition unit. Finally, the processed multimodal data, combined with the entity recognition results and relationship recognition results, is stored in a distributed graph database in a unified index format. An entity relationship network with the exhibits as the core is constructed in the graph database.

[0015] Furthermore, the cross-modal semantic alignment module is implemented using a two-stream neural network architecture, which includes a visual encoding branch and a text encoding branch.

[0016] The visual encoding branch takes exhibit image data as input, extracts the block embedding sequence of the image through a pre-trained visual transformer model, and calculates the global feature vector of the image through a multi-layer self-attention mechanism.

[0017] The text encoding branch takes exhibit-related text data as input, obtains the token embedding sequence of the text through a pre-trained language model, and obtains the text context feature vector through bidirectional long short-term memory network encoding.

[0018] The outputs of the visual encoding branch and the text encoding branch interact through a cross-modal attention layer to calculate the attention weight between image features and text features, and then perform weighted fusion of the two features based on this weight to generate a unified multimodal joint embedding vector.

[0019] Furthermore, the research path generation module includes a topic analysis unit and a path planning unit.

[0020] The topic parsing unit uses a keyword extraction algorithm to identify core entities and relational phrases from the research questions input by the user, and calculates the distribution representation of these entities and phrases in the semantic space based on the joint embedding vectors provided by the cross-modal semantic alignment module.

[0021] Based on the semantic distribution output by the topic parsing unit, the path planning unit uses a graph-based exploration algorithm to find related nodes directly connected to the core entity in the museum knowledge graph. It then sorts the related paths according to the semantic similarity score of the edges between nodes and outputs the top three paths with the highest scores as candidate research paths for the user to choose from.

[0022] Furthermore, the dynamic argument construction module includes an evidence retrieval unit, a logical relationship identification unit, an evidence cross-validation unit, and an argument graph construction unit. The evidence retrieval unit performs multi-hop queries in the graph database based on the research path determined by the research path generation module, retrieving image region descriptions, literature citations, and historical event records related to the path nodes as evidence material. The logical relationship identification unit performs semantic role labeling and dependency parsing on the evidence material returned by the evidence retrieval unit, identifying supporting, refuting, or supplementary logical relationships in the evidence, and assigning a confidence score based on statistical learning to each relationship. The evidence cross-validation unit cross-compares evidence material from different sources, calculating content consistency, source authority, and timeline matching between evidence, eliminating contradictory or low-credibility evidence, and generating a cross-validation report. The argument graph construction unit constructs a tree-like argument structure graph with the research hypothesis as the root node, the identified logical relationships as directed edges, and the evidence material as leaf nodes, and prunes and optimizes the argument branches based on the confidence scores and cross-validation results.

[0023] Furthermore, the research conclusion generation module is implemented using a text generation model based on an encoder-decoder architecture.

[0024] The encoder part of the model converts the argument structure graph output by the dynamic argument building module into a fixed-dimensional graph embedding vector.

[0025] The decoder uses graph embedding vectors as the initial hidden state and generates the research conclusion text word by word through an autoregressive approach. During the generation process, key phrases and data are copied from the original evidence material through a pointer network mechanism to ensure the accuracy and verifiability of the conclusion text.

[0026] The final output text from the research conclusion generation module includes a statement of the research question, a summary of the argumentation process, a summary of the core findings, and the corresponding index number of the evidence sources.

[0027] Furthermore, the system also includes a research process tracing module, which records all intermediate decisions and data operations of the research path generation module, dynamic argumentation construction module, and research conclusion generation module throughout the research process.

[0028] The research process tracing module stores the user-selected research path, the evidence list retrieved by the system, the constructed argumentation graph structure, and the attention distribution data of the conclusion generation model in chronological order as an interactive log file, allowing users to trace back the entire derivation process of any research conclusion at any time.

[0029] Furthermore, the system runs on a cloud-based containerized platform, and through a microservice architecture, it deploys the multimodal data aggregation module, cross-modal semantic alignment module, research path generation module, dynamic argumentation construction module, research conclusion generation module, and research process tracing module as independent service units that exchange data through a lightweight communication protocol.

[0030] The research process traceability module also includes a research process compliance check unit. This unit automatically verifies the compliance of the research path selection logic, evidence citation norms, and the rationality of the argumentation structure based on the academic norms and industry standards of museum exhibit research. It marks non-compliant links and provides correction suggestions. The compliance check results are stored synchronously with the log file.

[0031] The system front end provides a web-based visual research workbench, which integrates a visual view of the research path, an argumentation structure editing interface, and a conclusion report preview panel, supporting researchers to conduct interactive exploration and manual revisions.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows: by constructing a multimodal data aggregation module and a cross-modal semantic alignment module, the present invention realizes a unified semantic representation and deep association of museum exhibit images and text data, overcoming the semantic gap problem caused by modal isolation in traditional systems.

[0033] The research path generation module can automatically plan multiple research paths based on semantic associations, reducing the time cost and cognitive load for researchers to manually design research plans.

[0034] The dynamic argumentation construction module builds a structured argumentation support system through automated evidence retrieval and logical relationship identification, thereby enhancing the systematicness and rigor of the research process.

[0035] The research conclusion generation module automatically converts the argument structure into standardized text and ensures the traceability of conclusions to original evidence, greatly improving the efficiency and quality of academic research output.

[0036] Through modular design and cloud deployment, the entire system provides intelligent research support throughout the entire process, from data integration to conclusion generation, promoting the development of in-depth research on museum exhibits towards automation, systematization, and interpretability. Attached Figure Description

[0037] Figure 1 is a schematic diagram of the overall technical solution architecture of the museum exhibit in-depth research system integrating multimodal large models proposed in this invention; Figure 2 is a schematic diagram of the core principle framework of the cross-modal semantic alignment module in this invention; Figure 3 is a logical flow framework diagram of the research path generation module in this invention; Figure 4 is a schematic diagram of the multi-level interaction relationship and data flow of the dynamic argumentation construction module in this invention. Detailed Implementation

[0038] Please refer to Figure 1. This embodiment details the specific technical implementation of a system for in-depth research of museum exhibits that integrates a multimodal large model. The system is deployed on a cloud-based containerized platform and employs a microservice architecture, deploying core functional modules as independent service units that exchange data via lightweight communication protocols.

[0039] The system front end provides a web-based visual research workbench, which integrates a visual view of the research path, an argumentation structure editing interface, and a conclusion report preview panel, supporting researchers to conduct interactive exploration and manual revisions.

[0040] The core of the system includes a multimodal data aggregation module, a cross-modal semantic alignment module, a research path generation module, a dynamic argumentation construction module, a research conclusion generation module, and a research process tracing module.

[0041] The multimodal data aggregation module is responsible for collecting and structuring all raw data related to museum exhibits. This module specifically includes an image acquisition unit, a text acquisition unit, and a data preprocessing unit.

[0042] The image acquisition unit acquires multi-angle image data of exhibits using a high-resolution digital scanning device. The scanning device typically has a resolution of 6000×4000 pixels, and at least eight images from different angles are acquired for each exhibit. The image acquisition unit has a built-in pre-trained multimodal embedding model (such as the CLIP model) that automatically extracts the semantic embedding feature vector of each image, which has a dimension of 512.

[0043] The text acquisition unit automatically retrieves textual materials related to the exhibits from digital libraries, academic databases, and the museum's internal archives. The search process is based on a pre-defined list of keywords, including exhibit names, production dates, material types, excavation sites, and relevant historical figures. The textual materials include full-text academic papers, monograph chapters, scanned texts of original historical documents, and curatorial documents.

[0044] The data preprocessing unit performs a multi-level processing flow on the raw text acquired by the text acquisition unit.

[0045] First, word segmentation is performed using a word segmentation algorithm that combines a dictionary and a statistical model, with a dictionary size exceeding 500,000 entries. Then, stop words are removed; the stop word list contains approximately 2,000 commonly used function words and high-frequency words without substantive meaning. Next, named entity recognition is performed to identify entities such as personal names, place names, organization names, and time expressions in the text, and each entity is labeled with its type and position offset in the text.

[0046] For image data, the data preprocessing unit normalizes the feature vectors output by the image acquisition unit, scaling the numerical range of each feature vector to between 0 and 1. Then, principal component analysis is used to reduce the dimensionality of the features, reducing the 2048-dimensional deep convolutional neural network features to 512 dimensions while retaining more than 95% of the original information variance.

[0047] Finally, all processed multimodal data is stored in a distributed graph database using a unified index format. This graph database adopts the Neo4j architecture, where each data node contains a unique identifier, data type label, feature vector field, and original data reference links. It also labels the entity type of nodes (e.g., exhibits, historical figures, historical events) based on entity recognition results and constructs relationship edges between nodes based on relationship recognition results. Relationship types include belonging to the same exhibit, originating from the same document, belonging to the same era, and figures associated with exhibits.

[0048] Please refer to Figure 2. The cross-modal semantic alignment module receives the image features and text features output by the multimodal data aggregation module and maps them to a unified semantic vector space.

[0049] This module is implemented using a two-stream neural network architecture, which includes a visual encoding branch and a text encoding branch.

[0050] The visual coding branch takes exhibit image data as input, specifically receiving the 512-dimensional image feature vector output by the data preprocessing unit.

[0051] This branch employs a pre-trained visual transformer model containing 12 encoding layers, each with 8 self-attention heads. The image feature vector is first segmented into 16 block embedding sequences, each with a dimension of 64; then, it undergoes multi-layer self-attention mechanism computation, with the hidden layer dimension of the self-attention layer being 512; finally, a global average pooling operation is used to obtain the global feature vector of the image, which also has a dimension of 512.

[0052] The text encoding branch takes exhibit-related text data as input, specifically receiving the segmented text sequence output by the data preprocessing unit. This branch employs a pre-trained language model, with a Transformer-based bidirectional encoder representation model and a vocabulary size of 30522.

[0053] The text sequence is first converted into a token embedding sequence, with each token embedding having a dimension of 768; then it is encoded by a bidirectional long short-term memory network with two layers and a hidden layer size of 512; finally, the text context feature vector is obtained by mean pooling of the last hidden state, and this vector has a dimension of 512.

[0054] The outputs of the visual encoding branch and the text encoding branch interact through a cross-modal attention layer. The cross-modal attention layer calculates the attention weights between image features and text features, specifically through a query key-value mechanism.

[0055] Using image features as the query vector and text features as the key and value vectors, the formula for calculating the attention score is: ; where query vector The image's global feature vector, key vector For text context feature vectors, value vectors Similarly, for text context feature vectors, the scaling factor is square root. In The value is 512.

[0056] The two features are weighted and fused based on the calculated attention weights to generate a unified multimodal joint embedding vector with a dimension of 512.

[0057] This joint embedding vector enables a deep association between visual features and textual concepts, making images and texts with similar semantics close in distance in the vector space.

[0058] Please refer to Figure 3. The research path generation module automatically plans research paths based on the unified semantic representation output by the cross-modal semantic alignment module. This module includes a topic parsing unit and a path planning unit.

[0059] The topic parsing unit uses a keyword extraction algorithm to identify core entities and relational phrases from the research questions input by the user.

[0060] The keyword extraction algorithm adopts a method that combines word frequency inverse document frequency statistics and graph ranking algorithm. First, the word frequency inverse document frequency score of each word in the input text is calculated. Then, a word co-occurrence graph is constructed, in which the nodes are candidate words and the edge weights are the number of co-occurrences. Finally, the PageRank value of each node is calculated iteratively, and the top 5 words are selected as core entities and relational phrases.

[0061] The topic parsing unit computes the distribution representation of these entities and phrases in the semantic space based on the joint embedding vectors provided by the cross-modal semantic alignment module.

[0062] Specifically, for each identified keyword, the corresponding node is found in the museum knowledge graph, and all multimodal joint embedding vectors associated with that node are obtained. The centroids of these vectors are then calculated as the semantic distribution representation of the keyword.

[0063] Based on the semantic distribution output by the topic parsing unit, the path planning unit uses a graph-based exploration algorithm to find related nodes directly connected to the core entity in the museum knowledge graph.

[0064] The exploration algorithm starts from each core entity node, traverses all its first-degree neighbor nodes, and calculates the cosine similarity score between the neighbor nodes and the core entity node in the semantic space.

[0065] The formula for calculating cosine similarity is: Where vector The semantic vector of the core entity node, the vector This represents the semantic vector of the neighboring nodes. The path planning unit sorts the associated paths based on the semantic similarity score of the edges between nodes, with the edge score being the average of the semantic similarity scores of the two endpoints.

[0066] The top three highest-scoring paths are then output as candidate research paths for users to choose from. Each path contains a sequence of 2 to 4 nodes, including exhibits, historical events, and documents.

[0067] Please refer to Figure 4. The dynamic argumentation construction module automatically constructs an argumentation structure that supports the research hypothesis based on the research path selected by the research path generation module.

[0068] The dynamic argumentation construction module includes an evidence retrieval unit, a logical relationship identification unit, an evidence cross-verification unit, and an argumentation diagram construction unit.

[0069] The evidence retrieval unit performs multi-hop queries in the graph database based on the research path determined by the research path generation module.

[0070] The query statement was written in Cypher query language. Starting from the starting node of the path, it traversed along the relation edges to the node with a depth of 3, and retrieved the image region description, literature citations and historical event records related to the path node as evidence material.

[0071] Each query returns a maximum of 100 pieces of evidence, each piece of evidence including the content text, source identifier, timestamp, and initial confidence value.

[0072] The logical relationship identification unit performs deep semantic analysis on the evidence materials returned by the evidence retrieval unit. First, semantic role labeling is performed to identify the predicate argument structure in each sentence and label semantic roles such as agent, patient, time, and place. Then, dependency parsing is performed to construct the dependency tree of the sentence and identify grammatical dependencies such as subject-predicate, verb-object, and attributive-head relations. Finally, based on the analysis results, supporting, refuting, or supplementary logical relationships in the evidence are identified.

[0073] Relationship identification employs a combination of rule-based and statistical learning methods. The rule base contains 200 logical relation patterns, and the statistical model is a support vector machine classifier trained on an argument corpus. The classifier assigns a confidence score based on statistical learning to each relation, with the score ranging from 0 to 1.

[0074] The argumentation diagram construction unit uses the research hypothesis as the root node, the identified logical relationships as directed edges, and the evidence materials as leaf nodes to construct a tree-like argumentation structure diagram.

[0075] The argument graph is stored using a directed acyclic graph (DAG) data structure. Node attributes include node type, content summary, and confidence score; edge attributes include relation type, weight value, and creation time. The argument graph construction unit performs pruning optimization on the argument branches based on the confidence score, removing edges and their subtrees with a confidence score below 0.6, and retaining at least 3 main argument branches to ensure the completeness of the argument.

[0076] The cross-validation unit cross-compares evidence materials from different sources. First, it extracts the core information summary and source attributes (such as journal level and archive preservation level) of each piece of evidence, calculates the semantic similarity at the content level to verify consistency, and checks whether the historical timeline corresponding to the evidence matches. For evidence with a semantic similarity of less than 0.5 or contradictory timelines, it is marked as low credibility and excluded from the evidence material library. Finally, it generates a cross-validation report that includes verification dimensions, results and explanations of abnormal evidence.

[0077] The research conclusion generation module receives the complete argument structure output by the dynamic argument construction module and generates the final research conclusion text. This module is implemented using a text generation model based on an encoder-decoder architecture.

[0078] The model encoder part converts the argument structure graph output by the dynamic argument building module into a fixed-dimensional graph embedding vector.

[0079] The encoder is implemented using a graph attention network with 3 layers, each containing 4 attention heads.

[0080] The graph attention network takes the feature vectors of all nodes in the argument graph as input, and through multiple layers of message passing and aggregation operations, finally generates a 256-dimensional graph embedding vector through a readout function. The decoder part uses the graph embedding vector as the initial hidden state and generates the research conclusion text word by word through an autoregressive approach.

[0081] The decoder is a Transformer-based generative model with a vocabulary size of 50,000 and a maximum generation length of 1,000 tokens.

[0082] During the generation process, the model copies key phrases and data from the original evidence material through a pointer network mechanism.

[0083] The pointer network calculates the matching probability between generated words and words in the evidence material. When the matching probability exceeds the threshold of 0.7, the words in the evidence material are directly copied to the output sequence to ensure the accuracy and verifiability of the conclusion text.

[0084] The final output text from the research conclusion generation module contains four fixed parts: the research question statement section directly restates the research question entered by the user; the argumentation process summary section summarizes the main argumentation branches generated by the dynamic argumentation construction module and their supporting evidence; the core findings summary section extracts the key conclusions in the argumentation chain; and the evidence source index number section lists the unique identifiers of all cited evidence in the database, in the format of evidence type abbreviation plus an 8-digit number.

[0085] The research process tracing module records all intermediate decisions and data operations made by the research path generation module, dynamic argumentation construction module, and research conclusion generation module throughout the entire research process.

[0086] This module employs an event-driven architecture, storing the user-selected research path, the evidence list retrieved by the system, the constructed argument graph structure, and the attention distribution data of the conclusion generation model in chronological order as an interactive log file. The log file is in JSON format, and each operation record includes a timestamp, module identifier, operation type, input data hash value, and output data summary.

[0087] The research process traceability module also includes a research process compliance check unit. This unit has a built-in academic norms library for museum exhibit research, including evidence citation format standards and logical guidelines for research path design. This unit scans the research process log data in real time, verifying whether evidence citations are properly attributed, whether there are logical jumps in the research path, and whether the argumentation structure conforms to academic argumentation paradigms. Non-compliant steps are marked in red and specific correction suggestions are pushed out. The compliance check results are appended in JSON format to the corresponding operation record in the log file.

[0088] The system front-end provides a research process replay function, allowing users to review the entire derivation process of any research conclusion at any time. The replay speed is adjustable and supports pausing and single-step execution.

[0089] The system's modules exchange data via a lightweight communication protocol based on the gRPC framework, with data serialization using Protocol Buffers format. Each microservice is deployed in an independent Docker container, with a container resource quota of 4 CPU cores and 16GB of memory.

[0090] The system utilizes a service mesh for traffic management, fault recovery, and security policy enforcement, ensuring high availability and scalability. The front-end research workbench, developed using the React framework, supports real-time collaborative editing, allowing up to five researchers to work on the same research project simultaneously.

[0091] The research workbench maintains a long-term connection with the backend microservices via the WebSocket protocol, enabling real-time updates of the research path visualization view, bidirectional synchronization of the argumentation structure editing interface, and instant rendering of the conclusion report preview panel.

[0092] This embodiment provides an alternative implementation scheme for a museum exhibit in-depth research system that integrates multimodal large models, focusing on its differentiated design in the cross-modal semantic alignment module and dynamic argument construction module.

[0093] In the cross-modal semantic alignment module, this embodiment uses a multimodal representation learning method based on contrastive learning to replace the two-stream neural network architecture.

[0094] This method performs modality alignment directly at the raw data level, without the need to pre-extract image and text features.

[0095] Specifically, the module receives the original exhibit images and related text pairs output by the multimodal data aggregation module, and generates positive and negative sample pairs through data augmentation technology.

[0096] Positive sample pairs consist of an image of the same exhibit and the correct descriptive text, while negative sample pairs consist of randomly combined images and texts of different exhibits.

[0097] The module employs a symmetric encoder structure, with the image encoder using the VisionTransformer model and the text encoder using the BERT model. Both encoders output vectors have a dimension of 768. The contrastive learning objective function uses normalized temperature-scaled cross-entropy loss, with the temperature parameter set to 0.05.

[0098] By maximizing the similarity of positive sample pairs while minimizing the similarity of negative sample pairs, the model learns a joint representation space across modalities. The advantage of this method lies in its end-to-end training, avoiding error accumulation in the feature extraction and alignment stages.

[0099] In the dynamic argument construction module, this embodiment adopts an argument construction method based on knowledge graph reasoning instead of the method based on semantic role labeling and dependency parsing.

[0100] This method views argument construction as a reasoning process on a museum knowledge graph.

[0101] The evidence retrieval unit performs path-based graph queries, starting from the research path nodes and traversing along predefined argumentation relationship edges. These relationship edges include supporting relationships, rebuttal relationships, illustrative relationships, etc., and the relationship definitions are based on domain ontology.

[0102] The logical relationship recognition unit uses a graph neural network model to directly infer the logical relationships between nodes on the graph structure.

[0103] The model takes a subgraph structure as input, aggregates neighbor node information through multi-layer graph convolution operations, and finally outputs the probability distribution of various logical relationships between any two nodes.

[0104] The argument graph construction unit adopts the argument structure optimization method based on the minimum spanning tree algorithm, using the reciprocal of the confidence of the logical relationship as the edge weight to construct the minimum cost argument tree connecting all relevant evidence.

[0105] This method is particularly suitable for scenarios where argument elements are scattered across multiple corners of a knowledge graph, and it can discover non-local argument support relationships.

[0106] The implementation of the remaining modules of the system, including the multimodal data aggregation module, research path generation module, research conclusion generation module, and research process tracing module, is basically the same as described above, but the parameter configurations have been adjusted.

[0107] The path planning unit of the research path generation module outputs the top 5 candidate research paths, increasing the breadth of research exploration.

[0108] The decoder of the research conclusion generation module uses a sequence generation model based on convolutional neural networks instead of the Transformer model, which has higher parallel efficiency when generating long texts.

[0109] In addition to recording operation logs, the research process traceability module also records the inference time and resource consumption data of each module, providing a basis for system performance optimization.

[0110] The system deployment architecture is the same as described above, both adopting a cloud-based containerized microservice architecture. However, in terms of service discovery and load balancing strategies, this embodiment uses a request routing algorithm based on consistent hashing to ensure that requests from the same research session are always routed to the same group of service instances, maintaining the consistency of the research status.

[0111] The front-end research workbench has added an argument strength visualization function, using lines of different colors and thicknesses to represent the confidence level of argument branches, helping researchers to intuitively assess the robustness of arguments.

[0112] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0113] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A system for in-depth research of museum exhibits integrating a multimodal large model, characterized in that, include: The multimodal data aggregation module is used to collect and structure the storage of exhibit image data, exhibit basic attribute data, relevant academic literature full-text data, and historical archive text data; The cross-modal semantic alignment module performs joint embedding representation learning on image and text data output by the multimodal data aggregation module, mapping data from different modalities to a unified semantic vector space to achieve a deep association between visual features and textual concepts. The research path generation module, based on the unified semantic representation output by the cross-modal semantic alignment module, automatically identifies key elements of the research topic using a pre-defined research objective parsing algorithm and dynamically generates multiple candidate research paths based on the semantic correlation strength between elements. The dynamic argument construction module retrieves relevant evidence fragments from the multimodal data aggregation module in real time according to the research path selected by the research path generation module, and automatically constructs an argument structure supporting the research hypothesis based on the logical relationships and credibility weights between the evidence. The research conclusion generation module receives the complete argument structure output by the dynamic argument construction module, summarizes and refines the argument chain using a pre-defined conclusion synthesis algorithm, and generates a final research conclusion text with a standardized format and evidence citations.

2. The system for in-depth research of museum exhibits integrating multimodal large models according to claim 1, characterized in that, The multimodal data aggregation module includes an image acquisition unit, a text acquisition unit, and a data preprocessing unit. The image acquisition unit acquires multi-angle image data of the exhibits through a high-resolution digital scanning device and extracts semantic embedding feature vectors of the images using a pre-trained multimodal embedding model. The text acquisition unit retrieves academic papers, monograph chapters, historical records, and curatorial descriptions related to the exhibits from digital libraries, academic databases, and the museum's internal archives. The data preprocessing unit performs word segmentation, stop word removal, and named entity recognition operations on the raw text acquired by the text acquisition unit, and normalizes and reduces the dimensionality of the image features output by the image acquisition unit. Finally, the processed multimodal data, combined with the entity recognition results and relationship recognition results, is stored in a distributed graph database according to a unified index format. An entity relationship network with the exhibits as the core is constructed in the graph database.

3. The system for in-depth research of museum exhibits integrating a multimodal large model as described in claim 1, characterized in that, The cross-modal semantic alignment module is implemented using a two-stream neural network architecture, which includes a visual encoding branch and a text encoding branch. The visual encoding branch takes exhibit image data as input, extracts the block embedding sequence of the image through a pre-trained visual transformer model, and calculates the global feature vector of the image through a multi-layer self-attention mechanism. The text encoding branch takes exhibit-related text data as input, obtains the token embedding sequence of the text through a pre-trained language model, and obtains the text context feature vector through bidirectional long short-term memory network encoding. The outputs of the visual encoding branch and the text encoding branch interact through a cross-modal attention layer to calculate the attention weight between image features and text features, and then performs weighted fusion of the two features based on this weight to generate a unified multimodal joint embedding vector.

4. The system for in-depth research of museum exhibits integrating multimodal large models according to claim 1, characterized in that, The research path generation module includes a topic parsing unit and a path planning unit. The topic parsing unit uses a keyword extraction algorithm to identify core entities and relational phrases from the research questions input by the user, and calculates the distribution representation of these entities and phrases in the semantic space based on the joint embedding vectors provided by the cross-modal semantic alignment module. The path planning unit uses a graph-based exploration algorithm to find related nodes directly connected to the core entities in the museum knowledge graph based on the semantic distribution output by the topic parsing unit, and sorts the related paths according to the semantic similarity score of the edges between nodes, outputting the top 3 paths with the highest scores as candidate research paths for the user to choose from.

5. The system for in-depth research of museum exhibits integrating multimodal large models according to claim 1, characterized in that, The dynamic argumentation construction module includes an evidence retrieval unit, a logical relationship identification unit, an evidence cross-verification unit, and an argumentation graph construction unit. The evidence retrieval unit performs multi-hop queries in the graph database according to the research path determined by the research path generation module, and retrieves image region descriptions, literature citations, and historical event records related to the path nodes as evidence materials. The logical relationship identification unit performs semantic role labeling and dependency parsing on the evidence materials returned by the evidence retrieval unit, identifies supporting, refuting, or supplementary logical relationships in the evidence, and assigns a confidence score based on statistical learning to each relationship. The evidence cross-validation unit cross-compares evidence materials from different sources, calculates the consistency of content, authority of sources, and timeline matching between evidence, eliminates contradictory or low-credibility evidence, and generates a cross-validation report. The argument graph construction unit constructs a tree-like argument structure graph with the research hypothesis as the root node, the identified logical relationships as directed edges, and the evidence materials as leaf nodes, and prunes and optimizes the argument branches based on the confidence scores and cross-validation results.

6. The system for in-depth research of museum exhibits integrating a multimodal large model as described in claim 1, characterized in that, The research conclusion generation module is implemented using a text generation model based on an encoder-decoder architecture. The encoder part of the model converts the argument structure graph output by the dynamic argument construction module into a fixed-dimensional graph embedding vector. The decoder part uses the graph embedding vector as the initial hidden state and generates the research conclusion text word by word through an autoregressive method. During the generation process, key phrases and data are copied from the original evidence material through a pointer network mechanism to ensure the accuracy and verifiability of the conclusion text. The final output text of the research conclusion generation module includes a statement of the research question, a summary of the argumentation process, a summary of the core findings, and the corresponding evidence source index number.

7. The system for in-depth research of museum exhibits integrating multimodal large models according to claim 1, characterized in that, The system for in-depth research on museum exhibits, which integrates a multimodal large model, also includes a research process tracing module. This module records all intermediate decisions and data operations of the research path generation module, dynamic argumentation construction module, and research conclusion generation module throughout the entire research process. The research process tracing module stores the user-selected research path, the evidence list retrieved by the system, the constructed argumentation graph structure, and the attention distribution data of the conclusion generation model in chronological order as an interactive log file, allowing users to review the entire derivation process of any research conclusion at any time. The research process tracing module also includes a research process compliance check unit. This unit automatically verifies the compliance of the research path selection logic, evidence citation norms, and the rationality of the argumentation structure based on the academic norms and industry standards of museum exhibit research. It marks non-compliant links and provides correction suggestions. The compliance check results are stored synchronously with the log file.

8. The system for in-depth research of museum exhibits integrating multimodal large models according to claim 1, characterized in that, The system, which integrates a multimodal large model for in-depth research on museum exhibits, runs on a cloud-based containerized platform. Through a microservice architecture, it deploys the multimodal data aggregation module, cross-modal semantic alignment module, research path generation module, dynamic argument construction module, research conclusion generation module, and research process tracing module as independent service units that exchange data via a lightweight communication protocol. The system front end provides a web-based visual research workbench, which integrates a visual view of the research path, an argument structure editing interface, and a conclusion report preview panel, supporting researchers to conduct interactive exploration and manual correction.

9. The system for in-depth research of museum exhibits integrating a multimodal large model according to claim 3, characterized in that, The cross-modal attention layer calculates the attention weights between image features and text features, specifically through a query key-value mechanism. Using image features as query vectors and text features as key and value vectors, attention scores are calculated, and the two features are weighted and fused based on these scores to generate a unified multimodal joint embedding vector.

10. The system for in-depth research of museum exhibits integrating multimodal large models according to claim 5, characterized in that, The logical relationship identification unit uses a combination of rule-based and statistical learning methods to identify relationships. The rule base contains multiple logical relationship patterns, and the statistical model is a support vector machine classifier trained on an argument corpus. The classifier assigns a confidence score based on statistical learning to each relationship, with the score ranging from 0 to 1.