Exhibit image customized interpretation generation system and method oriented to exhibition and display scenes

By using a customized explanation generation system for exhibit images designed for exhibition scenarios, and leveraging multimodal interaction and thought chain reasoning technologies, the system addresses the issues of low user customization, insufficient knowledge accuracy, and unnatural and fluent explanation styles found in existing technologies. This results in the generation of highly accurate, rich, and fluent explanation content.

CN121958593APending Publication Date: 2026-05-01INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF COMPUTING TECH CHINESE ACAD OF SCI
Filing Date
2025-12-23
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing exhibit description generation technologies suffer from low levels of user customization, insufficient accuracy and richness of knowledge, and a lack of fluency and naturalness in the description style.

Method used

A customized explanation generation system for exhibit images designed for exhibition scenarios was developed. The system receives user information through interactive units, combines user profiles and basic databases, acquires background knowledge through multimodal interaction, and generates explanations through a multimodal large language model. The system also introduces thought chain reasoning technology to improve the coherence and interpretability of the explanation content.

Benefits of technology

It enables the generation of highly customized commentary content, improving the accuracy and richness of the commentary, while also enhancing the fluency of the commentary and user trust.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121958593A_ABST
    Figure CN121958593A_ABST
Patent Text Reader

Abstract

The invention provides an exhibit image customization interpretation generation system and method oriented to an exhibition and display scene, and the system comprises an interaction unit which is used for receiving explanation customization information provided by a user, and the explanation customization information at least comprises a user portrait and a specified exhibit image; the basic database comprises images of all exhibits in the target exhibition and display scene, an introduction text of each exhibit and an image of each exhibit; the query prediction unit constructs a personalized database based on the explanation customization information and predicts a plurality of questions in which the user may be interested in the specified exhibit; the exhibit knowledge retrieval unit is used for retrieving a plurality of introduction text segments most related to the specified exhibit and all predicted problems from the basic database by adopting a preset knowledge retrieval method to form background knowledge of the specified exhibit, and each introduction text segment is a paragraph or a sentence of an introduction text; and the explanation generation unit is used for generating an explanation text of the specified exhibit based on the explanation customization information and the background knowledge.
Need to check novelty before this filing date? Find Prior Art

Description

A customized explanation generation system and method for exhibit images for exhibition scenarios Technical Field

[0001] This invention relates to the field of exhibition scene interpretation generation, and more specifically, to a customized interpretation generation system and method for exhibit images for exhibition scenes. Background Technology

[0002] In museums, exhibition halls, science and technology museums, art galleries, and other exhibition settings, exhibit explanations play a crucial role as a bridge for visitors to acquire knowledge and information. Through explanations, visitors can better understand the historical background, cultural value, and scientific significance of exhibits, thus gaining a richer and deeper experience and enhancing their interest and motivation to learn. Traditional human explanations can be tailored to the exhibit content and the needs of the visitors, but this is costly in terms of manpower and time. Electronic guides can provide standardized explanations, but they usually rely on pre-set content and cannot be customized or interactive based on the actual needs of the visitors. With the development of artificial intelligence, especially multimodal technology, generating customized exhibit explanations for visitors is becoming increasingly possible. By combining exhibit data with user profiles and understanding users' individual needs through real-time interaction, a tailored explanation experience can be provided, demonstrating broad application prospects in various exhibition settings.

[0003] Currently, representative methods applicable to intelligent exhibit narration generation include narration generation methods based on explicit entity perception, narration generation methods based on knowledge retrieval enhancement, and narration generation methods based on visual-language model reasoning. Among these, the narration generation method based on explicit entity perception is implemented using an encoder-decoder architecture. It extracts visual features through a CLIP image encoder, converts them into soft cues using a trained projector, and constructs hard cues through an entity classifier. Finally, the two cues are concatenated and input into a language model to generate accurate and fluent narration. While this method improves the accuracy of identifying unseen entities, it lacks the introduction of background knowledge beyond visual information, making it prone to creating illusions and resulting in low credibility. It also lacks explainability and does not support user interaction, leading to insufficient customization of the narration. The narration generation method based on knowledge retrieval enhancement automatically extracts structured information from an art text corpus to construct a context-aware knowledge graph. When generating narration, this graph is combined with multi-dimensional background information, thereby improving the accuracy and knowledge richness of the narration. While this method utilizes retrieval-enhanced generation techniques to improve the accuracy and richness of explanations, it lacks user interaction, personalization, and its explanations primarily list knowledge items, resulting in a less fluent and natural style. The vision-language model-based reasoning method decomposes the model's thought process into four stages: task summarization, image description, logical reasoning, and conclusion generation, enhancing the model's reasoning ability through structured thinking. This method uses Llama3.2-11B-Vision-Instruct as the baseline model and performs supervised fine-tuning to improve the model's accuracy and interpretability in question-answering tasks. By employing a stage-level bundle search method during the reasoning process, it further improves the accuracy and efficiency of reasoning. Although this method enhances user interactivity through question-answering, the interaction is limited to text and does not support multimodal interaction. Furthermore, the lack of background knowledge integration results in poor knowledge richness and accuracy, limiting its effectiveness in scenarios requiring the introduction of external knowledge.

[0004] In summary, existing exhibit description generation technologies suffer from problems such as low user customization, insufficient accuracy and richness of knowledge, and a lack of fluency and naturalness in the description style.

[0005] It should be noted that the background information provided is only for illustrating relevant information about the present invention to aid in understanding the technical solutions of the present invention, and does not imply that the relevant information is necessarily prior art. In the absence of evidence indicating that the relevant information was disclosed before the filing date of this invention, the relevant information should not be considered prior art. Summary of the Invention

[0006] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide a customized explanation generation system and method for exhibit images in exhibition scenarios.

[0007] The objective of this invention is achieved through the following technical solution:

[0008] A customized explanation generation system for exhibit images in exhibition scenarios is disclosed. The system provides personalized text explanations for exhibits within a target exhibition scenario. The system includes an interaction unit, a basic database, a query prediction unit, an exhibit knowledge retrieval unit, and an explanation generation unit. The interaction unit receives customized explanation information provided by the user, which includes at least a user profile and a specified exhibit image. The basic database includes images of all exhibits in the target exhibition scenario, introductory text for each exhibit, and images of each exhibit. The query prediction unit is configured with a personalized database construction module and a question prediction module. The personalized database construction module is used to select multiple introductory texts matching the user profile from the basic database to form a personalized explanation. The database includes a question prediction module for predicting multiple questions that users might be interested in regarding a specific exhibit based on a personalized database and customized explanation information. The exhibit knowledge retrieval unit uses a preset knowledge retrieval method to retrieve multiple introductory text segments most relevant to the specified exhibit and all predicted questions from the basic database to form background knowledge for the specified exhibit. Each introductory text segment is a paragraph or sentence within the introductory text. The explanation generation unit is equipped with a pre-trained explanation generation model for generating explanatory text for the specified exhibit based on customized explanation information and background knowledge. This explanatory generation model is a multimodal large language model trained by taking customized explanation information and background knowledge as input, sequentially executing multiple preset inference tasks, and outputting the explanatory text for the specified exhibit.

[0009] Optionally, the personalized database construction module can also construct a personalized database in the following way: using a preset large language model to rewrite all the introductory text in the basic database into text that matches the expression style of the user profile, and then assemble the rewritten text into a personalized database.

[0010] Optionally, the question prediction module is configured to: use a preset large language model to filter introductory text segments from a personalized database that match the customized information being explained, generate multiple questions with the filtered introductory text segments as answers, and use all generated questions as predictions of what users might be interested in regarding the specified exhibits.

[0011] Optionally, in the explanation generation unit, the explanation text generation model is constructed in the following manner: A training dataset is constructed, comprising multiple labeled samples, wherein each sample includes an image of an exhibit, a question of interest to the user, and background knowledge of the exhibit constructed using a preset knowledge retrieval method, and the label is the explanation text corresponding to the sample; a multimodal large language model is obtained, and the multimodal large language model is trained based on the samples in the training dataset as input, reasoning using a preset thought chain, and the explanation text corresponding to the sample as output, to obtain the explanation text generation model, wherein the preset thought chain consists of multiple task instructions and is used to guide the multimodal large language model to sequentially execute multiple preset reasoning tasks to generate explanation text for the exhibit.

[0012] Optionally, the preset knowledge retrieval enhancement method is as follows: coarse-grained retrieval, which selects a first preset number of introductory texts with the highest similarity to the images of the exhibits from the basic database to form a candidate text set; fine-grained retrieval, which divides all the introductory texts in the candidate text set into multiple introductory text segments, using the images of the exhibits and questions of interest to users as the joint information of the exhibits, and selects a second preset number of introductory text segments with the highest similarity to the joint information of the exhibits from all the divided introductory text segments, and then combines all the selected introductory text segments into background knowledge.

[0013] Optionally, the coarse-grained retrieval is as follows: the specified exhibit image is encoded into a feature vector using a preset multimodal encoder; the cosine similarity between the feature vector of the specified exhibit image and the feature vector of each introductory text is obtained to obtain the similarity between the specified exhibit image and each introductory text; and a first preset number of introductory texts with the highest similarity are selected to form a candidate text set.

[0014] Optionally, the fine-grained retrieval is as follows: each introductory text in the candidate text set is divided into multiple introductory text segments based on sentences or paragraphs of a preset length; a multimodal retrieval tool is obtained and configured to calculate the similarity between each introductory text segment and the joint information of the exhibits, and to select a second preset number of introductory text segments with the highest similarity from all introductory text segments; all introductory text segments selected by the multimodal retrieval tool are used to form background knowledge.

[0015] Optionally, the multiple preset reasoning tasks include an exhibit description generation task, a sub-question-answer generation task, and an explanatory text generation task, wherein: the image description generation task is a task of generating visual information about the exhibit related to questions of interest to the user based on an image of the specified exhibit; the sub-question-answer generation task is a task of progressively decomposing the questions of interest to the user into a series of logically progressive sub-questions and generating text answers for each sub-question based on background knowledge of the specified exhibit; and the explanatory text generation task is a task of generating explanatory text for the specified exhibit based on the visual information of the exhibit and the text answers.

[0016] According to a second aspect of the present invention, a method for generating customized explanations of exhibit images for exhibition scenarios is proposed. The method includes: constructing customized explanation information for a specified exhibit based on user needs; generating explanatory text for the specified exhibit based on the customized explanation information using a customized explanation system as described in any one of claims 1-8; and providing an explanation of the specified exhibit based on the generated explanatory text.

[0017] Compared with the prior art, the advantages of the present invention are as follows:

[0018] The customized explanation system and method for exhibition scenarios proposed in this invention have the following three advantages: First, users can freely upload self-taken images of exhibits and ask questions instantly using language or drawing. The system combines user profiles to generate highly relevant question candidates and style-based explanations in real time, with a significantly higher degree of customization than template-based guides. Second, the two-stage multimodal retrieval first accurately retrieves question-related fragments from an authoritative knowledge base, and then expands related knowledge, effectively suppressing illusions and ensuring that the explanation content is accurate, rich, and traceable. Third, the introduction of thought chain reasoning first explicitly locates key visual areas and displays multi-step derivation processes, and then generates coherent and natural language explanations, which not only improves text fluency but also enhances the interpretability of the results and user trust. Attached Figure Description

[0019] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:

[0020] Figure 1 is a schematic diagram of an exhibit image customization explanation generation system for exhibition scenarios according to an embodiment of the present invention;

[0021] Figure 2 is a schematic diagram of a method for generating customized explanations of exhibit images for exhibition scenarios according to an embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention is further described in detail below through specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0023] As mentioned in the background technology section, existing methods for generating exhibit descriptions have relatively simple interactive forms, which cannot accommodate users' needs for more detailed and specific interactions and questions. The degree of customization of the description content is insufficient. Specifically, methods based on knowledge retrieval enhancement refer to preset exhibit knowledge when generating description content, which improves the accuracy of understanding the content. However, the emphasis on listing knowledge items makes the language style less fluent and natural when used for explanation. On the other hand, other description generation methods that do not refer to specific exhibit knowledge rely too much on model training, which is prone to illusion problems and cannot guarantee the accuracy and richness of knowledge in the explanation.

[0024] To address the aforementioned issues, this invention proposes a customized explanation generation system and method for exhibition scenarios. The system supports multimodal interaction using text and images, allowing users to ask questions about the overall structure or specific details of the exhibits. It can also predict potential interests and concerns based on user profiles. In generating customized explanations, this invention introduces retrieval enhancement technology to acquire background knowledge related to the exhibits and matching the user's questions, thereby improving the accuracy and richness of the explanations. Simultaneously, it employs thought chain reasoning technology to enhance the coherence and explainability of the explanation content, thus strengthening the user's interactive experience and sense of trust.

[0025] According to an embodiment of the present invention, a customized explanation generation system for exhibition scenarios is proposed. Referring to Figure 1, the system includes an interaction unit, a basic database, a query prediction unit, an exhibit knowledge retrieval unit, and an explanation generation unit. To facilitate understanding of the present invention, each unit in the system will be described in detail below with reference to specific embodiments.

[0026] I. Interactive Unit

[0027] The interactive unit is used to receive customized explanation information provided by the user. The customized explanation information includes at least the user profile and the image of the specified exhibit. In addition, the user can selectively input other information to describe the specified exhibit and user profile, such as exhibition type, text instructions and area labels.

[0028] According to one embodiment of the present invention, the interaction unit is a client or platform that provides users with an input information interface. Users can freely upload exhibit images or specify exhibits to be explained from preset images. Furthermore, the interaction unit can not only ask questions and identify areas of interest verbally, but also ask questions about the exhibit's overall or partial aspects by clicking or selecting areas of interest on the exhibit image. This provides users with richer and more direct interactive means to express their personalized needs. It should be noted that the user profile in the customized explanation information includes information such as age, professional background, and education level used to shape the user's personality and image. The interaction unit provides multimodal interaction methods including text and images. Users can ask questions or identify areas of interest regarding the exhibit's overall or partial aspects through any text, photos, or uploaded images. For example, they can ask questions about exhibit details, historical background, or describe areas of interest, such as "Please introduce the figure on the far left of this artwork." Images can be preset exhibit images or exhibit images uploaded by the user, and areas of interest can be drawn on the image in the form of dots or rectangles. The interactive unit uses multimodal input to acquire rich background knowledge, providing multimodal contextual basis for the subsequent generation of explanatory text, thereby improving the relevance, accuracy and personalization of the explanatory content.

[0029] II. Basic Database

[0030] The basic database stores images of exhibits, introductory text for each exhibit, and images of each exhibit, providing the system with visual and textual knowledge sources for exhibits to support enhanced retrieval and personalized explanation generation.

[0031] According to one embodiment of the present invention, the basic database consists of images of exhibits and introductory texts (textual knowledge documents) of exhibits. The images and introductory texts of exhibits are sourced from Baidu Encyclopedia, Wikipedia, and official exhibit descriptions from major museums. The basic database contains at least images of all exhibits in the target exhibition scene and their corresponding introductory texts. In addition, to further enhance the system's adaptability to multiple scenes, the database also synchronously includes images and introductory texts of exhibits in other exhibition scenes, enabling cross-scene knowledge reuse.

[0032] According to one embodiment of the present invention, the basic database also includes feature vectors for each exhibit image and each introductory text, wherein each feature vector is encoded by its corresponding exhibit image or introductory text through a visual-language contrastive learning encoder (such as the EVA-CLIP model) so that it can be called by the exhibit knowledge retrieval unit in the subsequent retrieval process to calculate similarity.

[0033] III. Query Prediction Unit

[0034] The query prediction unit is used to predict questions that users may be interested in regarding the exhibits based on the customized explanation information and the basic database. The query prediction unit is configured with a personalized database construction module and a question prediction module. The personalized database construction module is used to filter out multiple introductory texts that match the user profile from the basic database to form a personalized database. The question prediction module is used to predict multiple questions that users may be interested in regarding a specified exhibit based on the personalized database and the customized explanation information.

[0035] According to one embodiment of the present invention, to achieve customized explanation generation, the present invention constructs a personalized database of user profiles based on user profiles (such as age, professional background, and education level). Taking age as an example, the user group can be divided into three categories: children, students, and adults. Then, a corresponding personalized database is constructed based on the profiles to achieve accurate knowledge retrieval for different user profiles. Specifically, a large language model (such as GPT-4o) is used to rewrite the knowledge text in the basic database so that the knowledge content and expression style match the specific user profile. Alternatively, a personalized database for a specific group of people can be extracted from the basic database, such as encyclopedic knowledge text for children. In addition, to avoid rebuilding the personalized database every time an explanation is generated, the query prediction unit can pre-generate several personalized databases corresponding to high-frequency user profiles. When a high-frequency user calls the system, the query prediction unit directly matches the user profile with the personalized database and calls it immediately, eliminating the reconstruction step and significantly improving processing efficiency.

[0036] According to one embodiment of the present invention, the question prediction module uses a preset large language model (multimodal large language model) to filter introductory text segments that match the customized information in the personalized database, and generates multiple questions with the selected introductory text segments as answers. All generated questions are used to predict what the user might be interested in regarding a specified exhibit. Specifically, the multimodal large language model is used to convert the image and region annotations of the specified exhibit into text descriptions, and predicts the user's rough queries based on the user's prompts. Based on this, knowledge text segments related to the user profile are retrieved from the personalized database. Finally, the multimodal large language model is used to convert the knowledge text segments into a set of questions that the user is interested in. The prediction process can be represented as: Q=QuePred(I,C,T,R), where Q represents the user's questions of interest, I represents the image of the specified exhibit, C represents the user profile and scene type information, T represents the text instructions, and R represents the region annotations of the specified exhibit image by the user, i.e., the regions of interest specified by bounding boxes or points. The question prediction module adopts a strategy of generating questions based on a knowledge base (basic database), ensuring that the answers to all questions can be traced back to the knowledge base (basic database), thereby improving the accuracy of question answering.

[0037] IV. Exhibit Knowledge Retrieval Unit

[0038] The exhibit knowledge retrieval unit uses a preset knowledge retrieval method to retrieve multiple introductory text segments most relevant to a specified exhibit and all predicted questions from a basic database to form background knowledge for the specified exhibit. According to one embodiment of the invention, the exhibit knowledge retrieval unit employs a two-stage knowledge retrieval strategy: first, it filters introductory texts related to the specified exhibit from the basic database; then, it extracts the introductory text segments most relevant to the specified exhibit and all predicted questions from the filtered introductory texts to construct background knowledge for the specified exhibit. Here, an introductory text segment is a paragraph or a sentence of the introductory text. The exhibit knowledge retrieval unit obtains knowledge text segments related to questions from a personalized database as background knowledge for explanation generation, ensuring that the acquired knowledge is highly relevant to both the query image and the question, thereby effectively improving the accuracy of question answering. Compared to methods that generate content solely based on exhibition images, this method has significant advantages in the accuracy and richness of exhibit knowledge.

[0039] V. Explanation Generation Unit

[0040] The narration generation unit is used to generate narration text for a specified exhibit based on the narration customization information and background knowledge. The narration generation unit is equipped with a pre-trained narration generation model, which is used to generate narration text for a specified exhibit based on the narration customization information and background knowledge. The narration generation model is a multimodal large language model trained by taking the narration customization information and background knowledge as input, sequentially executing multiple preset reasoning tasks, and outputting the narration text of the specified exhibit.

[0041] According to an embodiment of the present invention, in the narration generation unit, the narration text generation model is constructed in the following manner: a training dataset is constructed, the training dataset including multiple labeled samples, wherein each sample includes an image of an exhibit, a question of interest to the user, and background knowledge of the exhibit constructed using a preset knowledge retrieval method, and the label is the narration text corresponding to the sample; a multimodal large language model (such as Gemini-2.5-Flash, GPT-4o) is obtained, and the multimodal large language model is trained and its parameters are fine-tuned based on the samples in the training dataset as input, reasoning using a preset thought chain, and the narration text corresponding to the sample as output, so as to obtain the narration text generation model, wherein the preset thought chain is composed of multiple task instructions and is used to guide the multimodal large language model to sequentially execute multiple preset reasoning tasks to generate the narration text of the exhibit.

[0042] According to an embodiment of the present invention, the preset reasoning tasks in the preset thought chain include an exhibit description generation task, a sub-question-answer generation task, and an explanation generation task. The image description generation task generates visual information about the exhibit related to questions of interest to the user based on an image of a specified exhibit. The sub-question-answer generation task decomposes the user's questions of interest into a series of logically progressive sub-questions and generates text answers for each sub-question based on background knowledge of the specified exhibit. The explanation generation task generates explanatory text for the specified exhibit based on the visual information and text answers. This invention relies on a thought chain reasoning mechanism to achieve a dual improvement in the fluency and interpretability of the explanatory text. By conducting targeted training on a multimodal large language model, it can generate explanatory content that matches the user's style, ensuring that the content is fluent, natural, and easy to understand. Addressing the common illusion problem in existing multimodal models when generating content, this solution embeds a thought chain in the model reasoning stage, prompting the model to simultaneously generate a focused description of the exhibit's local features and a multi-step reasoning process for problem solving while outputting explanatory text, thereby effectively enhancing the interpretability and credibility of the generated results.

[0043] According to one embodiment of the present invention, the preset knowledge retrieval method used in the exhibit knowledge retrieval unit and the text generation model includes coarse-grained retrieval and fine-grained retrieval. Coarse-grained retrieval involves selecting a first preset number of introductory texts with the highest similarity to the exhibit images from a basic database to form a candidate text set. Fine-grained retrieval involves dividing all introductory texts in the candidate text set into multiple introductory text segments, using the exhibit images and user-interested questions as joint information for the exhibit, selecting a second preset number of introductory text segments with the highest similarity to the joint information of the exhibit from all the divided introductory text segments, and combining all the selected introductory text segments into background knowledge.

[0044] According to one embodiment of the present invention, coarse-grained retrieval includes the following steps: encoding a specified exhibit image into a feature vector using a preset multimodal encoder; obtaining the cosine similarity between the feature vector of the specified exhibit image and the feature vector of each descriptive text to obtain the similarity between the specified exhibit image and each descriptive text; and selecting a first preset number of descriptive texts with the highest similarity to form a candidate text set. The method for calculating the cosine similarity is as follows: , To specify the feature vector of the exhibit image, Represents the first in the personalized database The feature vector of a knowledge text.

[0045] According to one embodiment of the present invention, fine-grained retrieval includes the following steps: dividing each introductory text in the candidate text set into multiple introductory text segments based on sentences or paragraphs of a preset length; acquiring a multimodal retrieval tool (such as a PreFLMR retrieval tool) and configuring it to calculate the similarity between each introductory text segment and the exhibit's joint information, and selecting a second preset number of introductory text segments with the highest similarity from all introductory text segments. All introductory text segments selected by the multimodal retrieval tool are then used to form background knowledge. This invention employs a two-stage knowledge retrieval method of coarse-grained and fine-grained retrieval, ensuring that the acquired knowledge is highly relevant to both the query image and the question, thereby effectively improving the accuracy of question answering.

[0046] According to one embodiment of the present invention, the present invention proposes a method for generating customized explanations of exhibit images for exhibition scenarios. The method includes: constructing customized explanation information for a specified exhibit based on user needs; generating explanatory text for the specified exhibit based on the customized explanation information using the customized explanation system proposed in this invention; and providing an explanation of the specified exhibit based on the generated explanatory text.

[0047] To facilitate understanding of the process of generating explanatory text using the customized explanatory system proposed in this invention, the following description will be provided in conjunction with the accompanying drawings and examples.

[0048] Referring to Figure 2, the process of generating explanatory text can be summarized into three steps: multimodal interaction and customized demand prediction, exhibit knowledge retrieval enhancement, and explanation generation based on thought chain reasoning. Each step will be explained in detail below.

[0049] The steps for multimodal interaction and customized demand prediction are as follows: First, obtain the customized explanation information input by the user and then perform customized user query prediction based on the customized explanation information and the basic database to predict the questions the user is interested in. Referring again to Appendix 2, the customized explanation information input by the user includes a labeled image of an exhibit, the exhibition theme (art), and a user profile (student). This step constructs a personalized database adapted to the user profile based on the user's customized information and the basic database. Then, customized user query prediction is performed based on this personalized database to obtain the questions the user is interested in, namely, "What does the man on the mountain symbolize?"

[0050] The steps to enhance exhibit knowledge retrieval are as follows: a two-stage multimodal retrieval is used to construct background knowledge of the exhibits. Referring to Figure 2, the two-stage multimodal retrieval includes coarse-grained retrieval and fine-grained retrieval. The coarse-grained retrieval filters the document set from the basic database based on the user's customized information to obtain multiple knowledge documents (knowledge texts) related to the exhibits. The fine-grained retrieval filters out knowledge paragraphs or sentences that match the user's customized information from the knowledge documents obtained from the coarse-grained retrieval to form the background knowledge of the exhibits, such as "It is an oil painting on canvas created by the German Romantic painter Ka in 1817, depicting a person with their back to the viewer...".

[0051] The explanation generation steps based on the thought chain reasoning are as follows: A multimodal large language model (explanation generation unit) is used as input, taking customized information, customized query results, and background knowledge of the exhibits as input, to generate explanatory text for the exhibits. The multimodal large language model is configured with a thought chain to guide the reasoning process. Referring to Figure 2, the thought chain includes three tasks executed sequentially: image description generation, sub-question answer generation, and explanatory text generation. The results of these three tasks are shown in Figure 2.

[0052] (1) Image description generation task

[0053] Task instruction: What information about the problem can be gleaned from the image?

[0054] Mission Result: A man in a dark long coat stands on a rocky cliff, gazing at the vast and hazy landscape.

[0055] (2) Sub-question-answer generation task

[0056] Task instructions: Break down the user's questions into a series of logically progressive sub-questions and generate text answers for each sub-question based on the background knowledge of the specified exhibits.

[0057] Sub-question 1: What is this painting?

[0058] Answer: Ka's "The Traveler on the Sea of ​​Fog".

[0059] Sub-question 2: Who is the man on the mountain?

[0060] Answer: The man's identity is uncertain; some analysts believe it may be a self-portrait of the author.

[0061] Sub-question 3: What do men symbolize?

[0062] Answer: He symbolizes the Romantic art's profound focus on nature, the individual, and the spiritual world.

[0063] (3) Explain the text generation task

[0064] Task Instructions: Summarize the above information (i.e., the execution results of the image description generation task and the sub-question-answer generation task) and generate an explanatory text for student visitors.

[0065] Task Result: The painting you uploaded is "Wanderer on a Sea of ​​Fog," created by [artist's name]. The man in the painting stands on a rocky cliff, gazing at the vast landscape, exuding an air of mystery. His identity is uncertain; he may represent the artist himself. His image reflects the artist's inner world and symbolizes the profound focus on nature, the individual, and the spiritual world in Romantic art.

[0066] In summary, the customized explanation system and method for exhibition scenarios proposed in this invention have the following three advantages: (1) High degree of customization in generating explanations. It supports free interaction between users in a multimodal manner, including text and images, and allows users to upload their own photographs of exhibits and raise questions or areas of interest in the form of language or markings on the images. In addition, personalized knowledge data is constructed by combining exhibit knowledge and user profiles. Based on this, questions of interest to users are accurately predicted for users to choose from, and explanation content suitable for the user style is generated. (2) Enhanced accuracy and richness of knowledge through knowledge retrieval. A two-stage multimodal knowledge retrieval method is designed to obtain knowledge related to questions from a preset basic database as a reference for explanation generation, ensuring that the knowledge is accurate and can be extended to other related knowledge points. Compared with the method of generating content based solely on exhibition images, the accuracy and richness of exhibit knowledge are better. (3) Improved fluency and interpretability of explanation text based on thought chain reasoning. Directly using the retrieved content as explanation content is too rigid. Training a visual-language model to generate explanation content suitable for the user style is more fluent, natural, and easier for users to understand. Existing multimodal models are prone to illusion problems when generating content. Introducing thought chains into the model's reasoning process allows the model to not only generate explanatory content, but also generate local content of exhibits it is interested in and multi-step reasoning processes for answering questions, which can improve the interpretability and credibility of the model's generated results.

[0067] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.

[0068] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.

[0069] Computer-readable storage media can be tangible devices that hold and store instructions for use by an instruction execution device. Computer-readable storage media can include, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof.

[0070] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A customized text description generation system for exhibit images in exhibition scenarios, wherein the system provides personalized text description services for exhibits in a target exhibition scenario, characterized in that, The system includes an interaction unit, a basic database, a query prediction unit, an exhibit knowledge retrieval unit, and an explanation generation unit. The interaction unit receives customized explanation information provided by the user, which includes at least a user profile and images of specified exhibits. The basic database includes images of all exhibits in the target exhibition scene, introductory text for each exhibit, and images of each exhibit. The query prediction unit is configured with a personalized database construction module and a question prediction module. The personalized database construction module is used to select multiple introductory texts matching the user profile from the basic database to form a personalized database. The question prediction module is used to generate explanations based on the personalized database and the user profile. The system includes several components: a customized information prediction unit that identifies multiple questions users might be interested in regarding a specific exhibit; an exhibit knowledge retrieval unit that uses a pre-defined knowledge retrieval method to retrieve multiple introductory text segments most relevant to the specified exhibit and all predicted questions from a basic database to form background knowledge for the specified exhibit, where each introductory text segment is a paragraph or sentence; and an explanation generation unit equipped with a pre-trained explanation generation model that generates explanatory text for the specified exhibit based on the customized information and background knowledge. The explanation generation model is a multimodal large language model trained by taking the customized information and background knowledge as input, sequentially executing multiple pre-defined reasoning tasks, and outputting the explanatory text for the specified exhibit.

2. The system according to claim 1, characterized in that, The personalized database construction module can also construct a personalized database in the following way: using a preset large language model, all the introductory texts in the basic database are rewritten into texts that match the expression style of the user profile, and the rewritten texts are combined into a personalized database.

3. The system according to claim 1, characterized in that, The question prediction module is configured to: use a preset large language model to filter introductory text segments from a personalized database that match the customized information being explained, generate multiple questions with the selected introductory text segments as answers, and use all generated questions as predictions of what users might be interested in regarding the specified exhibits.

4. The system according to claim 1, characterized in that, In the explanation generation unit, the explanation text generation model is constructed as follows: a training dataset is constructed, which includes multiple labeled samples. Each sample includes an image of an exhibit, a question of interest to the user, and background knowledge of the exhibit constructed using a preset knowledge retrieval method. The label is the explanation text corresponding to the sample. A multimodal large language model is obtained and trained based on the samples in the training dataset as input, reasoning using a preset thought chain, and the explanation text corresponding to the sample as output, to obtain the explanation text generation model. The preset thought chain consists of multiple task instructions and is used to guide the multimodal large language model to sequentially execute multiple preset reasoning tasks to generate explanation text for the exhibit.

5. The system according to claim 1 or 4, characterized in that, The preset knowledge retrieval enhancement method is as follows: coarse-grained retrieval, which selects a first preset number of introductory texts with the highest similarity to the images of the exhibits from the basic database to form a candidate text set; fine-grained retrieval, which divides all the introductory texts in the candidate text set into multiple introductory text segments, using the images of the exhibits and questions of interest to users as the joint information of the exhibits, and selects a second preset number of introductory text segments with the highest similarity to the joint information of the exhibits from all the divided introductory text segments, and then combines all the selected introductory text segments into background knowledge.

6. The system according to claim 5, characterized in that, The coarse-grained retrieval is as follows: a preset multimodal encoder is used to encode the specified exhibit image into a feature vector; the cosine similarity between the feature vector of the specified exhibit image and the feature vector of each introductory text is obtained to obtain the similarity between the specified exhibit image and each introductory text; a first preset number of introductory texts with the highest similarity are selected to form a candidate text set.

7. The system according to claim 5, characterized in that, The fine-grained retrieval is as follows: each introductory text in the candidate text set is divided into multiple introductory text segments based on sentences or paragraphs of a preset length; a multimodal retrieval tool is obtained and configured to calculate the similarity between each introductory text segment and the joint information of the exhibits, and to select the second preset number of introductory text segments with the highest similarity from all introductory text segments; all introductory text segments selected by the multimodal retrieval tool are combined to form background knowledge.

8. The system according to claim 1 or 4, characterized in that, The multiple preset reasoning tasks include an exhibit description generation task, a sub-question-answer generation task, and an explanatory text generation task, wherein: the image description generation task is to generate visual information of the exhibit related to questions of interest to the user based on the image of the specified exhibit; the sub-question-answer generation task is to break down the questions of interest to the user into a series of logically progressive sub-questions and generate text answers for each sub-question based on the background knowledge of the specified exhibit; and the explanatory text generation task is to generate explanatory text for the specified exhibit based on the visual information of the exhibit and the text answers.

9. A method for generating customized explanations of exhibit images for exhibition scenarios, characterized in that, The method includes: constructing customized explanation information for a specified exhibit based on user needs; generating explanatory text for the specified exhibit based on the customized explanation information using a customized explanation system as described in any one of claims 1-8; and providing an explanation of the specified exhibit based on the generated explanatory text.

10. A computer-readable storage medium, characterized in that, It stores a computer program that can be executed by a processor to implement the steps of the method of claim 9.

11. An electronic device, characterized in that, include: One or more processors; and a memory, wherein the memory is used to store executable instructions; the one or more processors are configured to implement the steps of the method of claim 9 by executing the executable instructions.