Intelligent exhibit identification and explanation method and system for exhibition scene
By collecting user data in real time to identify exhibits and analyze interests, the system can dynamically adjust visitor paths and explanations, solving the problem of monotonous exhibit identification and explanation, and achieving personalization and interactivity to enhance the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 江苏遨信科技有限公司
- Filing Date
- 2026-05-07
- Publication Date
- 2026-06-02
Smart Images

Figure CN122132544A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of exhibit recognition and explanation technology, and more specifically to intelligent exhibit recognition and explanation methods and systems for exhibition scenarios. Background Technology
[0002] Currently, many exhibition halls have introduced digital tour guide systems, such as audio guides, QR code scanning guides, and location-triggered automatic guide apps, which to some extent replace traditional manual guides and provide basic information services for visitors.
[0003] Existing technologies suffer from the following problems: the exhibit recognition trigger mechanism is simplistic, lacking accurate judgment of the user's true intentions, relying solely on spatial distance as the sole criterion, and failing to distinguish whether the user is merely passing by, making a brief stop, or genuinely showing interest; the visitor route is fixed and rigid, unable to meet the user's personalized needs, and unable to be dynamically adjusted according to the user's real-time interests and preferences; the explanation content structure is fixed, unable to be integrated with the user's dynamic visitor path and personalized interaction results, and the use of a standard order of explanations makes it impossible to connect and integrate the answers to user questions with the subsequent explanations of related exhibits, thus limiting the user experience; to solve at least one of the above problems, this application proposes an intelligent exhibit recognition and explanation method and system for exhibition scenarios. Summary of the Invention
[0004] To address the shortcomings of existing technologies, the purpose of this application is to provide an intelligent exhibit recognition and explanation method and system for exhibition scenarios, which can effectively solve the problems in the background technology. The specific technical solution of this application is as follows:
[0005] Intelligent exhibit recognition and explanation methods in exhibition settings include:
[0006] Based on real-time collected user terminal perception data, the exhibits that the user is currently interested in are identified, exhibit icons are obtained, the dwell time of the user at the corresponding exhibit icon location is analyzed, the corresponding first explanation content is provided and the explanation begins.
[0007] Based on real-time collected user behavior data, user interests are analyzed. Combined with the relevance of exhibits and real-time environmental data of the exhibition hall, user fatigue is analyzed and time constraints are constructed according to the user's visit duration and movement. The visit path is dynamically adjusted to obtain the first visit path.
[0008] During the explanation, responding to user voice questions, analyzing the intent of the question and determining the corresponding exhibits, and generating answer content based on user interests;
[0009] By combining the first tour route and the answers provided, and by adjusting the order of the explanations and increasing the coherence of the answers, the first explanation was optimized to obtain the second explanation.
[0010] Specifically, based on real-time collected user terminal perception data, the exhibits currently of interest to the user are identified, exhibit identifiers are obtained, the time the user spends at the corresponding exhibit identifier location is analyzed, and corresponding initial explanation content is provided and the explanation begins, including:
[0011] Based on real-time collected user terminal perception data, the exhibit area is detected and the duration of user stay in the exhibit area is calculated to identify the exhibits that the user is currently interested in and obtain exhibit identification.
[0012] Analyze the time users spend at the corresponding exhibit markers, filter out exhibits whose dwell time exceeds a preset dwell threshold, and generate an explanation trigger signal;
[0013] In response to the explanation trigger signal, the corresponding first explanation content is retrieved from the pre-built exhibit explanation knowledge base, and the explanation begins.
[0014] Specifically, the step of detecting and calculating the duration of user stay in the exhibition area based on real-time collected user terminal perception data, identifying the exhibits currently of interest to the user, and obtaining exhibit identifiers includes:
[0015] Acquire real-time user terminal perception data, which includes real-time image sequences, voice signal segments, location coordinate sequences, and motion trajectory data;
[0016] Identify exhibit regions in the real-time image sequence, analyze candidate exhibits in each frame of the image, and obtain a first exhibit set;
[0017] Filter out the exhibit feature keywords in the speech signal segments, determine the exhibits corresponding to the exhibit feature keywords, and obtain the second set of exhibits;
[0018] Based on the location coordinate sequence and motion trajectory data, calculate the duration of user stay in the exhibition area and extract corresponding behavioral features;
[0019] The intersection of exhibits in the first and second exhibit sets is selected, and the attention score of each exhibit in the intersection is calculated based on the behavioral characteristics. The exhibit with the highest attention score is then identified as the exhibit.
[0020] Specifically, based on real-time collected user behavior data, user interests are analyzed. Combined with exhibit relevance and real-time exhibition hall environment data, user fatigue is analyzed and time constraints are constructed based on visit duration and movement patterns. The visit path is then dynamically adjusted to obtain a first visit path, including:
[0021] Based on real-time collected user behavior data, we extract the attention weights of users to different exhibit attributes, analyze user interests, and obtain interest vectors.
[0022] Based on the interest vector, combined with the relevance of exhibits and real-time environmental data of the exhibition hall, the visitor path is dynamically adjusted to obtain the first visitor path.
[0023] Specifically, the first visitor path is obtained by dynamically adjusting the visitor path based on the interest vector, combined with the relevance of exhibits and real-time environmental data of the exhibition hall, including:
[0024] Extract exhibit attribute information and association information between exhibits from the pre-constructed exhibit knowledge graph, calculate the association degree between exhibits, and construct an exhibit association matrix;
[0025] Based on user visit duration and movement patterns, analyze user fatigue levels and construct time constraints;
[0026] By combining the interest vector, exhibit association matrix, time constraints, and real-time environmental data of the exhibition hall, the visitor path is dynamically adjusted to obtain the first visitor path.
[0027] Specifically, during the explanation process, in response to user voice questions, the system analyzes the intent of the question and identifies the corresponding exhibits, and generates answers based on user interests, including:
[0028] During the explanation, in response to user voice questions, the voice signal is converted into text form based on the real-time collected user voice signal to obtain the question text;
[0029] Based on the analysis of the question text, the intent of the question is determined and the corresponding exhibits are identified. Then, the answer content is generated in combination with the user's interests.
[0030] Specifically, the step of analyzing the question text to determine the questioner's intent and the corresponding exhibits, and generating answer content based on user interests, includes:
[0031] Extract the intent category and entity information from the question text. The intent category includes inquiries about the era, craftsmanship, and related exhibits. The entity information includes the exhibit name, number, and exhibit characteristics.
[0032] The entity information is matched with the exhibit identifier currently being explained to determine the exhibit identifier corresponding to the question information;
[0033] Based on the intent category and the exhibit identifier in question, the corresponding exhibit knowledge is filtered from the pre-constructed exhibit knowledge graph to obtain the first knowledge set;
[0034] Calculate the similarity between exhibit knowledge and user interests in the first knowledge set, filter out exhibit knowledge with similarity greater than a preset similarity threshold to obtain a second knowledge set, and integrate the exhibit knowledge in the second knowledge set to generate answer content.
[0035] Specifically, by combining the first tour route and the answers provided, and by adjusting the order of the explanations and increasing the coherence of the inserted answers, the first explanation is optimized to obtain the second explanation, which includes:
[0036] Obtain the first visit path and the answer content. The first visit path includes a list of exhibit labels arranged in the order of visit and the corresponding dwell time and explanation time allocation for each exhibit.
[0037] The first part of the explanation is divided according to the exhibit labels, and the explanation segments corresponding to each exhibit are extracted to obtain a set of exhibit explanation segments;
[0038] Based on the first visitor route and the answers provided, the order of the exhibit explanation segments is adjusted and the continuity of the inserted answers is increased to obtain the second explanation content.
[0039] Specifically, based on the first visitor path and the answers provided, the order of the exhibit explanation segments is adjusted and the continuity of the inserted answers is increased to obtain the second explanation content, including:
[0040] Based on the order of exhibits in the first visitor route, the set of exhibit explanation segments is sorted to obtain an updated set of exhibit explanation segments;
[0041] Analyze the exhibit identifier corresponding to the answer content, locate the corresponding explanation segment in the updated exhibit explanation segment set, and match it with the explanation segment to determine the insertion position of the answer content;
[0042] According to the insertion position, the answer content is inserted into the corresponding explanation segment, and the coherence between the answer content and the explanation segment is analyzed. The explanation content is then optimized to obtain the second explanation content.
[0043] An intelligent exhibit recognition and explanation system for exhibition scenarios, used to implement the intelligent exhibit recognition and explanation method for the aforementioned exhibition scenarios, including:
[0044] The exhibit identification module identifies the exhibits that the user is currently interested in based on real-time user terminal perception data, obtains exhibit identifiers, analyzes the time the user spends at the corresponding exhibit identifier location, provides the corresponding initial explanation content, and begins the explanation.
[0045] The user interest analysis module analyzes user interests based on real-time collected user behavior data. It combines exhibit relevance and real-time exhibition hall environment data to analyze user fatigue and construct time constraints based on user visit duration and movement. It then dynamically adjusts the visit path to obtain the first visit path.
[0046] The question-and-answer module responds to user voice questions during the explanation process, analyzes the intent of the question, identifies the corresponding exhibits, and generates answers based on user interests.
[0047] The explanation optimization module combines the first tour route and the answers provided. By adjusting the order of the explanation content and increasing the coherence of the inserted answers, the first explanation content is optimized to obtain the second explanation content.
[0048] The beneficial effects of this application are as follows: By fusing and analyzing real-time image, voice, location, and motion trajectory data, the exhibits that users are interested in can be identified, enabling accurate triggering of explanations and avoiding accidental or missed triggers, thereby improving user comfort and the effective reach of information; by analyzing users' interest vectors and combining exhibit relevance with real-time exhibition hall environmental data for path planning, the efficiency of viewing and user experience can be improved by combining interest matching, content relevance, and environmental comfort; by recognizing user questions and matching and filtering knowledge sets with users' interest vectors, the resulting explanation content can match users' knowledge background and interests, enhancing the user's experience during the exhibition visit; and by optimizing the semantic coherence of answers according to the first viewing path, the completeness and logical coherence of the explanation content can be improved, enhancing the user's immersion and experience during the exhibition visit. Attached Figure Description
[0049] Figure 1 This is a flowchart illustrating the intelligent exhibit recognition and explanation method in the exhibition scene of this application embodiment;
[0050] Figure 2 This is a flowchart illustrating the process of generating answer content in an embodiment of this application.
[0051] Figure 3 This is a schematic diagram of the structure of the intelligent exhibit recognition and explanation system for an exhibition scene in this application embodiment. Detailed Implementation
[0052] The present application will be further described in detail below with reference to the accompanying drawings and embodiments.
[0053] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0054] Hereinafter, the terms "first," "second," and other generic terms are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0055] refer to Figure 1 The image shows a specific implementation method for the intelligent exhibit recognition and explanation method in the exhibition scenario of this application, including:
[0056] S101. Based on the real-time collected user terminal perception data, identify the exhibits that the user is currently interested in, obtain the exhibit identifier, analyze the time the user stays at the corresponding exhibit identifier location, provide the corresponding first explanation content, and begin the explanation.
[0057] S102. Based on real-time collected user behavior data, analyze user interests, combine exhibit relevance and real-time exhibition hall environment data, analyze user fatigue and construct time constraints according to user visit duration and movement, dynamically adjust the visit path, and obtain the first visit path.
[0058] S103. During the explanation, respond to the user's voice questions, analyze the intention of the question and determine the corresponding exhibits, and generate the answer content based on the user's interests;
[0059] S104. Combining the first tour route and the answers, the first explanation content is optimized by adjusting the order of the explanation content and increasing the coherence of the answer content insertion, resulting in the second explanation content.
[0060] In this embodiment, a smartphone or smart guide terminal is used to capture continuous images from the previous few seconds as a real-time image sequence, ambient sound as a speech signal segment, location coordinate sequence, and motion trajectory data. During data processing, a convolutional neural network object detection model trained and optimized with a large dataset of exhibit images is used. The model's training dataset consists of images of multiple exhibits taken from multiple angles, with manually labeled bounding boxes and categories. During training, a stochastic gradient descent optimization algorithm is employed, with an initial learning rate of 0.01, a momentum factor of 0.9, a weight decay of 0.0005, and a loss function that comprehensively considers classification error and bounding box regression error (cross-over ratio loss). The model converges after 200 training epochs. For each input image frame, the convolutional neural network object detection model outputs multiple candidate boxes along with their corresponding exhibit categories and confidence scores, resulting in the first exhibit set.
[0061] On the other hand, the speech signal is converted into text using an acoustic model based on a deep feedforward sequence memory network and a large language model based on a weighted finite state converter. A natural language understanding module, finely tuned from a large corpus of museum explanatory texts and user question-and-answer dialogues, extracts feature keywords related to the exhibits from the text, and obtains a second set of exhibits through keyword matching. Based on location coordinate sequences and motion trajectory data, a Kalman filter is used to smooth and predict user trajectories, calculating the duration of continuous stay within the preset exhibit area's electronic fence, and extracting the rate of change of body orientation and foot speed as behavioral features.
[0062] Preferably, exhibits that appear in both the first and second exhibit sets are selected and included in the candidate intersection. For each exhibit in the candidate intersection, an attention score is calculated by combining the confidence score of image recognition, the voice matching score, and the attention weight calculated based on dwell time and body orientation. The exhibit with the highest attention score is identified as the exhibit identifier, and the corresponding standard explanation content is retrieved from the pre-built exhibit explanation knowledge base as the first explanation content to be played.
[0063] It should be noted that by using multimodal data fusion, we can combine the user's visual focus, verbal content, and spatial behavior to identify the user's attention behavior, solve the problems of false triggering and missed triggering caused by single positioning or timed triggering mechanisms, ensure that the start of the explanation matches the user's attention intention, improve the user's interactive experience, and provide an accurate data foundation for the subsequent analysis process.
[0064] Specifically, the system collects and records user behavior data in real time, including attention scores for each exhibit, the number of times different sections of the exhibit's explanation are listened to, zoom-in views of exhibit detail images, and the intent categories involved in user questions via voice. The behavioral data is processed using a matrix factorization model based on collaborative filtering. The training dataset for this model includes: historical exhibition viewing data, including historical user behavior records on various exhibits; and real-time behavioral data generated by the current user during their current visit.
[0065] Preferably, the matrix factorization model takes user ID and exhibit ID as input and outputs the user's predicted rating for the exhibit. It is trained using a stochastic gradient descent optimization algorithm with a learning rate of 0.001 and a regularization parameter of 0.01, along with a root mean square error loss function. The model maps user interests to a multi-dimensional interest vector, where each dimension represents the strength of the user's preference for exhibit attributes. Relationship information between exhibits is extracted from a pre-constructed exhibit knowledge graph. The shortest path length and relationship type between any two exhibit nodes are calculated to obtain the correlation degree between exhibits, and an exhibit association matrix is constructed. Real-time environmental data of the exhibition hall, including pedestrian density, temperature, and humidity in each area, is collected to plan the user's visit path, including a list of exhibit identifiers and the corresponding dwell time and explanation time allocation for each exhibit, serving as the first visit path.
[0066] It should be noted that by integrating user interests, exhibit relevance, and real-time environment, the planned visitor path includes exhibits that users are interested in, and can proactively avoid congestion, thereby improving visitor efficiency, comfort, and the accuracy of knowledge acquisition.
[0067] While the system is playing a narration, users can interrupt and ask questions at any time. After the terminal's microphone captures the speech signal, it uses an end-to-end automatic speech recognition model based on connectionist temporal classification. This model is trained using a speech corpus containing various accents and museum ambient noise, employs a stochastic gradient descent optimizer, has an initial learning rate of 0.01, and uses connectionist temporal classification loss as the loss function. The model converts the speech into text to obtain the question text.
[0068] Preferably, a joint model for intent recognition and entity extraction based on BERT is invoked. This model is fine-tuned on a corpus of exhibition hall question-and-answer dialogues containing multiple labeled intent categories and entity information. The AdamW optimizer is used for fine-tuning, with a learning rate set to 3×10⁻⁶. -5 The batch size is 16, and the loss function is a weighted sum of cross-entropy loss for intent classification and conditional random field loss for entity recognition. The model's input is the question text, and the output is the predicted intent category and entity information extracted from the text. The extracted entity information is matched with the currently explained exhibit identifier to determine the exhibit identifier to which the user's question is directed.
[0069] Furthermore, by combining the intent category and the exhibit identifier in the question, all relevant knowledge points are retrieved to obtain the first knowledge set. Each knowledge point in the first knowledge set is mapped to a semantic vector, and cosine similarity is calculated between this vector and the user's interest vector. A preset similarity threshold is set, and knowledge points with similarity scores higher than the threshold are retained to obtain the second knowledge set. By combining user interests and preferences to filter information that matches individual needs, user engagement and cognitive depth can be enhanced.
[0070] Specifically, combining the first visitor path and the answers provided, the first explanation content is processed. The complete explanation of each exhibit is divided into explanation segments, resulting in a set of explanation segments. Following the exhibit order in the first visitor path, all segments in the explanation segment set are sorted to obtain an updated set of explanation segments. All answers are iterated through, and based on the associated exhibit identifier and intent category, the corresponding explanation segment is located in the updated explanation segment set. The corresponding insertion point is determined, and the corresponding explanation content is inserted at each insertion point to obtain the second explanation content. By integrating the user's visitor path and Q&A content, the relevance and coherence of the second explanation content can be improved, transforming the user's viewing experience from passively receiving information to actively participating, thus enhancing the user experience.
[0071] This application integrates and analyzes real-time image, voice, location, and motion trajectory data to identify exhibits that users are interested in, accurately triggering explanations and avoiding accidental or missed triggers, thus improving user comfort and effective information delivery. It analyzes user interest vectors and combines exhibit relevance with real-time exhibition hall environmental data for path planning, combining interest matching, content relevance, and environmental comfort to improve viewing efficiency and user experience. It identifies user questions by matching and filtering knowledge sets with user interest vectors, resulting in explanations that match user knowledge background and interests, enhancing the user's experience during the exhibition visit. Finally, it optimizes the semantic coherence of responses according to the first viewing path, improving the completeness and logical coherence of the explanations, and enhancing the user's immersion and experience during the exhibition.
[0072] Furthermore, based on real-time collected user terminal perception data, the exhibits currently of interest to the user are identified, exhibit identifiers are obtained, the dwell time of the user at the corresponding exhibit identifier location is analyzed, and corresponding initial explanatory content is provided and the explanation begins, including:
[0073] S201. Based on real-time collected user terminal perception data, detect the exhibit area and calculate the duration of user stay in the exhibit area, identify the exhibit currently being viewed by the user, and obtain exhibit identification.
[0074] S202. Analyze the time users spend at the corresponding exhibit markers, filter out exhibits whose dwell time exceeds a preset dwell threshold, and generate an explanation trigger signal.
[0075] S203. In response to the explanation trigger signal, retrieve the corresponding first explanation content from the pre-built exhibit explanation knowledge base and start the explanation.
[0076] In this embodiment, based on real-time user terminal perception data, the exhibit area is detected and the duration of user stay in the exhibit area is calculated to identify the exhibit currently of interest to the user, thus obtaining exhibit identification. By fusing multimodal data, combining user visual focus, verbal content, and spatial behavior information, the accuracy and robustness of identifying user attention points in complex exhibition environments are improved. This provides accurate data support for subsequent explanation services, ensuring that the system can accurately understand the user's intent, avoiding information interference or service misalignment caused by identification errors, and improving the user experience.
[0077] Specifically, after identifying exhibit identifiers, the system determines the associated dwell time for each identifier and constructs a preset dwell time threshold library based on exhibit type, complexity, and estimated viewing time. For example, for relatively simple and easy-to-understand two-dimensional calligraphy and painting exhibits, the set dwell time threshold is 8 seconds. The system compares the current cumulative dwell time corresponding to an exhibit identifier with the corresponding threshold in the preset dwell time threshold library. Once the cumulative dwell time of an exhibit identifier exceeds the preset dwell time threshold, the system immediately triggers a narration trigger signal. This signal includes the corresponding exhibit identifier, the trigger timestamp, and the threshold type on which the trigger is based.
[0078] It should be noted that by setting the cumulative dwell time, the system can intelligently filter out users' unconscious short pauses and trigger the explanation service when users show interest. This can effectively avoid accidental triggering caused by users' unintentional wandering or environmental interference, ensuring that the explanation is started in sync with the user's understanding needs, thus improving the system's intelligence and user experience.
[0079] Specifically, the system maintains an exhibit explanation knowledge base, using the exhibit's unique identifier as the primary key to associate and store explanation content in various formats and versions for that exhibit, serving as the primary explanation content. Upon receiving an explanation trigger signal, the system parses the exhibit identifier and quickly locates and retrieves all pre-associated explanation content from the exhibit explanation knowledge base. Based on the user's terminal type and current network status, the system selects the appropriate explanation content version for delivery. By building this exhibit explanation knowledge base, the system can respond instantly to trigger signals and push corresponding content, improving service timeliness and accuracy, and enhancing the efficiency and quality of information delivery.
[0080] Furthermore, based on real-time collected user terminal perception data, the system detects and calculates the duration of user stay in the exhibition area, identifies the exhibits currently of interest to the user, and obtains exhibit identifiers, including:
[0081] S301. Acquire real-time user terminal perception data, wherein the user terminal perception data includes real-time image sequences, voice signal segments, location coordinate sequences, and motion trajectory data;
[0082] S302. Identify the exhibit areas in the real-time image sequence, analyze the candidate exhibits in each frame of the image, and obtain the first exhibit set;
[0083] S303. Filter out the exhibit feature keywords in the speech signal segments, determine the exhibits corresponding to the exhibit feature keywords, and obtain the second set of exhibits;
[0084] S304. Calculate the duration of user stay in the exhibition area based on the location coordinate sequence and motion trajectory data, and extract the corresponding behavioral features;
[0085] S305. Filter out the intersection of exhibits in the first exhibit set and the second exhibit set, and calculate the attention score of each exhibit in the exhibit intersection in combination with the behavioral features, and take the exhibit with the highest attention score as the exhibit identifier.
[0086] In this embodiment, the system captures a real-time image sequence by calling the preview data stream from the terminal's rear camera at a rate of 25 to 30 frames per second, showing the scene in front of the user. It also obtains an audio signal segment by continuously recording ambient sound using the microphone at a sampling rate of 16 kHz. Furthermore, it obtains a position coordinate sequence by calling the terminal's positioning module to acquire the user's high-precision latitude and longitude coordinates within the indoor space at a frequency of 10 times per second, and obtains motion trajectory data by reading data from the terminal's built-in inertial measurement unit. By collecting multimodal user terminal perception data, the system provides rich data information for subsequent user intent recognition, improving the accuracy of the analysis process when facing complex scenarios such as changing lighting, noisy environments, and positioning drift.
[0087] Specifically, a convolutional neural network (CNN) object detection model is configured. The training process includes: constructing a dataset of exhibit images, including images of exhibits collected from various exhibition halls, such as paintings, calligraphy, ceramics, bronzes, and jade. Each image is precisely annotated by professional annotators, marking the bounding box of the exhibit and its corresponding identification code. During training, a stochastic gradient descent optimizer is used, with an initial learning rate of 0.01, a momentum factor of 0.9, and a weight decay of 0.0005. The model's loss function includes cross-entropy loss for evaluating classification accuracy and smoothing L1 loss for evaluating bounding box regression accuracy. The input to the CNN object detection model is each frame of the image, and the output is a list of all detected candidate exhibits in that frame. Each candidate exhibit includes its identification code and the model's confidence score. Integrating the detection results from multiple consecutive frames yields the first set of exhibits. The object detection model can quickly filter candidate exhibits from the user's field of vision. The structure of the CNN object detection model includes a backbone network, a neck network, and a head network. The backbone network employs structures such as depthwise separable convolutions or cross-stage local networks to extract multi-scale feature maps from each input frame. The neck network fuses feature maps from different levels of the backbone network through connection methods such as feature pyramids or path aggregation networks to enhance the detection capability for small and multi-scale targets. The head network, based on the fused feature maps, predicts the target category confidence and bounding box position offset of each candidate region through a series of convolutional and fully connected layers. The model internally achieves feature reuse and fusion through operations such as skip connections, upsampling, and element-wise addition, and finally outputs the exhibit identification code and confidence score of each candidate exhibit.
[0088] Specifically, a cascaded model consisting of a speech recognition module and a semantic understanding module is invoked. The speech recognition module employs an acoustic model based on a deep feedforward sequential memory network and a large language model based on a weighted finite state converter. The input to the acoustic model is the spectral features of the speech signal segment, and the output is the posterior probability of the phoneme state. The large language model is responsible for decoding the phoneme sequence into a text sequence. The large language model is trained on a large speech dataset, using connectionist temporal classification as the loss function and a stochastic gradient descent optimizer during training, with an initial learning rate set to 0.01. The speech signal is then converted into a text-based question text. The semantic understanding module is a pre-trained language model based on BERT. The input is the question text, and the output is all the featured keywords of the exhibits identified in the text and their linked exhibit identification codes. All extracted exhibit identification codes are aggregated and deduplicated to obtain a second set of exhibits.
[0089] The acoustic model is built upon a Deep Feedforward Sequential Memory Network (DFSMN). DFSMN captures the long-term dependencies of speech signals by stacking multiple memory blocks. Each memory block contains forward and backward linear projection layers, and information flow is enhanced through skip connections and delayed connections. The model's input is the spectral features of speech signal segments, and its output is a phoneme sequence. The large-scale language model is represented using a Weighted Finite State Transformer (WFST), consisting of a series of states and weighted transition arcs. It integrates a pronunciation dictionary, a language model, and acoustic context constraints. The WFST structure fuses multiple subgraphs into an optimized search graph through combination operations. The input is the phoneme sequence output by the acoustic model, which, after passing through a Viterbi decoding algorithm, outputs the most probable text sequence.
[0090] It should be noted that by analyzing user voice in real time, it is possible to identify the points of interest that users actively express or reveal in conversations with peers, thereby enhancing the accuracy of judging user attention and improving the comprehensiveness and accuracy of understanding user intent.
[0091] Specifically, based on the position coordinate sequence, a smoothed trajectory is obtained through a Kalman filter. This smoothed trajectory is then overlaid and analyzed with a pre-stored electronic fence map of the exhibit area. The electronic fences, pre-defined by the exhibition organizers, are polygonal areas bound to a specific exhibit identifier, including the exhibit's physical location and corresponding viewing area. The system determines if each user's location falls within the electronic fence of the corresponding exhibit and records the duration of their stay at that location. By combining instantaneous velocity, direction angle, and other information from the motion trajectory data, behavioral characteristics of the user during their stay are extracted. Analyzing the stay duration and behavioral characteristics quantifies the intensity of interest and the user's level of focus, enabling accurate analysis of user attention levels.
[0092] Specifically, the intersection of the first and second exhibit sets is calculated. For each exhibit in the intersection, an attention score is calculated. For exhibit P in the intersection, the average confidence score, denoted as V(P), is obtained by averaging the confidence scores of exhibit P detected in multiple consecutive frames. The strength of the speech matching is analyzed and denoted as A(P), a Boolean value. If exhibit P is explicitly mentioned by the user's speech, this value is 1; otherwise, it is 0. The behavioral feature value, denoted as B(P), is calculated based on the dwell time and body orientation. V(P), A(P), and B(P) are then weighted and summed to obtain the attention score. The exhibit with the highest attention score is selected as the exhibit identifier. By integrating visual, auditory, and spatial behavioral information to calculate the attention score, the robustness and accuracy of the system in complex environments can be improved. The obtained exhibit identifiers provide accurate data support for subsequent analysis.
[0093] When calculating the behavioral characteristic value B(P), it is quantified based on the user's dwell time in front of the exhibit and body orientation: if the dwell time exceeds the average attention time of this type of exhibit and the angle between the body orientation and the center of the exhibit is less than 30 degrees, then B(P) is 1.0; if only one of them is satisfied, it is 0.5; if neither is satisfied, it is 0.
[0094] Furthermore, based on real-time collected user behavior data, user interests are analyzed. Combined with exhibit relevance and real-time exhibition hall environment data, user fatigue is analyzed and time constraints are constructed based on user visit duration and movement patterns. The visit path is then dynamically adjusted to obtain the first visit path, which includes:
[0095] S401. Based on real-time collected user behavior data, extract the user's attention weight for different exhibit attributes, analyze user interests, and obtain interest vectors.
[0096] S402. Based on the interest vector, combined with the relevance of exhibits and real-time environmental data of the exhibition hall, the visitor path is dynamically adjusted to obtain the first visitor path.
[0097] In this embodiment, the system continuously collects user behavior data, including but not limited to the user's attention score for each exhibit, the number of times the user repeatedly listens to different chapters of the exhibit's explanation, the user's zoom-in view of the exhibit's detailed images, and the intent categories involved in the user's voice questions during the exhibition. The system uses a matrix factorization model based on collaborative filtering to analyze user interests. The training dataset for this model includes the system's built-in historical exhibition data and the real-time behavior data generated by the current user during the current visit.
[0098] Preferably, the specific architecture of the matrix factorization model includes: the input layer receives user IDs and exhibit IDs; the model internally comprises two low-dimensional dense vector matrices: a user latent factor matrix, where each row represents a user and each column represents a latent interest factor; and an exhibit latent factor matrix, where each row represents an exhibit and each column also represents a latent factor. The model training process employs a stochastic gradient descent optimization algorithm with a learning rate of 0.001 and a regularization parameter of 0.01. The loss function is the root mean square error loss. While learning to predict user behavior, the model learns the weight distribution of each user on the latent factors. The weight distribution is parsed and mapped, aligned with predefined exhibit attributes, and an interest vector is output.
[0099] It should be noted that the matrix factorization model can transform fragmented user behavior data into interest vectors, quantifying user interests and the specific exhibit attributes that interest them. This provides an accurate data foundation for setting up personalized services, improves the matching degree between the explanation content and user interests, and enhances the user experience.
[0100] Specifically, based on interest vectors, combined with exhibit relevance and real-time exhibition hall environmental data, the visitor path is dynamically adjusted to obtain the primary visitor path. By combining exhibit relevance and real-time exhibition hall environmental data to output the primary visitor path, the user's personalized interests, the continuity of knowledge acquisition, and the comfort of the visiting environment can be balanced, thereby improving the user's exhibition experience.
[0101] Furthermore, based on the interest vector, combined with the exhibit relevance and real-time exhibition hall environmental data, the visitor path is dynamically adjusted to obtain the first visitor path, which includes:
[0102] S501. Extract exhibit attribute information and association information between exhibits from the pre-constructed exhibit knowledge graph, calculate the association degree between exhibits, and construct an exhibit association matrix;
[0103] S502. Analyze user fatigue based on visit duration and movement patterns, and establish time constraints.
[0104] S503. Combining the interest vector, exhibit association matrix, time constraints, and real-time environmental data of the exhibition hall, dynamically adjust the visitor path to obtain the first visitor path.
[0105] In this embodiment, knowledge is extracted from exhibition materials, academic literature, historical archives, and other data. Each exhibit is treated as an independent node, and each node contains corresponding exhibit attribute information. Connection edges are established based on the association information between exhibit nodes, and different weight coefficients are set for different types of association edges. The weight coefficients are preset according to the tightness of the association. For example, the weight coefficient for the same author is set to 1.0, indicating a strong association; the weight coefficient for those from similar eras is set to 0.6, indicating a medium association; and the weight coefficient for those from similar places of origin is set to 0.3, indicating a weak association.
[0106] Specifically, the shortest path between two nodes is calculated, the weight coefficients of all edges on the path are summed, and this sum is divided by the path length to obtain the corresponding correlation score. A higher value indicates a closer connection between the two exhibits in terms of history, culture, art, etc. The correlation scores are calculated pairwise for all exhibits to construct an exhibit correlation matrix. The rows and columns of the matrix represent unique identifiers for the exhibits. The value in the i-th row and j-th column of the matrix represents the correlation score between exhibit i and exhibit j, ranging from 0 to 1. By constructing the exhibit correlation matrix, the connections between exhibits are transformed into corresponding numerical values. During exhibit recommendation, user interests and the correlation between exhibits can be considered to improve the logical flow of exhibit explanations during visits, enhancing the systematic nature and comprehension of the explanation process.
[0107] Furthermore, the user's physiological state is analyzed to ensure that the recommended tour itinerary is within the user's physical and mental capacity. The total time elapsed from the user's entry into the exhibition hall to the current moment is continuously recorded as the cumulative tour duration, and the Euclidean distance between adjacent coordinate points in the location coordinate sequence is accumulated as the cumulative travel distance. A fatigue analysis model is constructed based on prior knowledge and experimental data. This model includes, but is not limited to, a linear weighted function. The model uses the cumulative tour duration and cumulative travel distance as input variables, sets corresponding weight coefficients, and then sums the two input variables with the weight coefficients. The model outputs a fatigue index.
[0108] For example, two thresholds are set to analyze fatigue levels. When the fatigue index is below 0.5, the user is considered energetic; between 0.5 and 0.8, the user is considered to be starting to feel fatigued, and the pace of the visit needs to be appropriately slowed down; above 0.8, the user is considered to be highly fatigued, and a rest stop needs to be arranged. Time constraints are constructed based on the fatigue index, setting an upper limit for the remaining visit time or adding rest stops along the path, including rest areas within the exhibition hall as mandatory nodes. By analyzing fatigue and constructing time constraints, the constructed first visit path can consider the user's physiological endurance, proactively adapting to changes in the user's physical condition, avoiding user fatigue due to excessively long routes or fast pace, thus improving the comfort of the visit and enhancing the user experience.
[0109] Specifically, during the path planning process, the system combines the user's current physical location, interest vector, exhibit association matrix, time constraints, and real-time environmental data of the exhibition hall acquired through the IoT platform. The exhibits throughout the exhibition hall are treated as a set of nodes to be visited, and the actual walking paths and distances between exhibits are used as edges to construct a search graph. The planning process starts from the user's current location and uses the A* algorithm to search towards the user-defined destination. At each step of the exploration, when considering adding a candidate exhibit to the current path, a comprehensive scoring function is called to calculate the corresponding incremental value.
[0110] Preferably, the comprehensive scoring function calculates the matching degree between candidate exhibit attributes and user interest vectors, obtaining an interest score by calculating the dot product of the exhibit attribute vector and the interest vector; the correlation between the candidate exhibit and the previous exhibit in the path is read from the exhibit association matrix to obtain a coherence score; real-time environmental data of the candidate exhibit area is queried, and if the crowd density is higher than a preset density threshold (which can be set to 2 people per square meter), a penalty score is deducted based on the degree of crowding, with higher penalties for higher crowding; the interest score, coherence score, and environmental penalty score are weighted and combined to obtain a comprehensive score, and the total time of the current path is checked to see if it exceeds the time constraint limit. The exhibit sequence with the highest comprehensive score and that meets the time constraint is selected, and the first visiting path is output according to the exhibit sequence.
[0111] It should be noted that by combining and integrating interest vectors, exhibit association matrices, time constraints, and environmental data, the planned first visit route can balance users' personalized interests, the continuity of knowledge acquisition, the comfort of the visiting environment, and the users' own physiological state, thereby improving the users' exhibition experience and the continuity of the explanation process.
[0112] Furthermore, during the explanation process, in response to user voice questions, the system analyzes the intent of the question and identifies the corresponding exhibits, generating answers based on user interests, including:
[0113] S601. During the explanation, in response to the user's voice question, the user's voice signal is converted into text form based on the real-time collected user voice signal to obtain the question text;
[0114] S602. Analyze the question intent based on the question text and determine the corresponding exhibits, and generate answer content based on user interests.
[0115] In this embodiment, when the system is playing narration content, the terminal microphone is always in a low-power listening state, and the user's questions are identified through a speech activity detector based on short-time energy and zero-crossing rate. Once the user starts speaking, the acquired audio stream is immediately used as the user's speech signal, and language recognition is performed through a speech recognition model. The speech recognition model adopts a deep learning model, including but not limited to a convolutional recurrent neural network based on connectionist temporal classification. The input of the model is the spectral features of the speech signal. The original audio is pre-emphasized to enhance high-frequency signals, and frame-by-frame processing is performed. After applying a Hamming window function to each frame signal, a fast Fourier transform is performed to obtain the spectrum, and then the corresponding Mel frequency cepstral coefficient features are extracted through a Mel filter bank as the model input. The model consists of a multi-layer convolutional neural network and a bidirectional long short-term memory network. The convolutional layers are used to extract local spectral features, and the recurrent layers are used to capture the temporal dependencies in the speech signal.
[0116] Preferably, the training data for the speech recognition model includes tens of thousands of hours of annotated speech data from hundreds of thousands of different speakers. During training, a stochastic gradient descent optimizer is used, with an initial learning rate set to 0.01 and an exponential decay strategy. The momentum factor is 0.9, and the loss function adopts connectionist temporal classification loss. This allows the model to be trained end-to-end without the need for precise alignment of the input and output beforehand, outputting a probability distribution sequence. Based on linguistic knowledge, the probability of various word sequences is analyzed, and the probability distribution sequence is decoded into a coherent text sequence using the Viterbi decoding algorithm to obtain the question text in text form.
[0117] It should be noted that the speech recognition model can accurately convert users' verbal questions into corresponding text data, providing accurate input data for subsequent intent understanding analysis. This ensures that the system's subsequent analysis will not be biased due to errors in speech recognition, thereby improving the accuracy and reliability of the question-and-answer process.
[0118] Specifically, based on the analysis of the question text, the system identifies the question's intent and determines the corresponding exhibits, then generates answers by combining these with user interests. By integrating intent understanding, knowledge graph queries, and user interest profiling, the system can accurately understand the deeper intent behind user questions and filter out information that users are interested in, satisfying their thirst for knowledge and interest needs, and improving user satisfaction throughout the process.
[0119] like Figure 2 As shown, based on the analysis of the question text, the questioner's intent is determined, and the corresponding exhibits are identified. Then, the answer content is generated based on the user's interests, including:
[0120] S701. Extract the intent category and entity information from the question text. The intent category includes inquiries about the year, craftsmanship, and related exhibits. The entity information includes exhibit name, exhibit number, and exhibit characteristics.
[0121] S702. Match the entity information with the exhibit identifier currently being explained to determine the exhibit identifier corresponding to the question information;
[0122] S703. Based on the intent category and the exhibit identifier in question, filter out the corresponding exhibit knowledge from the pre-constructed exhibit knowledge graph to obtain the first knowledge set;
[0123] S704. Calculate the similarity between exhibit knowledge and user interests in the first knowledge set, filter out exhibit knowledge with similarity greater than a preset similarity threshold, obtain a second knowledge set, integrate the exhibit knowledge in the second knowledge set, and generate answer content.
[0124] In this embodiment, a joint intent recognition and entity extraction model is invoked. This model includes a pre-trained masked language model masked on a massive general corpus and a Chinese BERT-base model trained on a next-sentence prediction task. The Chinese BERT-base model is then fine-tuned using professional questions from the museum domain. The training dataset used for fine-tuning includes a large number of historical question samples, labeled with intent category labels and entity information labels by museum experts. The AdamW optimizer is used during model fine-tuning, with a learning rate set to 3×10. -5The batch size is 16, and the training rounds are 5. The loss function is a joint loss function for multi-task learning, which is obtained by weighted summation of cross-entropy loss for intent classification and conditional random field loss for entity recognition, with weight coefficients of 0.6 and 0.4, respectively. During training, the model continuously adjusts its parameters through backpropagation to learn the mapping relationship from input text to intent and entity labels.
[0125] After the question text is input into the trained masked language model, the model propagates forward, and the output layer provides the prediction results. A softmax layer outputs a probability distribution indicating the probability that the text belongs to each predefined intent category. One or more of the highest probabilities are selected as the final intent. The model then uses a conditional random field layer to predict an entity label for each character or word unit in the text, obtaining entity information. Through a deep learning model fine-tuned with domain data, the model can accurately extract the user's intent category and corresponding entity information from the user's natural language questions. This provides accurate data support for subsequent explanations in the knowledge base, avoiding mismatched answers due to misunderstandings of intent.
[0126] Specifically, the system maintains the current user's dialogue state, including the currently identified exhibit identifier and the user's recently viewed exhibit history. When the entity information contains a clear exhibit name or exhibit number, the system queries the database and matches the exhibit name with a standard exhibit identifier library based on contextual information to determine the exhibit identifier in question. By analyzing the user's dialogue state and contextual information, the system can accurately point the question to the corresponding exhibit, avoiding knowledge retrieval errors caused by unclear exhibit references, improving the continuity and accuracy of the question-and-answer process, and enhancing the fluency and user experience of the question-and-answer dialogue.
[0127] The standard exhibit identification database is a pre-built structured database used to store the unique identification code of each exhibit and its associated information. The database uses the exhibit identification code as the primary key and stores the exhibit's name, number, category, year, author, location coordinates and other attribute information. The database is built based on the official collection archives provided by the exhibition hall and is imported after data cleaning and standardization to ensure that each exhibit has a globally unique identification code.
[0128] Furthermore, the system includes a pre-built exhibit knowledge graph, stored in a graph database. Nodes represent entities, including exhibit nodes and auxiliary nodes such as people, locations, and events. Nodes are connected by semantic relationship edges. The graph construction process involves a team of domain experts extracting and structuring knowledge from exhibit materials, academic papers, and archaeological reports, and importing this information into the graph database. Based on the intent category and the exhibit identifier in the question, a query statement is set for the graph database. The database is then searched to retrieve knowledge entries directly related to the user's query, including text, images, and links, forming the first knowledge set. Through the exhibit knowledge graph and graph query language, the system can quickly respond to complex and multi-dimensional user queries, rapidly filtering out corresponding explanations and other information related to the answer, thus improving the comprehensiveness and accuracy of the first knowledge set.
[0129] Specifically, the sentence encoder model represents each piece of knowledge in the first knowledge set as a vector form whose similarity can be calculated by the computer. Based on a pre-trained BERT architecture, the sentence encoder model consists of multiple layers of bidirectional Transformer encoders stacked together. Through fine-tuning on tasks such as natural language inference and semantic text similarity, it learns to map text of arbitrary length to a fixed-length semantic vector space, making semantically similar texts closer in the vector space. The training dataset for the sentence encoder model contains a large number of text pairs, labeled with similarity scores between the text data. A contrastive learning loss function is used during training.
[0130] Preferably, each knowledge text in the first knowledge set is input into the sentence encoder model to obtain a corresponding semantic vector representation, which serves as the knowledge vector. Based on the user interest vector, the cosine similarity between each knowledge vector and the user interest vector is calculated. The cosine similarity value ranges from -1 to 1; a higher value indicates a better semantic match. A preset similarity threshold is set according to the desired level of personalization. For example, when highly personalized answers are desired, a higher similarity threshold of 0.7 is set; when more comprehensive information is desired, a lower similarity threshold of 0.4 is set. Knowledge with similarity scores below the similarity threshold is removed, while knowledge with similarity scores above the similarity threshold is retained, resulting in the second knowledge set.
[0131] Based on the second knowledge set, a GPT-2 model based on the Transformer architecture was invoked and fine-tuned on a question-answering corpus in the museum domain. The fine-tuning dataset consisted of a knowledge point set and standard answer pairs. During training, the Adam optimizer was used with a learning rate of 3×10⁻⁶. -5The loss function is cross-entropy loss, which maximizes the probability of generating the correct answer. The model's input is a list of knowledge points in the second knowledge set and the user's original question intent. The answer content is generated through autoregression. During the generation process, an attention mechanism is used to ensure that the generated content covers all knowledge points and to establish logical connections between different knowledge points. The answer content is then obtained and played to the user.
[0132] It should be noted that by combining user interest vectors for semantic filtering and using deep learning for language generation, question-and-answer interactions can be conducted based on users' personal interests, ensuring that the content explained is the information that users need, improving the efficiency and effectiveness of information acquisition, and enhancing the user's question-and-answer experience.
[0133] Furthermore, combining the first tour route and the answers provided, the first explanation was optimized by adjusting the order of the explanations and increasing the coherence of the answers, resulting in the second explanation, which includes:
[0134] S801. Obtain the first visit path and answer content. The first visit path includes a list of exhibit labels arranged in visit order and the corresponding stay time and explanation time allocation for each exhibit.
[0135] S802. Divide the first explanation content according to the exhibit labels, extract the explanation segments corresponding to each exhibit, and obtain a set of exhibit explanation segments;
[0136] S803. Based on the first visitor path and the answers provided, adjust the order of the exhibit explanation segments and increase the coherence of the inserted answers to obtain the second explanation content.
[0137] In this embodiment, the first visit path includes a list of exhibit identifiers, a dwell time, and a time allocation for explanations. The exhibit identifier list is arranged according to the visit order planned by the system for the user. The dwell time is an estimated dwell time allocated to each exhibit identifier in the list, derived from the user's historical average attention time for that type of exhibit. The time allocation for explanations specifies the proportion of time that should be allocated to different explanation segments when playing the explanation content for each exhibit identifier. This proportion can be adjusted according to the weight of the corresponding attribute in the user's interest vector. The answer content during this visit is retrieved from the dialogue history. Each answer includes the exhibit identifier asked, the intent category, the text content, and the corresponding audio file address. The path information and question-and-answer information are associated and encapsulated according to the user's session ID.
[0138] It should be noted that by acquiring and integrating the first tour route and the answers, an accurate data foundation is provided for subsequent content analysis and arrangement. The route information clarifies the order of explanation and time allocation, while the Q&A information provides user interest information, including user points of interest. This ensures that the explanation content can accurately locate the content elements that need to be rearranged and inserted, avoiding arrangement errors caused by missing data or structural confusion.
[0139] Specifically, the system accesses a backend exhibit explanation knowledge base. This knowledge base contains the initial explanation content for all exhibits. This initial explanation content is then divided according to the explanation content index table in the knowledge base. The index table records the start and end timestamps of each logical paragraph or knowledge point in the complete explanation content for each exhibit, the corresponding knowledge point tags, and the association between that segment and the exhibit identifier. The system iterates through the list of all exhibit identifiers included in the first visitor path. For each exhibit identifier in the list, based on the explanation content index table, all relevant explanation segments are extracted from the corresponding complete initial explanation content and temporarily stored. All explanation segments extracted from the exhibits along the first visitor path are then integrated to obtain a set of exhibit explanation segments. By dividing the explanation content, the explanation order can be adjusted, and different knowledge points can be combined to synchronize the generated explanation content with user interests, improving the accuracy of the explanation process.
[0140] Specifically, based on the first visitor path and the responses received, the order of the exhibit explanation segments is adjusted, and the coherence of the responses is enhanced to obtain the second set of explanation content. By adjusting and optimizing the order of the exhibit explanation segments, and combining the user's physical movement trajectory and cognitive interaction results, the explanation content is generated. The explanation content is then adjusted and optimized based on the user's viewing experience to improve the user experience.
[0141] Furthermore, based on the first visitor path and the responses, the order of the exhibit explanation segments is adjusted and the coherence of the inserted responses is increased to obtain the second explanation content, including:
[0142] S901. Sort the set of exhibit explanation segments according to the order of exhibits in the first visitor path to obtain an updated set of exhibit explanation segments.
[0143] S902. Analyze the exhibit identifier corresponding to the answer content, locate the corresponding explanation segment in the updated exhibit explanation segment set, and match it with the explanation segment to determine the insertion position of the answer content.
[0144] S903. According to the insertion position, insert the answer content into the corresponding explanation segment, analyze the coherence between the answer content and the explanation segment, optimize the explanation content, and obtain the second explanation content.
[0145] In this embodiment, a list of exhibit identifiers arranged in sequence is parsed from the first visitor path, representing the order in which the user will visit the exhibits. A set of exhibit explanation segments is obtained, with each segment bound to an exhibit identifier. The exhibit identifier list in the first visitor path is traversed. For each exhibit identifier in the list, all explanation segments bound to that exhibit identifier are selected from the exhibit explanation segment set, sorted, and then extracted sequentially and appended to a new, empty list. This process is repeated until all exhibits on the path have been processed, resulting in an updated set of exhibit explanation segments.
[0146] It should be noted that by arranging the explanatory segments, it is possible to ensure that the content of the explanation is synchronized with the viewing of the exhibits. By maintaining consistency between auditory and visual experiences, users can focus their attention on the current exhibit, thereby improving the efficiency of information absorption and the smoothness of the viewing experience.
[0147] Specifically, the process iterates through the answer content. For each answer, it retrieves the associated exhibit identifier. Using this identifier as an index, it searches the updated exhibit explanation segment set, filtering out the start and end indices of all explanation segments belonging to that exhibit within the sequence. After identifying the segment range belonging to the exhibit, it analyzes the intent category corresponding to the answer content and compares it with the knowledge point tags corresponding to each explanation segment of the exhibit. It measures the matching degree by calculating the cosine similarity between the intent category text and the knowledge point tag text. It also calculates the similarity between the answer content intent and each candidate segment tag, selecting the segment with the highest similarity as the target segment for insertion. The insertion position is determined for each answer, and this position information is recorded after the target segment.
[0148] It should be noted that by matching intent and tag semantics, the most suitable position can be selected for adding content to the user's personalized Q&A, so that the Q&A content is closely integrated with relevant knowledge points, ensuring the logical coherence of the explanation and the integrity of the knowledge structure, making the question and the current explanation correspond, and improving the consistency of the explanation process and the user experience.
[0149] Specifically, based on the insertion position determined for each answer, content is spliced. Starting with updating the set of exhibit explanation segments, after adding the segment to the new sequence, all answer content is inserted into the corresponding positions to obtain a preliminary spliced sequence. The coherence optimization model is then used to process the newly added splicing points. This model is a sequence-to-sequence model based on the Transformer architecture, containing an encoder and a decoder. The encoder consists of multiple stacked multi-head self-attention layers and feedforward network layers. Each layer includes residual connections and layer normalization, used to semantically encode the preceding and following sentences of the input. The decoder also consists of multiple layers, each containing a self-attention layer, a cross-attention layer, and a feedforward network layer. The training dataset for this model includes a large amount of transition sentences connecting two sentences written by language experts. During training, the model learns to generate transition sentences based on the semantics of the preceding and following sentences. The Adam optimizer is used during training, with a learning rate set to 3×10⁻⁶. -5 The batch size is 32, and the loss function is cross-entropy loss, aiming to maximize the probability of generating the correct transition sentence.
[0150] During the optimization process, for each splicing point, the coherence optimization model takes as input the ending text of the previous segment and the beginning text of the next segment. The model outputs one or more candidate transition statements. The fluency of each transition statement is scored, and the one with the highest score is inserted between the two original segments. After processing and optimizing all insertion points, the second explanation content is obtained. By optimizing the transition process, the fluency of the explanation content is improved, enhancing user comprehension and experience during the explanation.
[0151] like Figure 3 As shown, the intelligent exhibit recognition and explanation system for exhibition scenes is used to implement intelligent exhibit recognition and explanation methods in exhibition scenes, including:
[0152] The exhibit identification module identifies the exhibits that the user is currently interested in based on real-time user terminal perception data, obtains exhibit identifiers, analyzes the time the user spends at the corresponding exhibit identifier location, provides the corresponding initial explanation content, and begins the explanation.
[0153] The user interest analysis module analyzes user interests based on real-time collected user behavior data. It combines exhibit relevance and real-time exhibition hall environment data to analyze user fatigue and construct time constraints based on user visit duration and movement. It then dynamically adjusts the visit path to obtain the first visit path.
[0154] The question-and-answer module responds to user voice questions during the explanation process, analyzes the intent of the question, identifies the corresponding exhibits, and generates answers based on user interests.
[0155] The explanation optimization module combines the first tour route and the answers provided. By adjusting the order of the explanation content and increasing the coherence of the inserted answers, the first explanation content is optimized to obtain the second explanation content.
[0156] The above description is merely a preferred embodiment of this application. The scope of protection of this application is not limited to the above embodiments. All technical solutions falling within the scope of this application's concept are within the scope of protection of this application. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of this application should also be considered within the scope of protection of this application.
Claims
1. A method for intelligent exhibit recognition and explanation in exhibition scenarios, characterized in that, include: Based on real-time collected user terminal perception data, the exhibits that the user is currently interested in are identified, exhibit icons are obtained, the dwell time of the user at the corresponding exhibit icon location is analyzed, the corresponding first explanation content is provided and the explanation begins. Based on real-time collected user behavior data, user interests are analyzed. Combined with the relevance of exhibits and real-time environmental data of the exhibition hall, user fatigue is analyzed and time constraints are constructed according to the user's visit duration and movement. The visit path is dynamically adjusted to obtain the first visit path. During the explanation, responding to user voice questions, analyzing the intent of the question and determining the corresponding exhibits, and generating answer content based on user interests; By combining the first tour route and the answers provided, and by adjusting the order of the explanations and increasing the coherence of the answers, the first explanation was optimized to obtain the second explanation.
2. The intelligent exhibit recognition and explanation method for exhibition scenarios according to claim 1, characterized in that, The process involves identifying the exhibits currently of interest to the user based on real-time collected user terminal perception data, obtaining exhibit identifiers, analyzing the user's dwell time at the corresponding exhibit identifier location, providing corresponding initial explanation content, and beginning the explanation, including: Based on real-time collected user terminal perception data, the exhibit area is detected and the duration of user stay in the exhibit area is calculated to identify the exhibits that the user is currently interested in and obtain exhibit identification. Analyze the time users spend at the corresponding exhibit markers, filter out exhibits whose dwell time exceeds a preset dwell threshold, and generate an explanation trigger signal; In response to the explanation trigger signal, the corresponding first explanation content is retrieved from the pre-built exhibit explanation knowledge base, and the explanation begins.
3. The intelligent exhibit recognition and explanation method for exhibition scenarios according to claim 2, characterized in that, The process of detecting and calculating the duration of user stay in the exhibition area based on real-time collected user terminal perception data, identifying the exhibits currently of interest to the user, and obtaining exhibit identifiers includes: Acquire real-time user terminal perception data, which includes real-time image sequences, voice signal segments, location coordinate sequences, and motion trajectory data; Identify exhibit regions in the real-time image sequence, analyze candidate exhibits in each frame of the image, and obtain a first exhibit set; Filter out the exhibit feature keywords in the speech signal segments, determine the exhibits corresponding to the exhibit feature keywords, and obtain the second set of exhibits; Based on the location coordinate sequence and motion trajectory data, calculate the duration of user stay in the exhibition area and extract corresponding behavioral features; The intersection of exhibits in the first and second exhibit sets is selected, and the attention score of each exhibit in the intersection is calculated based on the behavioral characteristics. The exhibit with the highest attention score is then identified as the exhibit.
4. The intelligent exhibit recognition and explanation method for exhibition scenarios according to claim 1, characterized in that, The process involves analyzing user interests based on real-time collected user behavior data, combining exhibit relevance and real-time exhibition hall environment data, analyzing user fatigue based on visit duration and movement patterns, constructing time constraints, and dynamically adjusting the visit path to obtain the first visit path, which includes: Based on real-time collected user behavior data, we extract the attention weights of users to different exhibit attributes, analyze user interests, and obtain interest vectors. Based on the interest vector, combined with the relevance of exhibits and real-time environmental data of the exhibition hall, the visitor path is dynamically adjusted to obtain the first visitor path.
5. The intelligent exhibit recognition and explanation method for exhibition scenarios according to claim 4, characterized in that, The first visitor path is obtained by dynamically adjusting the visitor path based on the interest vector, combined with the relevance of exhibits and real-time environmental data of the exhibition hall, including: Extract exhibit attribute information and association information between exhibits from the pre-constructed exhibit knowledge graph, calculate the association degree between exhibits, and construct an exhibit association matrix; Based on user visit duration and movement patterns, analyze user fatigue levels and construct time constraints; By combining the interest vector, exhibit association matrix, time constraints, and real-time environmental data of the exhibition hall, the visitor path is dynamically adjusted to obtain the first visitor path.
6. The intelligent exhibit recognition and explanation method for exhibition scenarios according to claim 5, characterized in that, During the explanation process, in response to user voice questions, the system analyzes the intent of the question, identifies the corresponding exhibits, and generates answers based on user interests, including: During the explanation, in response to user voice questions, the user voice signal is converted into text form based on the real-time collected user voice signal to obtain the question text; Based on the analysis of the question text, the intent of the question is determined and the corresponding exhibits are identified. Then, the answer content is generated in combination with the user's interests.
7. The intelligent exhibit recognition and explanation method for exhibition scenarios according to claim 6, characterized in that, The process of analyzing the question text to determine the questioner's intent and the corresponding exhibits, and generating answer content based on user interests, includes: Extract the intent category and entity information from the question text. The intent category includes inquiries about the era, craftsmanship, and related exhibits. The entity information includes exhibit name, exhibit number, and exhibit characteristics. The entity information is matched with the exhibit identifier currently being explained to determine the exhibit identifier corresponding to the question information; Based on the intent category and the exhibit identifier in question, the corresponding exhibit knowledge is filtered from the pre-constructed exhibit knowledge graph to obtain the first knowledge set; Calculate the similarity between exhibit knowledge and user interests in the first knowledge set, filter out exhibit knowledge with similarity greater than a preset similarity threshold to obtain a second knowledge set, and integrate the exhibit knowledge in the second knowledge set to generate answer content.
8. The intelligent exhibit recognition and explanation method for exhibition scenarios according to claim 1, characterized in that, The second set of explanations is obtained by combining the first tour route and the answers provided, adjusting the order of the explanations and increasing the coherence of the answers, and optimizing the first explanation. This second explanation includes: Obtain the first visit path and the answer content. The first visit path includes a list of exhibit labels arranged in the order of visit and the corresponding dwell time and explanation time allocation for each exhibit. The first part of the explanation is divided according to the exhibit labels, and the explanation segments corresponding to each exhibit are extracted to obtain a set of exhibit explanation segments; Based on the first visitor route and the answers provided, the order of the exhibit explanation segments is adjusted and the continuity of the inserted answers is increased to obtain the second explanation content.
9. The intelligent exhibit recognition and explanation method for exhibition scenarios according to claim 8, characterized in that, The second set of explanation content is obtained by adjusting the order of the exhibit explanation segments and increasing the coherence of the inserted explanation content based on the first visitor path and the answers provided, including: Based on the order of exhibits in the first visitor route, the set of exhibit explanation segments is sorted to obtain an updated set of exhibit explanation segments; Analyze the exhibit identifier corresponding to the answer content, locate the corresponding explanation segment in the updated exhibit explanation segment set, and match it with the explanation segment to determine the insertion position of the answer content; According to the insertion position, the answer content is inserted into the corresponding explanation segment, and the coherence between the answer content and the explanation segment is analyzed. The explanation content is then optimized to obtain the second explanation content.
10. An intelligent exhibit recognition and explanation system for exhibition scenes, characterized in that: The intelligent exhibit recognition and explanation method for implementing the exhibition scene as described in any one of claims 1 to 9 includes: The exhibit identification module identifies the exhibits that the user is currently interested in based on real-time user terminal perception data, obtains exhibit identifiers, analyzes the time the user spends at the corresponding exhibit identifier location, provides the corresponding initial explanation content, and begins the explanation. The user interest analysis module analyzes user interests based on real-time collected user behavior data. It combines exhibit relevance and real-time exhibition hall environment data to analyze user fatigue and construct time constraints based on user visit duration and movement. It then dynamically adjusts the visit path to obtain the first visit path. The question-and-answer module responds to user voice questions during the explanation process, analyzes the intent of the question, identifies the corresponding exhibits, and generates answers based on user interests. The explanation optimization module combines the first tour route and the answers provided. By adjusting the order of the explanation content and increasing the coherence of the inserted answers, the first explanation content is optimized to obtain the second explanation content.