Exhibition Hall Explanation System, Method, Electronic Device, Storage Medium and Product

Through multimodal input and output, high-precision synchronization control and personalized service mechanism, the problem of single information acquisition method and inaccurate multimedia display synchronization in the exhibition hall explanation system is solved, and the flexibility and interactivity of the exhibition hall explanation service is improved, and the level of intelligence is improved.

CN119066163BActive Publication Date: 2025-08-05SHENZHEN HAOYUAN DECORATION DESIGN ENG CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411014055.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2025-08-05
Estimated Expiration
2044-07-26

AI Technical Summary

Technical Problem

The existing exhibition hall explanation system lacks flexibility and interactivity, and cannot meet the diverse needs of different visitors for information acquisition methods and content depth. The synchronization of multimedia display and explanation information is inaccurate, resulting in impairment of the coherence and clarity of information transmission.

Method used

Multimodal input and output, high-precision synchronization control and personalized service mechanism are introduced. Multimodal information output is provided through the explanation module. The synchronization control module ensures the time synchronization of voice explanation and multimedia display content. The personalized service module generates personalized explanation content based on the input of visitors.

Benefits of technology

It improves the flexibility and interactivity of the exhibition hall explanation service, ensures the consistency and clarity of information transmission, meets the diverse information acquisition needs of visitors, and improves the intelligence level of the exhibition hall explanation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119066163B_ABST
    Figure CN119066163B_ABST
Patent Text Reader

Abstract

The present application discloses an exhibition hall explanation system, method, electronic device, storage medium and product, which relate to the field of exhibition hall management technology, wherein the system includes: an explanation module, a synchronization control module and a personalized service module, the explanation module is used to provide multimodal information output according to multimodal input information, wherein the multimodal information includes voice explanation content and multimedia display content; the synchronization control module is used to synchronize the voice explanation content with the multimedia display content; the personalized service module is used to generate personalized explanation content according to the multimodal input information. The present application can solve the problems that the existing system has the singleness and fixedness of the explanation content and the interaction mode when providing explanation services, which makes it difficult to meet the diverse information acquisition needs of visitors, and the inaccuracy of the synchronization between multimedia display and explanation information, which leads to the loss of coherence and clarity of information transmission, so as to improve the intelligence level of the exhibition hall explanation system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of exhibition hall management technology, and in particular to an exhibition hall explanation system, method, electronic equipment, storage medium and product. Background Art

[0002] With the continuous advancement of artificial intelligence technology, intelligent exhibition hall explanation systems have gradually become an important means to enhance visitors' visiting experience.

[0003] At present, exhibition hall explanation systems often rely on fixed explanation content and a single interactive method when providing explanation services. They lack flexibility and interactivity and cannot meet the diverse needs of different visitors for information acquisition methods and content depth. In addition, existing exhibition hall explanation systems also have shortcomings in the synchronization of multimedia display content and explanation information. The real-time synchronization between explanation content and displayed multimedia materials is often not accurate enough, which may cause visitors to feel confused or incoherent when obtaining information.

[0004] In summary, how to solve the above problems to improve the intelligence of the exhibition hall explanation system has become a technical problem that needs to be solved urgently in this field. Summary of the Invention

[0005] The main purpose of this application is to provide an exhibition hall explanation system, method, electronic equipment, storage medium and product, aiming to improve the intelligence of the exhibition hall explanation system.

[0006] To achieve the above objectives, the present application proposes an exhibition hall explanation system, which includes: an explanation module, a synchronization control module and a personalized service module;

[0007] The explanation module is used to provide multimodal information output according to the multimodal input information, wherein the multimodal information includes voice explanation content and multimedia display content;

[0008] The synchronization control module is used to synchronize the voice explanation content with the multimedia display content;

[0009] The personalized service module is used to generate personalized explanation content according to the multimodal input information.

[0010] In one embodiment, the synchronization control module includes: a content synchronization unit, a real-time annotation unit, and a status synchronization unit;

[0011] The content synchronization unit is used to update the multimedia display content in real time according to the voice explanation content;

[0012] The real-time annotation unit is used to identify the explanation progress of the voice explanation content and annotate the current explanation content in the multimedia display content according to the explanation progress;

[0013] The state synchronization unit is used to synchronize the lecture state between multiple terminals.

[0014] In one embodiment, the multimodal input information includes voice, text, image and / or video, and the explanation module includes: a front-end fusion unit, a middle fusion unit, a back-end fusion unit and an output unit;

[0015] The front-end fusion unit is used to convert the multimodal input information into standardized feature representations;

[0016] The intermediate fusion unit is used to process each of the feature representations and generate a decision result for each of the feature representations;

[0017] The back-end fusion unit is used to fuse the decision results according to the preset modal data weights to obtain a target decision result;

[0018] The output unit is used to generate the multimodal information according to the target decision result and output the multimodal information.

[0019] In one embodiment, the system further includes a user behavior collection module;

[0020] The user behavior collection module is used to collect the behavior data of visitors in the exhibition hall environment;

[0021] The personalized service module is used to analyze the visitor's visiting preferences and explanation preferences based on the multimodal input information and the behavioral data, recommend exhibition hall content based on the visiting preferences, and adjust the explanation content and interactive elements based on the explanation preferences.

[0022] In one embodiment, the personalized service module is further configured to:

[0023] Acquiring the image characteristics of the visitor and generating a virtual image based on the image characteristics;

[0024] The voice explanation content is outputted through the virtual image.

[0025] In one embodiment, the system further comprises:

[0026] A user positioning module is used to determine the location information of the visitor in the exhibition hall;

[0027] The navigation guidance unit is used to provide guidance for navigating to target exhibition hall content according to the position information, wherein the target exhibition hall content corresponds to the voice explanation content.

[0028] In addition, to achieve the above-mentioned purpose, the present application also proposes an exhibition hall explanation method, which is applied to the exhibition hall explanation system. The exhibition hall explanation method includes:

[0029] Providing multimodal information output according to multimodal input information, wherein the multimodal information includes voice explanation content and multimedia presentation content;

[0030] Synchronizing the voice explanation content and the multimedia presentation content;

[0031] Generate personalized explanation content based on the multimodal input information.

[0032] In addition, to achieve the above-mentioned purpose, the present application also proposes an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the exhibition hall explanation method as described above.

[0033] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the exhibition hall explanation method as described above are implemented.

[0034] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the exhibition hall explanation method as described above.

[0035] This application proposes an exhibition hall explanation system, which includes an explanation module, a synchronization control module and a personalized service module; wherein the explanation module is used to provide multimodal information output based on multimodal input information, and the multimodal information includes voice explanation content and multimedia display content, and can flexibly receive information input from different channels and respond; the synchronization control module is used to synchronize the voice explanation content and the multimedia display content to avoid misalignment or confusion of the output information, and ensure the precise time synchronization of the voice explanation and the multimedia display content; the personalized service module is used to generate personalized explanation content based on the multimodal input information, and fully meet the visiting and explanation needs of different visitors.

[0036] In this way, this application solves the problems of the existing system's singleness and fixedness of explanation content and interaction methods when providing explanation services, which makes it difficult to meet visitors' diverse information acquisition needs, as well as the inaccuracy of synchronization between multimedia display and explanation information, which leads to impaired continuity and clarity of information transmission, by introducing multimodal input and output, high-precision synchronous control and personalized service mechanisms. It improves the flexibility and interactivity of the exhibition hall's explanation services and the intelligence level of the exhibition hall's explanation system. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0038] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0039] Figure 1 The system structure diagram provided for the first embodiment of the exhibition hall explanation system of this application;

[0040] Figure 2 A schematic diagram of a virtual image provided in Example 1 of the exhibition hall explanation system of this application;

[0041] Figure 3 Another system structure diagram provided for the first embodiment of the exhibition hall explanation system of this application;

[0042] Figure 4 A flow chart of the second embodiment of the method for applying for exhibition hall explanation;

[0043] Figure 5 This is a schematic diagram of the equipment structure of the hardware operating environment involved in the exhibition hall explanation method in the embodiment of this application.

[0044] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0045] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0046] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0047] The main solution of the embodiment of the present application is: to provide an exhibition hall explanation system, the system includes: an explanation module, a synchronization control module and a personalized service module; the explanation module is used to provide multimodal information output based on multimodal input information, wherein the multimodal information includes voice explanation content and multimedia display content; the synchronization control module is used to synchronize the voice explanation content with the multimedia display content; the personalized service module is used to generate personalized explanation content based on the multimodal input information.

[0048] At present, exhibition hall explanation systems often rely on fixed explanation content and a single interactive method when providing explanation services. They lack flexibility and interactivity and cannot meet the diverse needs of different visitors for information acquisition methods and content depth. In addition, existing exhibition hall explanation systems also have shortcomings in the synchronization of multimedia display content and explanation information. The real-time synchronization between explanation content and displayed multimedia materials is often not accurate enough, which may cause visitors to feel confused or incoherent when obtaining information.

[0049] The embodiments of the present application provide a solution. By introducing multimodal input and output, high-precision synchronous control and personalized service mechanisms, it solves the problems of the existing system's singleness and fixedness of explanation content and interaction methods when providing explanation services, making it difficult to meet visitors' diverse information acquisition needs, as well as the inaccuracy of synchronization between multimedia display and explanation information, which leads to impaired continuity and clarity of information transmission. It improves the flexibility and interactivity of the exhibition hall's explanation services and the intelligence level of the exhibition hall's explanation system.

[0050] Based on this, the embodiment of the present application provides a pavilion explanation system, referring to Figure 1 , Figure 1 This is a schematic diagram of the system structure of the first embodiment of the exhibition hall explanation system of this application.

[0051] In this embodiment, the exhibition hall explanation system includes: an explanation module, a synchronization control module and a personalized service module;

[0052] The explanation module is used to provide multimodal information output according to the multimodal input information, wherein the multimodal information includes voice explanation content and multimedia display content;

[0053] It should be noted that multimodal input information includes voice, text, images and videos, etc. These input information are converted into a format that the system can understand through corresponding signal processing and feature extraction technologies. In addition, when the system is configured with AR / VR / MR (Augmented Reality / Virtual Reality / Mixed Reality) modules, multimodal input information can also include visitors' gesture operations, touch operations, etc.

[0054] The explanation module is responsible for providing multimodal information output based on the visitors' multimodal input information (such as voice questions, image queries, gesture operations, etc.). The output multimodal information mainly includes voice explanation content and multimedia display content (such as pictures, videos, animations, etc.), conveying exhibition information to visitors in a rich and intuitive form.

[0055] In a feasible implementation, taking visitors' voice questions as an example, the explanation module can convert the visitor's voice questions into text through speech recognition (ASR) technology; then use natural language processing (NLP) technology to perform semantic understanding on the recognized text, judge the visitor's query intention, and generate corresponding answer content; then use speech synthesis (TTS) technology to convert the generated answer content into voice to realize voice explanation. At the same time, through the integrated multimedia content library, appropriate pictures, videos and other multimedia materials are selected for display according to the explanation needs.

[0056] The synchronization control module is used to synchronize the voice explanation content with the multimedia display content;

[0057] The synchronization control module is responsible for ensuring that the voice explanation content and the multimedia display content are synchronized in time and rhythm to enhance the coherence and viewing experience of the explanation.

[0058] In one feasible implementation, the synchronization control module timestamps the audio explanation and multimedia presentation generated by the explanation module, and uses an algorithm to match the timestamps to synchronize the content. The progress of the audio explanation and multimedia presentation is then adjusted in real time based on visitor interactions (e.g., pause, resume, jump, etc.).

[0059] The personalized service module is used to generate personalized explanation content according to the multimodal input information.

[0060] The personalized service module generates personalized explanation content based on the visitors' multimodal input information to meet the needs of different visitors.

[0061] In a feasible implementation, the personalized service module collects and analyzes visitors' input information to build user profiles, understand visitors' interests, preferences and visiting habits, and then uses recommendation algorithms based on user profiles to recommend exhibition content of interest to visitors, provide targeted explanations or generate personalized explanation routes. At the same time, based on visitors' real-time interaction and feedback, the explanation content is dynamically adjusted to generate personalized explanation information that meets visitors' needs.

[0062] In a possible embodiment, Figure 3 As shown, the synchronization control module includes: a content synchronization unit, a real-time annotation unit and a status synchronization unit;

[0063] In this embodiment, in order to improve the coordination between the explanation content and the multimedia presentation, the system provides a synchronization control module including a content synchronization unit, a real-time annotation unit and a status synchronization unit, which aims to ensure the high consistency between the voice explanation content and the multimedia presentation content through precise synchronization technology, while enhancing the audience's sense of participation and understanding.

[0064] The content synchronization unit is used to update the multimedia display content in real time according to the voice explanation content;

[0065] The content synchronization unit monitors and analyzes the audio explanations output by the explanation module, dynamically adjusting and updating the multimedia displays within the exhibition hall based on real-time changes in the audio explanations. For example, when a specific exhibit is introduced, the unit triggers a detailed introduction, high-definition images, or video footage of the corresponding exhibit, ensuring that visitors can instantly access visual information that matches the audio explanation.

[0066] In a feasible implementation, the content synchronization unit adopts an event-driven mechanism. When the voice explanation module outputs new explanation content, a content update event is triggered. The multimedia content management system (CMS) interface is used to dynamically load and display multimedia resources related to the explanation content, and a timestamp or progress bar mechanism is further introduced to ensure accurate synchronization of the multimedia display content and the voice explanation content.

[0067] The real-time annotation unit is used to identify the explanation progress of the voice explanation content and annotate the current explanation content in the multimedia display content according to the explanation progress;

[0068] The real-time unit uses natural language processing and speech recognition technology to accurately identify the current progress and key information points of the voice explanation content. In the multimedia display content output by the explanation module, the real-time annotation unit will automatically add visual elements such as highlights, arrows or text boxes at the corresponding positions according to the identified explanation progress to indicate the specific content of the current explanation, so as to help visitors quickly locate and focus on the key information in the explanation.

[0069] The state synchronization unit is used to synchronize the lecture state between multiple terminals.

[0070] The status synchronization unit is responsible for synchronizing the status of the tours in real time across multiple devices (such as visitors' smartphones and touch screens within the exhibition hall). This ensures that visitors receive a consistent experience regardless of which device they use to access the tours. This synchronization unit operates via a cloud server or local area network communication mechanism, ensuring real-time and accurate data.

[0071] In a feasible implementation, the system can utilize synchronous control technology, real-time content analysis technology, and dynamic annotation algorithms to achieve real-time annotation of the current explanation content. Specifically:

[0072] 1. Synchronous control technology includes

[0073] Content synchronization and scheduling: Synchronize and update the audio explanation content with the multimedia presentation content. This utilizes an event-driven architecture (EDA), a distributed scheduling system (Apache Airflow), and message queues (such as RabbitMQ and Kafka).

[0074] Network synchronization and latency control: Ensure the synchronization of audio explanations and multimedia presentations during network transmission, reducing latency and asynchrony. This utilizes the real-time communication protocol (WebRTC), low-latency network transmission technology (QUIC), and edge computing.

[0075] 2. Real-time content analysis technology

[0076] The GPT (Generative Pre-trained Transformer) pre-trained language model is used for semantic understanding and content analysis.

[0077] Use keyword extraction technology (such as TF-IDF, TextRank) to identify important words and phrases in the voice explanation content.

[0078] Implement sentiment analysis and semantic role labeling to determine the emotional tendency and semantic role of the explanation content, providing richer contextual information for automatic labeling.

[0079] 3. Dynamic Labeling Algorithm

[0080] Based on timestamp synchronization technology, real-time content annotation is achieved by matching the timestamps of the explanation audio and multimedia content.

[0081] Use the Association Rule Learning algorithm to analyze the relationship between the explanation content and the multimedia display content, and dynamically generate annotations.

[0082] Real-time data stream processing: Process real-time data streams to enable simultaneous annotation of explanation and presentation content. Use real-time data processing frameworks (such as Apache Flink and Kafka Streams) for efficient data processing and content annotation.

[0083] Text rendering technology: Render text and image annotations in real time on large screens. Use Canvas and WebGL technologies for efficient text and image rendering.

[0084] In a feasible embodiment, the multimodal input information includes voice, text, image and / or video, such as Figure 3 As shown, the explanation module includes: a front-end fusion unit, a middle fusion unit, a back-end fusion unit and an output unit;

[0085] The front-end fusion unit is used to convert the multimodal input information into standardized feature representations;

[0086] The front-end fusion unit is responsible for converting multimodal input information into standardized feature representations. For example, it uses speech recognition technology to convert speech into text, and uses natural language processing technology to extract text features such as keywords, emotional tendencies, etc.; performs natural language processing on the input text to extract semantic features, entity relationships, etc. in the text; uses computer vision technology to extract features from images and identify objects and scenes in the images; decomposes the video into continuous image frames, and performs image processing on each frame, while extracting the motion features, timing information, etc. of the video. Then, all extracted features are converted into a unified format or vector space for subsequent processing.

[0087] The intermediate fusion unit is used to process each of the feature representations and generate a decision result for each of the feature representations;

[0088] It should be noted that the decision result refers to the specific response or explanation generated by the intermediate fusion unit after analyzing and processing the feature representation of each modal input information. These responses or explanations are the system's understanding of the user input and are used to guide the generation of explanation content.

[0089] The intermediate fusion unit trains or selects appropriate machine learning models (such as classifiers, regressors, etc.) for the feature representations of different modalities, and generates respective decision results through the machine learning models.

[0090] In addition, in actual application scenarios, the intermediate fusion unit can also design a feature interaction mechanism to allow feature representations of different modalities to influence each other to a certain extent, so as to enhance the accuracy of decision making.

[0091] The back-end fusion unit is used to fuse the decision results according to the preset modal data weights to obtain a target decision result;

[0092] It should be noted that module data weight refers to the contribution of different modal data to the final decision result during the multimodal information fusion process. These weights reflect the importance of different modal data in specific tasks. The weight can be pre-set by the system designer as a fixed value, or it can be dynamically adjusted according to specific scenarios and user feedback.

[0093] The back-end fusion unit integrates the decision results generated by the intermediate fusion unit according to the preset modal data weights to obtain the final target decision result.

[0094] The output unit is used to generate the multimodal information according to the target decision result and output the multimodal information.

[0095] The output unit generates multimodal content containing rich information based on the target decision results using text generation technology, image synthesis technology or video editing technology. According to user needs or application scenarios, it selects appropriate output forms, such as voice explanation, text display, image display, video playback, etc., or combines multiple forms to provide a richer explanation experience.

[0096] In this way, the explanation module first converts the diverse input information into a standardized feature representation through the front-end fusion unit, ensuring that information of different modalities can be effectively processed under the same framework; then, the intermediate fusion unit performs independent analysis and decision-making on each feature representation to generate their own preliminary decision results, and then, the back-end fusion unit intelligently fuses these preliminary decision results according to the preset modal data weights, and finally obtains a comprehensive and optimized target decision result; finally, the output unit generates and outputs multimodal information based on this target decision result, and presents it to the user in an intuitive and vivid way.

[0097] In a possible embodiment, Figure 3 As shown, the system also includes a user behavior collection module;

[0098] The user behavior collection module is used to collect the behavior data of visitors in the exhibition hall environment;

[0099] The user behavior collection module is responsible for comprehensively capturing the behavioral data of visitors in the exhibition hall. The behavioral data includes but is not limited to the visiting path, residence time, interaction frequency, etc. This data is collected through sensors installed in the exhibition hall (such as infrared sensors, cameras combined with AI image recognition technology) and smartphones carried by visitors (tracked via Bluetooth / Wi-Fi signals). Deep learning algorithms are then used to analyze the collected data in real time to identify visitors' behavioral patterns and points of interest.

[0100] The personalized service module is used to analyze the visitor's visiting preferences and explanation preferences based on the multimodal input information and the behavioral data, recommend exhibition hall content based on the visiting preferences, and adjust the explanation content and interactive elements based on the explanation preferences.

[0101] The personalized service module analyzes visitors' visiting preferences and explanation preferences based on multimodal input information and user behavior data, and then provides customized exhibition content recommendations based on visitors' visiting preferences. It adjusts the explanation content and interactive elements according to visitors' explanation preferences, for example, adjusting the explanation speed, tone, explanation focus and interactive question and answer frequency, etc., to ensure the targeted nature of the explanation.

[0102] In a feasible embodiment, the personalized service module is further used to:

[0103] Acquiring the image characteristics of the visitor and generating a virtual image based on the image characteristics;

[0104] The personalized service module can accurately obtain the image characteristics of visitors through facial recognition technology or photo information actively provided by visitors. Image characteristics include but are not limited to facial contours, skin color, hairstyle, clothing style, etc. Then, using computer graphics and artificial intelligence algorithms, the personalized service module can automatically generate a highly realistic virtual image based on these image characteristics. This virtual image is not only similar to the visitor in appearance, but can also achieve a certain degree of personalized customization in terms of expressions, movements, etc., so as to better reflect the visitor's emotions and willingness to interact.

[0105] The voice explanation content is outputted through the virtual image.

[0106] The personalized service module uses a generated avatar as the medium for explanations, and uses it as a virtual tour guide to output audio explanations. During the avatar's explanations, not only will the audio be played synchronously, but its expression, speaking speed, and tone will also be adjusted based on the context and visitor's reactions, making the explanations more vivid and engaging.

[0107] For example, in a feasible implementation, the initial virtual image generated by the system is as follows: Figure 2 As shown, the initial virtual image has multimodal interaction capabilities, intelligent analysis and understanding capabilities, emotional expression capabilities and action performance capabilities. Specifically, the initial virtual image has eyes, which can use computer vision technology to "see" visitors and exhibits and understand elements in the visual environment; the initial virtual image has ears, which can hear and recognize visitors' voice commands and questions through automatic speech recognition technology; the initial virtual image has a brain, which can use GPT technology to intelligently analyze and understand input information and provide accurate answers and explanations; the initial virtual image has expressions, which have rich facial expressions and can show emotional expressions according to the content of the explanation and the reactions of the visitors, thereby enhancing the realism of the communication; the initial virtual image has a mouth, which can "speak" by combining natural language processing and speech synthesis technology, and synchronize the lip shape through lip algorithm to make the explanation more vivid; the initial virtual image has movements, which can simulate real human movements and enhance the interactive experience.

[0108] Based on this, the personalized service module transforms the initial avatar into a virtual avatar of the visitor based on the acquired visitor's image characteristics. Specifically, it uses advanced GAN models such as Style GAN to generate the avatar, ensuring authenticity and rich details. Facial recognition and feature extraction technologies are used to accurately capture the visitor's facial features for personalized customization. Avatar replacement technology, combined with skeletal animation and expression capture, enables natural movement and expression changes in the avatar. This not only enhances the fun and interactivity of the tour, but also strengthens visitor participation and immersion.

[0109] In a possible embodiment, Figure 3 As shown, the system further includes:

[0110] A user positioning module is used to determine the location information of the visitor in the exhibition hall;

[0111] The navigation guidance unit is used to provide guidance for navigating to target exhibition hall content according to the position information, wherein the target exhibition hall content corresponds to the voice explanation content.

[0112] The user positioning module determines the visitor's location information in the exhibition hall in real time, and the navigation guidance unit provides visitors with the best path guidance to the target exhibition hall content based on this location information. The target exhibition hall content corresponds to the voice explanation content output by the current explanation module.

[0113] In a feasible implementation, by combining Bluetooth beacons, Wi-Fi positioning, GPS (outdoor assistance) and SLAM (simultaneous localization and mapping) technologies, the user positioning unit can achieve seamless high-precision positioning indoors and outdoors, and then display navigation information to visitors in an intuitive and easy-to-use manner through terminals such as visitors' smartphone APPs, touch screens in the exhibition hall, or AR glasses, guiding visitors to the target exhibition hall content.

[0114] In this way, the embodiments of the present application solve the problems of the singleness and fixedness of the explanation content and interaction methods of the existing system when providing explanation services, which makes it difficult to meet the diverse information acquisition needs of visitors, and the inaccuracy of the synchronization between multimedia display and explanation information, which leads to the loss of continuity and clarity of information transmission, by introducing multimodal input and output, high-precision synchronous control and personalized service mechanisms. It improves the flexibility and interactivity of the exhibition hall's explanation services and the intelligence level of the exhibition hall's explanation system.

[0115] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction, and no further details will be given later. On this basis, the embodiment of the present application provides a method for explaining an exhibition hall, referring to Figure 4 , Figure 4 This is a flowchart of the exhibition hall explanation method for this application.

[0116] In this embodiment, the exhibition hall explanation method is applied to the exhibition hall explanation system in the first embodiment. The exhibition hall explanation method includes steps S10 to S30:

[0117] Step S10, providing multimodal information output according to the multimodal input information, wherein the multimodal information includes voice explanation content and multimedia display content;

[0118] Multimodal information output is provided based on the visitors' multimodal input information (such as voice questions, image queries, gesture operations, etc.). The output multimodal information mainly includes voice explanation content and multimedia display content (such as pictures, videos, animations, etc.), conveying exhibition information to visitors in a rich and intuitive form.

[0119] Step S20, synchronizing the voice explanation content with the multimedia display content;

[0120] The system ensures that the voice explanation content and the multimedia display content are synchronized in time and rhythm to enhance the coherence and viewing experience of the explanation.

[0121] Step S30: Generate personalized explanation content based on the multimodal input information.

[0122] Based on the multimodal input information of visitors, personalized explanation content is generated to meet the needs of different visitors.

[0123] In this way, this embodiment solves the problems of the existing system's singleness and fixedness of explanation content and interaction methods when providing explanation services, which makes it difficult to meet visitors' diverse information acquisition needs, and the inaccuracy of synchronization between multimedia display and explanation information, which leads to impaired continuity and clarity of information transmission, by introducing multimodal input and output, high-precision synchronous control and personalized service mechanisms. It improves the flexibility and interactivity of the exhibition hall's explanation services and the intelligence level of the exhibition hall's explanation system.

[0124] In a feasible embodiment, step S20 may include steps S201 to S203:

[0125] Step S201, updating the multimedia display content in real time according to the voice explanation content;

[0126] Step S202: identifying the progress of the voice explanation content, and marking the current explanation content in the multimedia presentation content according to the explanation progress;

[0127] Step S203: Synchronize the lecture status among multiple terminals.

[0128] In a feasible embodiment, step S10 may include steps S101 to S104:

[0129] Step S101, converting the multimodal input information into standardized feature representations;

[0130] Step S102, processing each of the feature representations to generate a decision result for each of the feature representations;

[0131] Step S103, fusing the decision results according to the preset modal data weights to obtain a target decision result;

[0132] Step S104: generating the multimodal information according to the target decision result, and outputting the multimodal information.

[0133] In a feasible embodiment, step S30 may include steps S301 to S302:

[0134] Step S301, collecting visitor behavior data in the exhibition hall environment;

[0135] Step S302: analyzing the visitor's visiting preferences and explanation preferences based on the multimodal input information and the behavioral data, recommending exhibition hall content based on the visiting preferences, and adjusting explanation content and interactive elements based on the explanation preferences.

[0136] In a feasible embodiment, the exhibition hall explanation method may further include steps A10 to A20:

[0137] Step A10, obtaining the image characteristics of the visitor and generating a virtual image based on the image characteristics;

[0138] Step A20: output the voice explanation content through the virtual image.

[0139] In a feasible embodiment, the exhibition hall explanation method may further include steps B10 to B20:

[0140] Step B10, determining the location information of the visitor in the exhibition hall;

[0141] Step B20: providing navigation guidance to target exhibition hall content according to the location information, wherein the target exhibition hall content corresponds to the voice explanation content.

[0142] Compared with the existing technology, the beneficial effects of the exhibition hall explanation method provided by this application are the same as the beneficial effects of the exhibition hall explanation system provided by the above embodiment, and the other technical features in the exhibition hall explanation method are the same as the features disclosed in the above embodiment system, which will not be repeated here.

[0143] An embodiment of the present application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the exhibition hall explanation method in the above-mentioned embodiment 2.

[0144] Reference below Figure 5 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application can be a mobile electronic device in the form of a crawler robot, a quadruped robot, a humanoid robot, etc. Figure 5 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0145] like Figure 5 As shown, the electronic device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. Various programs and data required for the operation of the electronic device are also stored in RAM 1004. The processing device 1001, ROM 1002, and RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows an electronic device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have instead.

[0146] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0147] The electronic device provided in the embodiments of this application utilizes the exhibition hall explanation method of the above-described embodiments to enhance the intelligence of the exhibition hall explanation system. Compared with the prior art, the beneficial effects of the electronic device provided in the embodiments of this application are the same as those of the exhibition hall explanation method provided in the above-described embodiments, and the other technical features of the electronic device are the same as those disclosed in the exhibition hall explanation method of the above-described embodiments, and are not further described here.

[0148] It should be understood that the various parts disclosed in the embodiments of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any appropriate manner in any one or more embodiments or examples.

[0149] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0150] An embodiment of the present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, and the computer-readable program instructions are used to execute the exhibition hall explanation method in the above embodiment.

[0151] The computer-readable storage medium provided in the embodiments of the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0152] The computer-readable storage medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0153] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the electronic device: provides multimodal information output according to multimodal input information, wherein the multimodal information includes voice explanation content and multimedia presentation content; synchronizes the voice explanation content with the multimedia presentation content; and generates personalized explanation content according to the multimodal input information.

[0154] The computer program code for performing the operations of the embodiments of the present application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect via the Internet).

[0155] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0156] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0157] The computer-readable storage medium provided in the embodiments of this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned exhibition hall explanation method, thereby enhancing the intelligence of the exhibition hall explanation system. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in the embodiments of this application are the same as those of the exhibition hall explanation method provided in the aforementioned embodiments, and are not further elaborated here.

[0158] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned exhibition hall explanation method when executed by a processor.

[0159] The computer program product provided in this application can improve the intelligence of the exhibition hall explanation system. Compared with the existing technology, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the exhibition hall explanation method provided in the above embodiment, which will not be repeated here.

[0160] The above are only some embodiments of the present application and are not intended to limit the patent scope of the present application. All equivalent structural transformations made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. An exhibition hall explanation system, characterized in that: The system includes: an explanation module, a synchronization control module and a personalized service module; The explanation module is used to provide multimodal information output according to the multimodal input information, wherein the multimodal information includes voice explanation content and multimedia display content; The synchronization control module is used to synchronize the voice explanation content with the multimedia display content; The personalized service module is used to generate personalized explanation content according to the multimodal input information; The synchronization control module includes: a content synchronization unit, a real-time annotation unit and a status synchronization unit; The content synchronization unit is used to update the multimedia display content in real time according to the voice explanation content; The real-time annotation unit is used to identify the explanation progress of the voice explanation content and annotate the current explanation content in the multimedia display content according to the explanation progress; The state synchronization unit is used to synchronize the lecture state between multiple terminals; The multimodal input information includes voice, text, image and / or video, and the explanation module includes: a front-end fusion unit, a middle fusion unit, a back-end fusion unit and an output unit; The front-end fusion unit is used to convert the multimodal input information into standardized feature representations; The intermediate fusion unit is used to process each of the feature representations and generate a decision result for each of the feature representations; The back-end fusion unit is used to fuse the decision results according to the preset modal data weights to obtain a target decision result; The output unit is configured to generate the multimodal information according to the target decision result and output the multimodal information; The system also includes a user behavior collection module; The user behavior collection module is used to collect the behavior data of visitors in the exhibition hall environment; The personalized service module is configured to analyze the visitor's visiting preferences and explanation preferences based on the multimodal input information and the behavioral data, recommend exhibition hall content based on the visiting preferences, and adjust explanation content and interactive elements based on the explanation preferences; The personalized service module is further used to: Acquiring the image characteristics of the visitor and generating a virtual image based on the image characteristics; Outputting the voice explanation content through the virtual image; The system further comprises: A user positioning module is used to determine the location information of the visitor in the exhibition hall; The navigation guidance unit is used to provide guidance for navigating to target exhibition hall content according to the position information, wherein the target exhibition hall content corresponds to the voice explanation content.

2. A method for explaining an exhibition hall, characterized in that: Applied to the exhibition hall explanation system according to claim 1, the method comprises: Providing multimodal information output according to multimodal input information, wherein the multimodal information includes voice explanation content and multimedia presentation content; Synchronizing the voice explanation content and the multimedia presentation content; Generate personalized explanation content based on the multimodal input information.

3. An electronic device, characterized in that: The electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the exhibition hall explanation method according to claim 2.

4. A storage medium, characterized in that The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the exhibition hall explanation method according to claim 2 are implemented.

5. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the exhibition hall explanation method according to claim 2 are implemented.

Citation Information

Patent Citations

  • Exhibit information display method and device based on cloud platform, equipment and medium

    CN117235395A

  • Human-machine chess playing method and apparatus, device, storage medium and computer program product

    WO2023142472A1