Multi-modal cognitive method and system for digital heritage experience of silk road
By constructing virtual grid-based sensing areas and using multimodal presentation technology, the problem of traditional tour guide systems being unable to meet personalized needs has been solved. This has enabled personalized explanations that match user behavior perception with content, enhancing the immersive experience and cultural dissemination effect of the Silk Road digital heritage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-03-10
AI Technical Summary
Traditional audio guides or fixed looping videos lack interactivity and cannot meet the personalized needs of different knowledge backgrounds and interests. Users cannot actively explore the personalized experience of the Silk Road digital heritage, and the connection between physical cultural relics and real scenes is weakened.
A virtual grid-like sensing area is constructed, and the user's location and behavior are monitored in real time through a sensor network. Combined with the data acquisition module, the attributes of the display area are obtained. Relevant content is filtered using a similarity algorithm, and personalized explanations are triggered in real time through eye gaze and touch interaction. Multimodal presentation methods are adopted, including video, music and voice narration.
It enabled users to have a personalized immersive experience, enhanced visitors' sense of immersion and participation, optimized the efficiency of information transmission, strengthened their understanding and emotional resonance with the Silk Road cultural heritage, and improved the depth and effectiveness of cultural dissemination.
Smart Images

Figure CN121636726A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital analysis technology, specifically to a multimodal cognitive method and system for experiencing digital heritage along the Silk Road. Background Technology
[0002] As a shared cultural heritage of humankind, the value of the Silk Road lies not only in the artifacts themselves, but also in the history, art, and intellectual exchange behind them. With the maturity of technologies such as the Internet of Things, big data, artificial intelligence, and immersive media, a brand-new technological toolbox has been provided for the digital experience of cultural heritage. The public is no longer satisfied with passive viewing, but craves immersive, interactive, and personalized cognitive experiences. Under the influence of modern consumer electronics and entertainment products, users have put forward higher requirements for the smoothness of interaction, the accuracy of content, and the uniqueness of experience.
[0003] Traditional audio guides or fixed looping videos lack interactivity. Users passively receive information and cannot explore according to their own interests. All users hear the same recording, which cannot meet the personalized needs of different knowledge backgrounds and interests. The content such as images, music, and narration are often separate and do not form an organic whole to create an atmosphere and convey emotions. The system cannot actively perceive the user's interests, and the connection with physical cultural relics and real scenes is weakened. Summary of the Invention
[0004] (a) Technical problems to be solved In view of the above-mentioned shortcomings of the existing technology, the present invention provides a multimodal cognitive method and system for experiencing the digital heritage of the Silk Road, which can effectively solve the problems of the existing technology.
[0005] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: This invention discloses a multimodal cognitive method for experiencing digital heritage along the Silk Road, comprising the following steps: Step 1: Based on the physical display space of the Silk Road digital heritage, divide it into several logical units, each unit corresponding to a display area, and construct a virtual grid sensing area through a sensor network for real-time monitoring of user location and behavior; Step 2: Obtain the overall attributes of each display area and the attributes of individual items within the area through the data acquisition module, and store the attribute data in a preset database; Step 3: Based on the overall attributes and individual project attributes collected in Step 2, index the associated images, music and audio narration in the preset database, and filter out the most relevant content through a similarity algorithm; Step 4: Define an overview of each display area and a unique explanation for each individual project. Based on the different times users may spend on the site, divide the explanation into multiple duration levels, with each level corresponding to different content depth and playback duration. Step 5: Deploy the narration content defined in Step 4 to the corresponding virtual gridded sensing area, including the area overview and video, music or voice narration of individual projects, and set trigger conditions for each content; Step 6: Collect user behavior data in real time, including eye gaze data, dwell time, and touch interaction status, using sensors and interactive devices; Step 7: Based on the user's eye gaze data and touch interaction status collected in Step 6, trigger exclusive explanation content for the corresponding independent project, and trigger explanation content of corresponding duration level according to the user's dwell time.
[0006] Furthermore, the process of constructing the virtual gridded sensing area in step 1 includes: covering the physical display area with a wireless sensor network or camera array, dividing the area into uniform or non-uniform grid cells, assigning a unique identifier to each grid cell, and performing spatial mapping. The gridded sensing area supports dynamic adjustment.
[0007] Furthermore, the process of constructing the virtual gridded sensing area in step 1 includes: using a wireless sensor network or camera array to cover the physical display area, dividing the area into uniform or non-uniform grid units, assigning a unique identifier to each grid unit, and performing spatial mapping, so that the virtual grid can reflect changes in the user's position in real time, and the gridded sensing area supports dynamic adjustment.
[0008] Furthermore, the content indexing and filtering process in step 3 is as follows: Natural language processing algorithms are used to analyze attribute data, generate keyword vectors, and perform similarity matching with a multimodal content library; The filtering process is optimized based on users' historical preferences or contextual environment; The manual customization feature allows administrators to add new content or modify the association rules of existing content through a graphical user interface.
[0009] Furthermore, the formula for calculating the most relevant content using the similarity algorithm in step 3 is as follows: ; In the formula, Representative content item The relevance score indicates the degree of relevance; a higher score indicates a greater relevance. A vector representation of the overall attributes of the display area. Represents multimodal content items The vector representation of , Representative vector The Euclidean norm, Vector representations of attributes of independent projects. The weighting coefficient representing the overall attribute matching. The weighting coefficient represents the matching of attributes for independent projects.
[0010] Furthermore, the logic for setting the duration levels in step 4 includes: dividing the explanation content into a short version of less than 1 minute, a standard version of 1 to 3 minutes, and a detailed version of more than 3 minutes.
[0011] Furthermore, the user behavior data collection process in step 6 includes: Use an infrared eye-tracking device to collect eye gaze data, including gaze point and gaze duration; The duration of stay is calculated using location sensors and timers; Touch interaction status is obtained through a capacitive touchscreen or gesture recognition system; The data collection frequency is adjustable, and encrypted transmission is used to protect user privacy.
[0012] Furthermore, in step 7, when it is detected that a user is looking at a specific item for more than a threshold time or interacting with it by touch, exclusive explanation content is immediately triggered. The duration of the stay trigger adopts a multi-level judgment logic, including: short stay triggers a brief version, medium stay triggers a standard version, and long stay triggers a detailed version.
[0013] A multimodal cognitive system for experiencing digital heritage along the Silk Road includes: The region construction module is used to construct a virtual gridded sensing region based on several real-world display areas. It divides the physical space into multiple grid units, each unit corresponding to a display area or independent project. It locates and senses the user's position and maps it to a virtual coordinate system to associate space with content. The attribute collection module is used to collect the overall attributes of the corresponding display area and the attributes of individual items within the area; in the preset database, it indexes and filters related video, music and audio narration content based on these attributes; it supports manual customization of adding or modifying content, including updating attributes, adding new content or adjusting indexing rules; The content definition module is used to define the overview and explanation content of the display area, and to define exclusive explanation content for individual projects within the display area; and to set different duration levels of explanation content based on different dwell times; The content deployment module is used to deploy video, music, or audio explanations of the area overview and exclusive items corresponding to different dwell times for each sensing area. The data acquisition module is used to collect users' eye gaze data, dwell time and touch interaction status, and monitor user behavior in real time using visual sensors, touch screens or motion detection devices, and convert the data into a processable format; The content triggering module is used to trigger exclusive explanatory content for independent projects based on the collected user eye gaze data and touch interaction status, and to trigger explanatory content of corresponding duration based on the collected user dwell time.
[0014] Furthermore, the content definition module, region construction module, attribute acquisition module, and content deployment module are interconnected via a wireless network, and the data acquisition module is interconnected with the content deployment module and content triggering module via a wireless network.
[0015] (III) Beneficial Effects 1. Compared with the known prior art, the technical solution provided by this invention has the following beneficial effects: By constructing a virtual grid-like sensing area and collecting real-time data on users' eye gaze, dwell time, and touch interactions, the system accurately perceives users' real-time focus of interest and time constraints. This dynamically triggers explanations that match user behavior and initiates personalized explanations for specific artifacts that users are looking at. This not only ensures that each visitor receives tailored information, greatly enhancing their immersion and engagement in the digital heritage experience, but also effectively optimizes the efficiency of information delivery, helping users build a deeper understanding of the Silk Road cultural heritage within a limited time.
[0016] 2. By performing correlation indexing and filtering in the database, and organically integrating multiple media elements into a coherent narrative when triggered, a strong sense of context and historical atmosphere is created. Through an immersive experience, users not only deepen their understanding of the connotation of cultural heritage and their emotional resonance, but also make the historical stories of the ancient Silk Road more vivid, intuitive and easy to perceive, thereby greatly enhancing the depth and effect of cultural dissemination. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0018] Figure 1 This is a flowchart illustrating the multimodal cognition method of the present invention; Figure 2 This is a schematic diagram of the framework of the multimodal cognitive system in this invention.
[0019] The numbers in the diagram represent: 1. Region Construction Module; 2. Attribute Collection Module; 3. Content Definition Module; 4. Content Deployment Module; 5. Data Collection Module; 6. Content Trigger Module. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0021] The present invention will be further described below with reference to embodiments.
[0022] This embodiment presents a multimodal cognitive method for experiencing digital heritage along the Silk Road, such as... Figure 1 As shown, it includes the following steps: Step 1: Based on the physical display space of the Silk Road digital heritage, divide it into several logical units, each corresponding to a display area. A virtual grid sensing area is constructed using a sensor network to monitor user location and behavior in real time. This virtual grid sensing area is aligned with the physical display area via coordinate mapping, ensuring that user interaction data is accurately linked to the corresponding display content. The construction process of the virtual grid sensing area includes: covering the physical display area with a wireless sensor network or camera array, dividing the area into uniform or non-uniform grid units, assigning each grid unit a unique identifier, and performing spatial mapping. This allows the virtual grid to reflect changes in user location in real time. The grid sensing area supports dynamic adjustment, optimizing the grid size and shape based on updated display content or user traffic.
[0023] Step 2: Obtain the overall attributes of each display area and the attributes of individual items within the area through the data acquisition module. The overall attributes include the area theme, historical background, and cultural value, while the attributes of individual items include the name of the cultural relic, its age, and artistic characteristics. The acquisition methods include automatic tag scanning and manual input, and the attribute data is stored in a preset database. The acquisition of display content attributes includes: automatically obtaining the attributes of individual items through RFID tags, QR codes, or near-field communication devices, and supplementing the overall attributes through a manual input interface. The acquired attribute data includes text, image, and audio formats, and an attribute table is established in the preset database using a structured storage method. The table structure includes area ID, item ID, attribute type, and attribute value to ensure efficient data querying and updating.
[0024] Step 3: Based on the overall attributes and individual project attributes collected in Step 2, index related images, music, and audio narrations in the preset database, and filter out the most relevant content using a similarity algorithm. Provide a manual interface for administrators to add, modify, or delete indexed content to ensure accuracy and personalization. The indexing process employs keyword matching and semantic analysis technology. The content indexing and filtering process is as follows: Natural language processing algorithms are used to analyze attribute data, generate keyword vectors, and perform similarity matching with a multimodal content library; The filtering process is optimized based on users' historical preferences or contextual environment; The manual customization feature allows administrators to add new content or modify the association rules of existing content through a graphical user interface, ensuring the flexibility and accuracy of content indexing.
[0025] Step 4: Define an overview of each display area and specific content for each individual project. Based on potential user dwell time, divide the content into multiple duration levels, such as short, standard, and detailed versions. Each level corresponds to different content depth and playback duration to accommodate varying user interests and time constraints. The logic for setting duration levels includes: dividing the content into short versions (less than 1 minute), standard versions (1-3 minutes), and detailed versions (more than 3 minutes). Each version has a different level of content depth: the short version focuses on key highlights, the standard version includes a basic introduction, and the detailed version covers historical details and related stories. Duration level settings are based on user behavior statistical analysis and dynamically configured through the content management system.
[0026] Step 5: Deploy the narration content defined in Step 4 to the corresponding virtual gridded sensing area, including the area overview and video, music or voice narration of individual projects, and set trigger conditions for each content. The deployment process takes into account the content version corresponding to different dwell time to ensure that the system can dynamically call according to user behavior.
[0027] Step 6: Collect user behavior data in real time, including eye gaze data, dwell time, and touch interaction status, using sensors and interactive devices; the user behavior data collection process includes: Use an infrared eye-tracking device to collect eye gaze data, including gaze point and gaze duration; The duration of stay is calculated using location sensors and timers; Touch interaction status is obtained through a capacitive touchscreen or gesture recognition system; The data acquisition frequency is adjustable, and encrypted transmission is used to protect user privacy. The data is also preprocessed into a standardized format for subsequent analysis.
[0028] Step 7: Based on the user's eye gaze data and touch interaction status collected in Step 6, trigger the exclusive explanation content for the corresponding independent project. The explanation content is triggered according to the user's dwell time, with the triggering process using an event-driven mechanism to ensure that content playback is synchronized with user behavior, enhancing the immersive experience. When it is detected that a user's gaze at an independent project exceeds a threshold time or engages in touch interaction, the exclusive explanation content is immediately triggered. The dwell time trigger uses a multi-level judgment logic, including: a short dwell time triggers a brief version, a medium dwell time triggers a standard version, and a long dwell time triggers a detailed version. The triggering process prioritizes real-time user behavior to avoid content conflicts.
[0029] Compared with existing technologies, by integrating multi-dimensional behavioral perception of user eye gaze, dwell time, and touch interaction, the traditional one-way content push is transformed into two-way intelligent interaction, significantly improving the immersiveness and personalization of the user experience. By adopting an intelligent filtering algorithm based on attribute vector similarity and multi-time level content adaptation, the system solves the pain points of traditional tour guide systems, such as rigid content and inability to adapt to different tour paces. Through a combination of semantic analysis technology and manually customized indexing, the system constructs an evolving digital heritage knowledge base, which can ensure the accuracy of cultural dissemination and continuously optimize content strategies based on real-time feedback.
[0030] At other levels, in this embodiment, such as Figure 2 As shown, a multimodal cognitive system for experiencing digital heritage along the Silk Road includes: The region construction module 1 is used to construct a virtual gridded sensing region based on several real display areas. It divides the physical space into multiple grid units, each unit corresponding to a display area or independent project, locates and senses the user's position, and maps it to a virtual coordinate system to associate space with content. Attribute collection module 2 is used to collect the overall attributes of the corresponding display area and the attributes of individual items within the area; in the preset database, it indexes and filters related video, music and audio narration content based on these attributes; it supports manual customization of adding or modifying content, including updating attributes, adding new content or adjusting indexing rules; Content definition module 3 is used to define the overview and explanation content of the display area, and to define exclusive explanation content for individual projects within the display area; and to set different duration levels of explanation content based on different dwell times; Content deployment module 4 is used to deploy video, music or voice explanations of area overviews and exclusive items corresponding to different dwell times for each sensing area; Data acquisition module 5 is used to collect user eye gaze data, dwell time and touch interaction status, monitor user behavior in real time using visual sensors, touch screens or motion detection devices, and convert the data into a processable format; Content triggering module 6 is used to trigger exclusive explanatory content for independent projects based on the collected user eye gaze data and touch interaction status, and to trigger explanatory content of corresponding duration based on the collected user dwell time.
[0031] Content definition module 3, region construction module 1, attribute acquisition module 2, and content deployment module 4 are interconnected via a wireless network. Data acquisition module 5 is interconnected with content deployment module 4 and content triggering module 6 via a wireless network.
[0032] This embodiment provides a calculation formula for a similarity algorithm to filter the most relevant content, specifically: ; In the formula, Representative content item The relevance score indicates the degree of relevance; a higher score indicates a greater relevance. A vector representation of the overall attributes of the display area. Represents multimodal content items The vector representation of , Representative vector The Euclidean norm, Vector representations of attributes of independent projects. The weighting coefficient representing the overall attribute matching. The weighting coefficient represents the matching of attributes for independent projects.
[0033] Compared with existing technologies, by weighted fusion of overall and local attributes and combined with semantic vectorization calculation, adaptive and accurate matching of multi-dimensional content is achieved, which improves retrieval accuracy and scenario adaptability compared with traditional keyword matching technology.
[0034] In summary, this invention digitizes physical space by constructing a virtual grid-like sensing area, laying a technical foundation for accurately sensing user behavior. Combined with content attributes and a custom database, it ensures the richness, relevance, and scalability of explanatory resources. The dynamic content triggering mechanism based on real-time user behavior enables the system to intelligently determine user interests and needs, automatically matching and pushing video, music, or audio explanations of appropriate duration and theme, thereby providing a highly personalized immersive experience.
[0035] This invention not only greatly enhances the fun, interactivity, and educational depth of the visit, effectively extending the time users spend there, but also reduces the burden of human guides.
[0036] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-modal cognitive method for Silk Road digital heritage experience, characterized in that, The method comprises the following steps: Step 1: According to the physical exhibition space of the Silk Road Digital Heritage, it is divided into several logical units, each unit corresponds to an exhibition area, and a virtual grid sensing area for real-time monitoring of user location and behavior is constructed through a sensor network; Step 2: Obtain the overall attributes of each exhibition area and the attributes of the independent projects within the area through the data acquisition module, and store the attribute data in the preset database; Step 3: Based on the overall attributes and independent project attributes collected in step 2, index the associated images, music and voice commentary in the preset database, and select the most relevant content through a similarity algorithm; Step 4: Define the commentary content of the area overview for each exhibition area, and define the exclusive commentary content for the independent projects. According to the different possible residence time of the user, the commentary content is divided into multiple time length grades, each grade corresponds to different content depth and playing time; Step 5: Deploy the commentary content defined in step 4 to the corresponding virtual grid sensing area, including the image, music or voice commentary of the area overview and the independent projects, and set the trigger condition for each content; Step 6: Collect user's eye gaze data, residence time and touch interaction state of user behavior data in real time through sensors and interactive devices; Step 7: Trigger the exclusive commentary content of the corresponding independent project according to the user's eye gaze data and touch interaction state collected in step 6, and trigger the commentary content of the corresponding time length grade according to the user's residence time.
2. The multi-modal cognitive method for Silk Road digital heritage experience according to claim 1, wherein, The construction process of the virtual grid sensing area in step 1 includes: using a wireless sensor network or a camera array to cover the physical exhibition area, dividing the area into uniform or non-uniform grid units, giving each grid unit a unique identifier, and performing spatial mapping. The grid sensing area supports dynamic adjustment.
3. The multi-modal cognitive method for Silk Road digital heritage experience according to claim 1, wherein, The collection of exhibition content attributes in step 2 includes: automatically obtaining the attributes of independent projects through radio frequency identification tags, two-dimensional codes or near field communication devices, and supplementing the overall attributes through a manual input interface; The collected attribute data includes text, image and audio formats, and a structured storage method is used to establish an attribute table in the preset database, and the table structure includes region ID, project ID, attribute type and attribute value.
4. The multi-modal cognitive method for Silk Road digital heritage experience according to claim 1, wherein, The content indexing and filtering process in step 3 is: Use natural language processing algorithms to analyze attribute data, generate keyword vectors, and match similarity with multi-modal content library; The filtering process is optimized based on user historical preferences or context environment; The manual customization function allows administrators to add new content or modify the association rules of existing content through a graphical user interface.
5. The multi-modal cognitive method for Silk Road digital heritage experience according to claim 1, wherein, The calculation formula for selecting the most relevant content in the similarity algorithm in step 3 is: ; wherein, a relevance score of the content item, a vector representation of the overall properties of the presentation area, a vector representation of the multi-modal content item, a Euclidean norm of the vector a vector representation of the independent item properties, a weight coefficient of the overall property match, a weight coefficient of the independent item property match. 6. The multi-modal cognitive method for Silk Road digital heritage experience according to claim 1, wherein, The setting logic of time length grade in step 4 includes: dividing the commentary content into a short version less than 1 minute, a standard version of 1-3 minutes and a detailed version more than 3 minutes according to the time length.
7. The multi-modal cognitive method for Silk Road digital heritage experience according to claim 1, wherein, The collection process of user behavior data in step 6 includes: Use infrared eye tracking devices to collect eye gaze data, including gaze points and gaze duration; Calculate the residence time through the position sensor and the timer; The touch interaction state is obtained through a capacitive touch screen or a gesture recognition system; The data acquisition frequency is adjustable, and encrypted transmission is adopted to protect user privacy.
8. The multi-modal cognitive method for Silk Road digital heritage experience according to claim 1, wherein, When it is detected that the user gazes at a certain independent item for more than a threshold time or performs touch interaction, the exclusive explanation content is triggered immediately in step 7, and the stay duration triggering adopts a multi-level judgment logic, including: a short stay triggers a brief version, a medium stay triggers a standard version, and a long stay triggers a detailed version.
9. A multi-modal cognitive system for Silk Road digital heritage experience, the method is based on the implementation system of a multi-modal cognitive method for Silk Road digital heritage experience according to any one of claims 1-8, characterized in that, It comprises: A region construction module (1) for constructing a virtual grid sensing region based on a plurality of display regions in reality, dividing the physical space into a plurality of grid units, each unit corresponding to a display region or an independent item, positioning and sensing the user's position, and mapping it to a virtual coordinate system to associate the space and the content; An attribute acquisition module (2) for acquiring the overall attributes of the corresponding display region and the attributes of the independent items in the region; in a preset database, based on these attributes, the index and screening of the associated image, music and voice explanation content are performed; Supporting manual customization to add or modify content, including updating attributes, adding new content or adjusting index rules; A content definition module (3) for defining the region overview explanation content of the display region and defining the exclusive explanation content of the independent items in the display region; according to different stay times, set explanation content of different time length levels; A content deployment module (4) for deploying the image, music or voice explanation content of the region overview and exclusive items corresponding to different stay times for each sensing region; A data acquisition module (5) for acquiring the user's eye gaze data, stay duration and touch interaction state, using visual sensors, touch screens or motion detection devices to monitor user behavior in real time, and converting the data into a processable format; A content triggering module (6) for triggering the exclusive explanation content of the independent items according to the user's eye gaze data and touch interaction state obtained by acquisition, and triggering the explanation content corresponding to the time length according to the user's stay time obtained by acquisition.
10. The multi-modal cognitive system for Silk Road digital heritage experience according to claim 9, wherein, The content definition module (3), the region construction module (1), the attribute acquisition module (2) and the content deployment module (4) are connected through a wireless network, and the data acquisition module (5) is connected with the content deployment module (4) and the content triggering module (6) through a wireless network.