Historical scenic spot multi-modal interaction system based on large language model
By using a multimodal interactive system based on a large language model, the problems of discontinuity in the cultural heritage and personalized content provision in historical scenic area guide systems have been solved. This has enabled personalized historical content delivery and immersive experiences, thereby improving the effectiveness of cultural heritage transmission and the quality of tourists' understanding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING BOYUN TECHNOLOGY CO LTD
- Filing Date
- 2025-09-26
- Publication Date
- 2026-05-19
AI Technical Summary
Existing historical site tour guide systems are unable to identify gaps in the cultural heritage, leading to cognitive breaks in historical narratives for visitors. They also lack personalized content tailored to individual visitor differences, multi-dimensional analysis, and intelligent semantic completion capabilities, resulting in inefficient information delivery and poor cultural heritage preservation.
The historical scenic area multimodal interactive system based on a large language model acquires visitor location and gaze focus data through a data acquisition module. Combined with multi-dimensional feature extraction and fault identification modules, it identifies fault areas in the cultural heritage. Furthermore, through a coefficient calculation module, it calculates semantic completion coefficients and contextual coherence coefficients to construct a multi-path reasoning model for cross-temporal and spatial historical narratives. This generates multimodal fusion interactive scenes, dynamically adjusts the narrative rhythm and information density, and achieves immersive delivery and personalized interpretation.
It has significantly improved the cultural heritage preservation effectiveness and visitor experience quality of historical sites, enabled personalized historical content delivery, facilitated coherent narratives across cultural heritage gaps, enhanced visitors' immersive experience and historical awareness, and improved information transmission efficiency and cultural heritage preservation effectiveness.
Smart Images

Figure CN121392199B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal interaction technology, and more specifically, to a multimodal interaction system for historical scenic spots based on a large language model. Background Technology
[0002] Current historical site tour guide systems suffer from numerous technical limitations, severely impacting visitor experience and cultural preservation. Traditional guide methods fail to identify gaps in the cultural heritage, often resulting in abrupt "cliff-like" experiences for visitors. For example, a sudden jump from Ming and Qing dynasty architecture to Yuan dynasty ruins fails to effectively present the historical evolution in between, causing cognitive dissonance. Existing systems employ a standardized, one-size-fits-all approach, neglecting individual visitor differences. This leads to history enthusiasts and general tourists in the same group receiving the same level of content, resulting in the former feeling superficial and the latter finding it obscure and difficult to understand. Tour guides or electronic guides often only provide surface-level information about artifacts, such as "when was this built?" and "who did it belong to?", lacking multi-dimensional analysis of architectural features, historical records, and cultural symbol systems, preventing visitors from gaining a deeper understanding of the cultural connotations behind the artifacts. The interactive methods are simplistic and rigid, mostly relying on preset voice triggers or fixed route guidance. They fail to dynamically adjust content based on visitor dwell time, focus, and emotional responses. For example, even if a visitor showing keen interest in a Buddhist sculpture lingers for a considerable time, the system still plays the content for a fixed duration before abruptly ending the session. Furthermore, the limitations of a single sensory channel lead to inefficient information transmission. Relying solely on voice or text fails to recreate the rich sensory experience of historical scenes, making it difficult for visitors to develop an immersive historical understanding. In explanations spanning dynastic changes or cultural transitions, the existing system lacks intelligent semantic completion capabilities and cannot construct a coherent narrative based on historical correlation gradients. This results in visitors acquiring only fragmented knowledge after visiting the entire scenic area, failing to form a systematic historical cognitive framework and significantly reducing the effectiveness of cultural transmission.
[0003] In view of this, the present invention proposes a multimodal interactive system for historical scenic spots based on a large language model to solve the above problems. Summary of the Invention
[0004] To overcome the aforementioned deficiencies of the prior art and to achieve the above objectives, the present invention provides the following technical solution: a multimodal interactive system for historical scenic spots based on a large language model, comprising:
[0005] The data acquisition module is used to acquire real-time location data and line-of-sight trajectory data of tourists in the historical scenic area, and to collect sensor data streams from tourists' mobile devices;
[0006] The interest analysis module is used to construct a dynamic distribution map of tourist interest points based on real-time location data and gaze trajectory data.
[0007] The feature extraction module is used to extract multi-dimensional feature data of various cultural relics and historical sites within the historical scenic area, including architectural structural features, historical document records, and cultural symbol features.
[0008] The fault identification module is used to identify fault areas in the cultural heritage lineage based on the historical correlation gradient between adjacent cultural relics and historical sites in multi-dimensional feature data.
[0009] The coefficient calculation module is used to calculate the semantic completion coefficient and context coherence coefficient of the large language model in the fault area based on the knowledge density mutation characteristics of the cultural heritage fault area.
[0010] The reasoning construction module is used to build a multi-path reasoning model for cross-temporal and spatial historical narratives based on semantic completion coefficients and contextual coherence coefficients.
[0011] The scene generation module is used to generate multimodal fusion interactive scenes based on the multipath reasoning model and the dynamic attention distribution map of tourists' points of interest.
[0012] The pattern recognition module is used to identify the characteristic interaction patterns generated by the coupling of multimodal fusion interaction scenarios and tourists' cognitive needs by monitoring the semantic features and emotional tendencies of tourists' voice input signals in real time.
[0013] The dynamic adjustment module is used to dynamically adjust the narrative rhythm and information density of multimodal fusion interaction scenarios based on the semantic offset and emotional response changes of the feature interaction patterns.
[0014] The immersive delivery module is used to control interactive devices based on the adjusted multimodal fusion interactive scene parameters, so as to realize the immersive delivery and personalized interpretation of historical and cultural knowledge.
[0015] The technical effects and advantages of this invention's multimodal interactive system for historical scenic spots based on a large language model are as follows:
[0016] This invention significantly enhances the cultural heritage preservation effectiveness and visitor experience quality of historical sites. By intelligently sensing visitors' interests, the system can tailor historical content for visitors with different knowledge backgrounds, enabling history enthusiasts to engage in in-depth exploration while providing engaging interpretations for ordinary tourists. Where historical gaps are common in traditional explanations, the system cleverly connects preceding and following cultural threads, allowing visitors to gain a coherent and complete historical picture from scattered site visits, as if experiencing the evolution of civilization over thousands of years. Visitors are no longer passive recipients of information but active participants in historical exploration. Every pause, gaze, and even emotional shift guides the system to adjust the narrative direction and pace, creating a personalized experience akin to having a private tutor. The multi-sensory immersive presentation brings ancient architecture to life, allowing visitors to simultaneously experience visual reconstruction, environmental sound effects, and contextual explanations, transforming abstract historical knowledge into intuitive and vivid scene experiences. The system precisely strikes a balance between historical rigor and engaging content, harmoniously coexisting profound cultural symbols with accessible interpretations, satisfying the needs of academic inquiry while maintaining popular appeal. Scenic area managers benefit from the systematic analysis of visitor behavior and interests, enabling them to optimize exhibition design and cultural resource allocation, thereby improving overall service quality. Most importantly, this invention allows silent cultural relics to "speak," conveying the essence of traditional culture in a way that aligns with modern cognitive habits. It builds a bridge for dialogue across time and space, transforming history from a cold past into a living heritage that resonates with contemporary life. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the multimodal interactive system for historical scenic spots based on a large language model according to the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] This application provides a multimodal interactive system for historical scenic spots based on a large language model. The execution entities of the system include, but are not limited to, historical scenic spot guide platforms, cultural heritage interpretation systems, intelligent tour guide systems, immersive experience platforms, and multimodal interactive devices, which can be regarded as general computing nodes of this application. The data processing platform includes, but is not limited to, at least one of the following: a tourist behavior analysis system, a cultural relic feature extraction system, and a historical knowledge reasoning system.
[0020] Please see Figure 1This invention provides a multimodal interactive system for historical scenic spots based on a large language model, including a data acquisition module, an interest analysis module, a feature extraction module, a fault identification module, a coefficient calculation module, an inference construction module, a scene generation module, a pattern recognition module, a dynamic adjustment module, and an immersion delivery module.
[0021] The data acquisition module is used to acquire real-time location data and line-of-sight trajectory data of tourists in the historical scenic area, and to collect sensor data streams from tourists' mobile devices;
[0022] The interest analysis module is used to construct a dynamic distribution map of tourist interest points based on real-time location data and gaze trajectory data.
[0023] The feature extraction module is used to extract multi-dimensional feature data of various cultural relics and historical sites within the historical scenic area, including architectural structural features, historical document records, and cultural symbol features.
[0024] The fault identification module is used to identify fault areas in the cultural heritage lineage based on the historical correlation gradient between adjacent cultural relics and historical sites in multi-dimensional feature data.
[0025] The coefficient calculation module is used to calculate the semantic completion coefficient and context coherence coefficient of the large language model in the fault area based on the knowledge density mutation characteristics of the cultural heritage fault area.
[0026] The reasoning construction module is used to build a multi-path reasoning model for cross-temporal and spatial historical narratives based on semantic completion coefficients and contextual coherence coefficients.
[0027] The scene generation module is used to generate multimodal fusion interactive scenes based on the multipath reasoning model and the dynamic attention distribution map of tourists' points of interest.
[0028] The pattern recognition module is used to identify the characteristic interaction patterns generated by the coupling of multimodal fusion interaction scenarios and tourists' cognitive needs by monitoring the semantic features and emotional tendencies of tourists' voice input signals in real time.
[0029] The dynamic adjustment module is used to dynamically adjust the narrative rhythm and information density of multimodal fusion interaction scenarios based on the semantic offset and emotional response changes of the feature interaction patterns.
[0030] The immersive delivery module is used to control interactive devices based on the adjusted multimodal fusion interactive scene parameters, so as to realize the immersive delivery and personalized interpretation of historical and cultural knowledge.
[0031] This invention achieves precise capture of tourists' interests by acquiring real-time location data and gaze trajectory data. Based on the historical correlation gradient between adjacent cultural relics and historical sites in multi-dimensional feature data, it identifies gaps in the cultural heritage, enabling the system to identify historical knowledge gaps. The semantic completion coefficient and contextual coherence coefficient of the large language model are calculated based on the knowledge density mutation characteristics of these gaps, improving the scientific rigor of the system's knowledge completion. A multi-path reasoning model for cross-temporal and spatial historical narrative is constructed based on the semantic completion coefficient and contextual coherence coefficient, making historical interpretation more coherent. Multi-modal fusion interactive scenarios are generated based on the multi-path reasoning model and the dynamic attention distribution map of tourists' points of interest, enhancing the system's personalized responsiveness and immersive experience. Real-time monitoring of tourists' voice input signals identifies characteristic interaction patterns and dynamically adjusts the narrative rhythm and information density, improving the system's adaptability and interaction quality.
[0032] During visits to historical sites, data acquisition devices such as eye-tracking devices, positioning sensors, and built-in sensors in mobile devices are used to collect various types of data from tourists in real time. Real-time location data reflects the specific location and movement trajectory of tourists within the historical site; gaze focus trajectory data reflects the object the tourist is looking at, the duration of their gaze, and the path of their gaze migration; sensor data streams include posture and behavioral data provided by the gyroscopes, accelerometers, and orientation sensors of the tourist's mobile devices.
[0033] It should be noted that the data acquisition devices adopt a distributed deployment method to ensure comprehensive coverage of all areas of the historical scenic spot. The data acquisition devices continuously collect data at each sampling moment to ensure the continuity and integrity of the data.
[0034] In one implementation of this invention, the sampling frequency is set to once every 0.5 seconds.
[0035] In one implementation of this invention, the collected data undergoes preprocessing steps such as denoising, outlier detection, and standardization to ensure data quality. Wavelet transform is used for denoising, the 3σ principle is used for outlier detection, and Z-score standardization is employed for data standardization. The specific methods used for these preprocessing steps are not described here, as they are all techniques well-known to those skilled in the art. Other data preprocessing algorithms may also be used, and are not limited thereto.
[0036] The following steps all use preprocessed real-time location data, line-of-sight focus trajectory data, and sensor data streams for analysis.
[0037] In this embodiment of the invention, the method for constructing a dynamic attention distribution map of tourist points of interest includes:
[0038] Based on the real-time location data of tourists' dwell time and movement speed, a heat map of tourists' dwell time in various areas of the historical scenic spot is determined.
[0039] Based on the distribution density of fixation points and fixation duration in the gaze focus trajectory data, a heatmap of tourists' visual attention is determined;
[0040] Based on the spatiotemporal overlay relationship between the dwell heatmap and the visual attention heatmap, the comprehensive attention index of each point of interest is calculated;
[0041] Based on the changing trend of the comprehensive attention index over time, a dynamic attention distribution map of tourist interest points is generated.
[0042] In this embodiment, the spatiotemporal characteristics of real-time tourist location data are first analyzed to calculate the distribution of tourist dwell time in each area. Clustering algorithms (such as K-means, DBSCAN, etc.) are applied to identify hotspots where tourists frequently linger. The characteristics of tourist movement speed changes are analyzed; slow movement or lingering indicates potential interest. Combining dwell time and movement speed characteristics, a dwell heatmap reflecting the density of tourists' physical presence is generated. This heatmap visually displays tourists' activity preferences within the historical scenic area. Next, the gaze focus trajectory data is processed to analyze the spatial distribution characteristics of gaze points. The density distribution of gaze points in each area is calculated; high-density areas indicate more visual attention. The duration of each gaze point is statistically analyzed; longer gazes indicate higher levels of interest. Combining gaze density and duration generates a visual attention heatmap, which reflects the distribution of tourists' visual attention. Finally, the dwell heatmap and the visual attention heatmap are overlaid in a spatiotemporal dimension. This study analyzes and designs a weighted fusion algorithm to scientifically integrate data from two different modalities: physical dwell time and visual attention. Considering the temporal synchronicity of the two heatmaps, it processes the matching relationship between physical location and visual focus within the same time period. A spatial mapping algorithm is applied to transform heatmaps from different coordinate systems to a unified reference system. Based on the overlay analysis results, a comprehensive attention index is calculated for various points of interest (such as cultural relics, buildings, and exhibition areas) within the historical scenic area. This index comprehensively considers both physical dwell time intensity and visual attention intensity. Finally, the study tracks the temporal evolution characteristics of the comprehensive attention index, analyzes the fluctuation and trend characteristics of attention over time, and applies time-series analysis methods (such as moving window analysis and trend decomposition) to capture the dynamic characteristics of attention changes. Based on the temporal evolution analysis results, a dynamic attention distribution map of tourist points of interest is constructed. This distribution map not only displays the spatial dimension of attention distribution but also includes temporal dimension change information, providing data support for subsequent personalized interactions.
[0043] In this embodiment of the invention, the method for extracting multi-dimensional feature data includes:
[0044] Using computer vision technology, architectural structural features, including geometric configuration, spatial layout and material properties, are extracted from image data of cultural relics and historical sites.
[0045] Natural language processing technology is used to extract features of historical documents from historical documents, including chronological information and the relationship between historical events and figures.
[0046] Using semiotic analysis, we extract cultural symbolic features from the decorative elements of cultural relics and historical sites, including symbolic meaning, aesthetic tradition, and religious connotation.
[0047] The architectural structural features, historical document features, and cultural symbol features are integrated into a multi-dimensional feature vector to form multi-dimensional feature data.
[0048] In this embodiment, computer vision technology is first used to process the image data of cultural relics and historical sites. Edge detection and contour extraction algorithms (such as Canny algorithm, Sobel algorithm, etc.) are used to identify the geometric contours and structural lines of the buildings. 3D reconstruction technology (such as SfM, MVS, etc.) is applied to reconstruct a three-dimensional model of the cultural relics and historical sites from multi-angle images. The spatial structure and geometric characteristics of the buildings are analyzed. Texture analysis algorithms are used to extract characteristic parameters of the building materials, such as roughness and reflectivity. Color distribution features are extracted from the images to reflect the visual appearance of the buildings. Through the above steps, a model containing geometric configuration and spatial layout is formed. The system first establishes a set of architectural structural features based on the site's layout and material properties. These features directly reflect the physical form of cultural relics and historical sites. Then, natural language processing techniques are applied to analyze historical documents related to these relics and historical sites. Named Entity Recognition (NER) technology is used to extract key chronological information from the text, such as dynasties and specific years. Event extraction techniques are applied to identify descriptions of historical events related to the relics and historical sites, such as construction, reconstruction, and war damage. Relationship extraction techniques are used to identify the connections between the relics and historical sites and historical figures, such as designers, builders, and owners. Finally, topic modeling techniques (such as LDA and BERT) are applied to mine... Based on the above analysis, the thematic content related to cultural relics and historical sites in the literature is used to form a set of historical document records containing chronological information, historical events, and related figures. These features reflect the historical background of cultural relics and historical sites. Next, semiotic analysis methods are used to study the decorative elements of cultural relics and historical sites, and a decorative symbol recognition model is established to automatically identify specific decorative elements on cultural relics and historical sites, such as dragon patterns, lotus flowers, and ruyi patterns. A symbol-meaning knowledge base is constructed, mapping the identified decorative elements to their cultural symbolic meanings. The compositional characteristics and aesthetic features of the decorative elements are analyzed to reflect the artistic style of a specific period. Religious images and symbols in the decorative elements are identified, and their religious connotations and spiritual connotations are analyzed. Through the above analysis, a set of cultural symbol features containing symbolic meanings, aesthetic traditions, and religious connotations is formed. These features reveal the cultural connotations of cultural relics and historical sites. Finally, the three sets of features are scientifically integrated, and a feature fusion framework is designed to unify features of different dimensions and scales into a standardized representation form. A feature vector space is constructed to ensure the comparability and compatibility of different types of features. The integrated feature vectors are used as multi-dimensional feature data. These data comprehensively describe the physical form, historical background, and cultural connotations of cultural relics and historical sites, providing a foundation for subsequent analysis.
[0049] In this embodiment of the invention, based on the historical correlation gradient between adjacent cultural relics and historical sites in multi-dimensional feature data, regions with gaps in cultural heritage transmission are identified, including:
[0050] The historical correlation value of each cultural relic and historical site in the multi-dimensional feature data is subjected to time-series weighting to obtain the weighted historical correlation value;
[0051] Calculate the difference in weighted historical correlation values between adjacent cultural relics and historical sites, and use it as the gradient of historical correlation between adjacent cultural relics and historical sites;
[0052] Adjacent cultural relics and historical sites whose historical correlation gradient is greater than a preset gradient threshold are recorded as candidate fault interfaces.
[0053] A similarity analysis of the cultural symbol features of cultural relics and historical sites on both sides of the candidate fault interface is conducted to obtain the cultural continuity index corresponding to the candidate fault interface.
[0054] Candidate fault interfaces with a cultural continuity index less than a preset continuity threshold are marked as fault regions in the cultural heritage lineage.
[0055] In this embodiment, the historical correlation value of each cultural relic is first calculated based on multi-dimensional feature data. The historical correlation value is a quantitative indicator that measures the degree of correlation between cultural relics and historical sites in the historical context. Taking into account factors such as the frequency of historical records, the importance of related historical events, and the influence of historical figures, a weighted model is designed to perform time-series weighting on the historical correlation value. Modern cultural relics are given lower weights, and ancient cultural relics are given higher weights to balance the recording bias caused by differences in time span. A sliding time window method is used to locally strengthen the correlation value of specific periods to highlight the cultural inheritance characteristics of specific historical stages. Through time-series weighting, a weighted historical correlation value with a greater sense of historical depth is obtained. Then, spatial modeling is performed on the cultural relics and historical sites in the scenic area to determine the adjacency relationships between them, including physical adjacency (close spatial distance) and logical adjacency (high thematic relevance). The difference in weighted historical correlation values between each pair of adjacent cultural relics and historical sites is calculated. This difference is the historical correlation gradient. The larger the gradient value, the more significant the change in the historical correlation between adjacent cultural relics and historical sites, which may indicate a discontinuity in cultural inheritance. Then, a preset gradient is set. A threshold, determined through historical data analysis and expert knowledge, is typically set to 1.5 times the standard deviation of historical correlation. Adjacent cultural relics and historical sites with a historical correlation gradient greater than the preset threshold are statistically analyzed. These pairs form candidate fault interfaces, indicating potential breaks in the cultural heritage. Further analysis of the cultural symbol characteristics of the cultural relics and historical sites on both sides of the candidate fault interface is conducted. Feature similarity algorithms (such as cosine similarity and Jaccard similarity coefficient) are applied to calculate the similarity of the cultural symbol characteristics between the two sides, constructing a semantic network of cultural symbols. The semantic connections between the cultural symbols of the two sides are analyzed. Based on the combined similarity calculation results, a cultural continuity index is obtained for each candidate fault interface. This index measures the continuity of cultural heritage on both sides of the fault. Finally, a preset continuity threshold is set. When the cultural continuity index is less than this threshold, it indicates a significant break in cultural heritage. Candidate fault interfaces with a cultural continuity index less than the preset continuity threshold are formally marked as cultural heritage fault regions. These regions are key breakpoints in the historical and cultural heritage process, requiring systematic knowledge supplementation and narrative connection.
[0056] In this embodiment of the invention, based on the abrupt changes in knowledge density in the discontinuity regions of cultural heritage, the semantic completion coefficient and contextual coherence coefficient of the large language model in the discontinuity regions are calculated, including:
[0057] The knowledge density values of cultural relics and historical sites on both sides of the fault area of cultural heritage are obtained and recorded as the first knowledge density value and the second knowledge density value, respectively.
[0058] Calculate the knowledge density mutation coefficient of the cultural heritage fault area based on the ratio of the first knowledge density value to the second knowledge density value.
[0059] Based on the knowledge density mutation coefficient and the semantic understanding depth of the large language model, the semantic completion coefficient of the large language model in the fault region is calculated.
[0060] The context coherence coefficient of the large language model in the fault region is calculated by multiplying the semantic completion coefficient and the knowledge density mutation coefficient, and combining the context window length of the large language model.
[0061] In this embodiment, the knowledge density of cultural relics and historical sites on both sides of the cultural heritage discontinuity area is first assessed. Knowledge density refers to the number of historical and cultural knowledge points contained in a unit space or time. Knowledge points are extracted and statistically analyzed for cultural relics and historical sites in the earlier period of the discontinuity area to calculate the first knowledge density value. Similarly, knowledge points are extracted and statistically analyzed for cultural relics and historical sites in the later period of the discontinuity area to calculate the second knowledge density value. Knowledge points include information on historical events, figures, systems, crafts, and ideas. Then, the ratio of the first knowledge density value to the second knowledge density value is calculated. This ratio reflects the relative change in knowledge density on both sides of the discontinuity. When the ratio is much greater than 1, it indicates a sharp decrease in knowledge density from the early to the late period, possibly indicating a lack of historical records. When the ratio is much less than 1, it indicates a sharp increase in knowledge density from the early to the late period, possibly indicating a historical abrupt change or cultural innovation. This ratio is defined as the knowledge density mutation coefficient to quantify the discontinuity of knowledge transmission in the discontinuity area. Next, the semantic understanding characteristics of the large language model are analyzed to assess the depth of semantic understanding of the large language model for specific historical periods and cultural fields. The semantic understanding depth is determined by the amount of pre-trained data, fine-tuning degree, and test performance of the model in the relevant domain. A mapping relationship between semantic understanding depth and knowledge density mutation coefficient is constructed, and a completion ability evaluation function is designed. This function takes semantic understanding depth and knowledge density mutation coefficient as input and outputs the semantic completion coefficient of the large language model in the gap region. The semantic completion coefficient reflects the model's ability to fill knowledge gaps, and its value range is usually between 0 and 1, where 0 indicates that it is completely impossible to complete and 1 indicates that it is perfectly completed. Finally, the context processing ability of the large language model is considered, and the context window length of the model is evaluated, that is, the length of the text sequence that the model can effectively process. The degree of context dependence of the semantic completion task is analyzed, and a context coherence evaluation function is designed. This function comprehensively considers the semantic completion coefficient, knowledge density mutation coefficient, and context window length to calculate the context coherence coefficient of the large language model in the gap region. The context coherence coefficient reflects the model's ability to establish a coherent narrative between the text before and after the gap, and its value range is usually between 0 and 1. These two coefficients will serve as key parameters for the subsequent construction of a multi-path reasoning model, guiding the system in knowledge completion and narrative connection in cultural gap regions.
[0062] In this embodiment of the invention, a multi-path reasoning model for cross-temporal and spatial historical narratives is constructed based on semantic completion coefficients and contextual coherence coefficients, including:
[0063] Based on multi-dimensional feature data, a spatiotemporal knowledge graph of historical scenic spots is constructed;
[0064] Mark the location of the discontinuity in the cultural heritage transmission line in the spatiotemporal knowledge graph, and generate a historical narrative path map containing the discontinuity area;
[0065] Based on the historical narrative path map, the information distribution ratio of the main narrative path and the branch narrative path in each fault region of the large language model is calculated.
[0066] Based on the information allocation ratio, semantic completion coefficient, and context coherence coefficient, a multi-path reasoning model for cross-temporal and spatial historical narratives is constructed. The multi-path reasoning model includes the causal relationship distribution of historical events and the superimposed distribution of time clues.
[0067] In this embodiment, a spatiotemporal knowledge graph of the historical scenic area is first constructed based on multi-dimensional feature data. Cultural relics and historical sites are used as nodes in the knowledge graph, with each node containing key attributes from the multi-dimensional feature data. Multiple types of relationships are established between nodes, such as temporal relationships, spatial adjacency relationships, cultural heritage relationships, and historical event associations. Graph database technology is applied to store and manage the knowledge graph, ensuring efficient querying and reasoning. This knowledge graph forms the knowledge foundation for historical narratives. Then, in the constructed spatiotemporal knowledge graph, the locations of previously identified cultural heritage discontinuities are clearly marked. These discontinuities are special structures in the knowledge graph, representing potential breaks in the historical narrative. Based on the network structure of the knowledge graph, path planning algorithms (such as Dijkstra's algorithm and A* algorithm) are used to generate multiple possible narrative paths traversing the entire historical scenic area. These paths cross discontinuities, forming a historical narrative path map containing these discontinuities. The path map displays various possible historical narrative clues. Next, for each discontinuity, the completion capability of the large language model is analyzed. Based on the previously calculated semantic completion coefficients, the model's performance in that discontinuity region is evaluated. To ensure the reliability of a coherent narrative, this paper distinguishes between main and secondary narrative paths based on the structural characteristics of the narrative path. The main path typically connects important historical nodes, while secondary paths provide supplementary and expanded perspectives. An information allocation algorithm is designed to calculate the optimal information allocation ratio between the main and secondary narratives in each fault region of the large language model. Regions with high semantic completion coefficients can be allocated more main information, while regions with low semantic completion coefficients should increase the proportion of secondary information to balance the reliability and richness of the narrative. Finally, the results of the above analysis are integrated to construct a multi-path reasoning model for cross-temporal and spatial historical narratives. In this model, semantic completion coefficients are incorporated to guide the knowledge filling strategy in fault regions, contextual coherence coefficients are integrated to ensure narrative coherence across faults, and information allocation ratios are applied to balance the content weight of the main and secondary narratives. The model pays special attention to the causal relationship distribution of historical events, enhancing the explanatory power of historical narratives through causal reasoning. At the same time, a timeline overlay distribution is constructed to present the historical evolution process at different time scales. The final multi-path reasoning model can provide tourists with a rich, coherent, and personalized historical narrative experience while ensuring historical accuracy.
[0068] In this embodiment of the invention, a multimodal fusion interaction scenario is generated based on a multi-path inference model and a dynamic attention distribution map of tourist points of interest, including:
[0069] Based on the dynamic attention distribution map of tourists' points of interest, the weight of interactive content demand in each area of the historical scenic spot is determined.
[0070] Based on the multi-path reasoning model, the initial presentation sequence and media combination of multimodal content that meet the interactive content requirements weight are calculated.
[0071] The initial presentation sequence is temporally orchestrated to generate a temporally orchestrated multimodal content stream;
[0072] The multimodal content stream, arranged chronologically, is mapped onto the physical space of historical scenic spots to generate multimodal fusion interactive scenes. The information transmission continuity of multimodal fusion interactive scenes is enhanced in areas where the cultural heritage is interrupted.
[0073] In this embodiment, the spatial distribution characteristics of the dynamic attention distribution map of tourist points of interest are first analyzed to identify high-attention areas, which are where tourist interests are concentrated. The temporal variation characteristics of the attention distribution are then analyzed to identify the migration trend of the focus of attention. Combining spatial distribution and temporal trends, interactive content demand weights are assigned to each area within the historical scenic area. High-attention areas receive higher weights, indicating a need for richer and more in-depth interactive content. These weights reflect the personalized content needs of different areas. Then, based on a multi-path reasoning model, historical narrative content related to each area is extracted, including main narrative and sub-narrative, taking into account regional... The interactive content demand weights are determined to ascertain the level of detail and depth of the content. A multimodal expression strategy is designed to transform historical narrative content into various media formats, including text, images, audio, video, 3D models, and AR / VR scenes. Based on content characteristics and expressive effects, the optimal media combination is determined; for example, high-definition images are suitable for details of cultural relics, narrative videos are suitable for historical events, and interactive 3D models are suitable for cultural symbols. Based on the above analysis, an initial presentation sequence is calculated, including the sequential arrangement of content units and the selection of media formats. This sequence is the preliminary organizational form of the multimodal content. Then, the initial presentation sequence is optimized through temporal arrangement, considering the continuity of the narrative. To ensure consistency and a smooth logical transition between content units, attention is paid to the rationality of knowledge progression, moving from simple to complex and from concrete to abstract. An adaptive branching structure is constructed to adjust the narrative path according to changes in visitor interests, with particular emphasis on handling transitions in areas of cultural heritage discontinuity. Through the semantic completion and contextual coherence capabilities of a large language model, smooth transitions in these areas are achieved. After chronological arrangement, a multimodal content flow with a well-structured timeline is formed. This content flow is dynamically generated and can respond in real-time to changes in visitor interests. Finally, the multimodal content flow is mapped onto the physical space of the historical site using smart projection, AR glasses, etc. Intelligent audio guides, interactive screens, and other devices present multimodal content, establishing a precise correspondence between virtual content and physical space, achieving seamless integration of the physical and virtual worlds. Special attention is paid to information transmission in areas where cultural heritage is disrupted, deploying richer interactive methods in these areas to enhance the immersiveness and coherence of the content. Ultimately, this generates multimodal fusion interactive scenarios that integrate physical space and virtual content. These scenarios respect the physical environment of historical sites while enhancing the immersive experience for visitors through multimodal technology. Particularly in areas where cultural heritage is disrupted, technology bridges the gaps in historical narratives, achieving a coherent and engaging cultural heritage.
[0074] In this embodiment of the invention, by real-time monitoring of the semantic features and emotional tendencies of tourists' voice input signals, the characteristic interaction patterns generated by the coupling of multimodal fusion interaction scenarios and tourists' cognitive needs are identified, including:
[0075] Semantic analysis is performed on the tourist's voice input signal to extract the core semantic feature vector of the voice input signal;
[0076] Identify the frequency distribution of keywords related to history and culture in the core semantic feature vector, and denote them as candidate interaction intentions;
[0077] Calculate the semantic similarity between the candidate interaction intent and the current narrative theme of the multimodal fusion interaction scene. The semantic similarity is determined by the cosine similarity algorithm.
[0078] Candidate interaction intentions with semantic similarity greater than a preset similarity threshold are marked as feature interaction patterns.
[0079] In this embodiment, the voice input signals of tourists during the interaction process are first collected. Speech recognition technology is used to convert the voice signals into text data to ensure recognition accuracy, especially for historical proper nouns. Natural language processing is then performed on the recognized text, including basic processing such as word segmentation, part-of-speech tagging, and named entity recognition. Semantic analysis technology is applied to extract key semantic components from the text, such as keywords, action words, and modifiers. Vectorization methods (such as Word2Vec and BERT) are used to convert the extracted semantic components into computable feature vectors, obtaining feature vectors representing the core semantics of the tourist input. Then, within these core semantic feature vectors, semantic elements related to history and culture are identified. A history and culture domain dictionary is established, containing professional terms related to historical periods, figures, events, architecture, and cultural relics. The domain dictionary is used to filter and weight the words in the feature vectors, calculate the frequency distribution of history and culture-related words, analyze the clustering characteristics of the word frequency distribution, identify semantic focal points, and then... The associated semantic elements are organized into candidate interaction intentions, which represent the potential information needs or interests of tourists. Next, the narrative theme currently presented in the multimodal fusion interaction scenario is obtained. The theme is typically represented as a semantic vector or a set of keywords. The cosine similarity algorithm is used to calculate the semantic similarity between the candidate interaction intentions and the current narrative theme. The cosine similarity formula is the dot product of two vectors divided by the product of their magnitudes, with values ranging from -1 to 1. Values closer to 1 indicate greater semantic similarity. A preset similarity threshold is set, typically between 0.6 and 0.8. The threshold setting needs to balance the system's sensitivity and accuracy. Finally, candidate interaction intentions with semantic similarity greater than the preset similarity threshold are formally marked as feature interaction patterns. These interaction patterns reflect the coupling point between tourists' cognitive interests and the system's current narrative content, serving as an important basis for the system's dynamic adjustments. Based on the identified feature interaction patterns, the system will specifically adjust subsequent content presentation strategies to improve the personalization and satisfaction of the interactive experience.
[0080] In this embodiment of the invention, the narrative rhythm and information density of a multimodal fusion interaction scenario are dynamically adjusted based on the semantic offset and emotional response changes of the feature interaction pattern, including:
[0081] The deviation between the semantic center of the calculated feature interaction pattern and the preset narrative thread of the multimodal fusion interaction scene is denoted as semantic offset.
[0082] The rate of polarity change of sentiment words in the statistical feature interaction mode within a preset time window is denoted as sentiment response change.
[0083] Based on the semantic offset and changes in emotional response, determine the amount of narrative rhythm adjustment and information density adjustment in multimodal fusion interaction scenarios;
[0084] Based on the adjustment of narrative rhythm and information density, the content presentation speed and knowledge point distribution density of multimodal fusion interactive scenarios are updated.
[0085] In this embodiment, the semantic structure of the characteristic interaction patterns is first analyzed, its core concepts and relationships are extracted, a semantic network representation is constructed, the centrality of the semantic network is calculated, and the semantic centroid is determined. This centroid represents the core focus of the interaction pattern. Simultaneously, the pre-defined narrative thread of the multimodal fusion interaction scenario is analyzed. The pre-defined narrative thread is the content presentation thread planned by the system, usually designed based on the logical structure of historical narrative. Semantic distance measurement methods (such as Wasserstein distance, Jensen-Shannon divergence, etc.) are used to calculate the degree of deviation between the semantic centroid and the pre-defined narrative thread. This degree of deviation is the semantic offset; the larger the offset, the stronger the visitor's semantic offset. The greater the difference between the focus of interest and the system's preset narrative direction, the more effective the analysis becomes. Then, the analysis examines the emotional vocabulary in the characteristic interaction patterns. Using an emotional dictionary and sentiment analysis model, the analysis identifies the emotional vocabulary and its polarity (positive, negative, neutral) in the interaction content. A preset time window is set, typically the interaction process over the last few minutes. The trend of change in the polarity of emotional vocabulary within the time window is calculated, including changes in polarity direction and intensity. The rate of polarity change is defined as the change in emotional response. This change reflects the dynamic characteristics of the tourist's emotional state; a rapid increase in positive polarity indicates increased interest, while a negative change indicates possible disappointment or dissatisfaction. Finally, based on semantic offset and emotional response changes, an adaptive adjustment mechanism is designed. The overall strategy involves constructing an adjustment decision matrix. Different combinations of offsets and emotional changes correspond to different adjustment strategies: Large offset + positive emotional change: moderately adjust the theme to follow visitor interests and increase the depth of related content; Large offset + negative emotional change: significantly adjust the narrative direction and quickly respond to changes in visitor interests; Small offset + positive emotional change: maintain the current narrative route and appropriately increase the richness of details; Small offset + negative emotional change: keep the theme unchanged, reduce content complexity, and increase interest. Based on the decision matrix, specific adjustments to narrative rhythm and information density are determined. Narrative rhythm adjustments control the speed and pace of content presentation, while information density adjustments control... The system determines the quantity and complexity of knowledge points within a given time frame. Based on this determined adjustment, it updates the parameter settings of the multimodal interactive scenario in real time, adjusting the content presentation speed, accelerating or slowing down the narrative pace to adapt to changes in tourists' acceptance and interests. It also optimizes the density of knowledge point distribution, increases or decreases the amount of information per unit time, adjusts the depth of content, and modulates the level of professionalism according to tourists' cognitive needs, dynamically balancing entertainment and knowledge. The system continuously monitors the adjustment effects, forming a closed-loop feedback mechanism to continuously optimize the interactive experience. Through this dynamic adjustment mechanism, the system can respond in real time to changes in tourists' cognitive needs, providing personalized and engaging historical and cultural experiences.
[0086] In this embodiment of the invention, the method for calculating semantic completion coefficients includes:
[0087] Based on the knowledge density mutation coefficient, the difficulty level of knowledge completion in the fault region is determined.
[0088] Analyze the coverage of pre-trained data of the large language model in relevant historical domains to determine the model's knowledge base score;
[0089] Assess the contextual understanding ability of large language models and determine their reasoning ability scores;
[0090] The semantic completion coefficient is calculated based on a weighted combination of the knowledge completion difficulty level, the model's basic knowledge score, and the model's reasoning ability score.
[0091] In this embodiment, the difficulty of knowledge completion in the fault region is first assessed based on the previously calculated knowledge density mutation coefficient. A difficulty grading standard is designed, for example, dividing the knowledge density mutation coefficient into four levels: low mutation (1.0-1.5), medium mutation (1.5-3.0), high mutation (3.0-10.0), and extremely high mutation (>10.0). Each level corresponds to a different level of knowledge completion difficulty. The characteristics of the historical period and cultural field involved in the fault region are assessed. Special periods (such as dynastic changes and periods of war) and special fields (such as religious changes and technological innovations) usually have higher completion difficulties. Taking all the above factors into account, the knowledge completion difficulty level of the fault region is determined. This level reflects the objective difficulty of filling the gaps in historical knowledge. Then, the knowledge reserves of the large language model in the relevant historical fields are analyzed, the document coverage of the corresponding historical period and cultural field in the model's pre-training data is assessed, and the model's mastery of key knowledge points such as relevant historical figures, events, and systems is calculated. The accuracy of the model's knowledge in this field is assessed through knowledge testing methods. Based on the above assessment results, the model's knowledge base score is determined. This score reflects the model's level of knowledge accumulation in the relevant historical field. Then, the large language model is evaluated. The model's contextual understanding and reasoning abilities were tested, including its performance on historical text comprehension tasks such as temporal relationship inference and causal relationship analysis. The model's ability to connect and interpret historical events was assessed, such as inferring possible historical development paths based on known historical facts. The model's ability to handle ambiguous or contradictory information was also tested, which is particularly important in historical interpretation. Based on the comprehensive test results, a reasoning ability score was determined, reflecting the model's level of ability to handle complex historical information. Finally, a formula for calculating the semantic completion coefficient was designed, which comprehensively considers the knowledge completion difficulty level (D) and the model's knowledge base. The basic score (K) and the model reasoning ability score (R) can be typically calculated using the formula: Semantic completion coefficient = (w1×K+w2×R) / (w3×D), where w1, w2, and w3 are weight parameters that are adjusted according to the specific application scenario. The calculated semantic completion coefficient is usually normalized to the range of 0 to 1. The closer the value is to 1, the stronger the semantic completion ability of the model in the fault region. This coefficient will guide the system's knowledge completion strategy in the fault region. High coefficient regions can rely more on model-generated content, while low coefficient regions need to cite more specific historical materials or provide multiple possible interpretations.
[0092] In this embodiment of the invention, the structure of the multi-path reasoning model for cross-temporal and spatial historical narratives includes:
[0093] A node network layer based on a spatiotemporal knowledge graph represents historical entities and their relationships.
[0094] A fault transition layer constructed based on semantic completion coefficients is used to handle knowledge connections in fault areas of cultural heritage.
[0095] A narrative coherence layer built on context coherence coefficients ensures the logical fluency of cross-temporal and spatial narratives;
[0096] A content organization layer built on the information allocation ratio balances the content proportions of the main narrative and sub-narratives.
[0097] An explanation generation layer built upon the reasoning capabilities of a large language model provides multi-faceted interpretations of historical events.
[0098] In this embodiment, a node network layer is first constructed. This layer forms the basic structure of the model. Based on a spatiotemporal knowledge graph, it includes all cultural relics, historical sites, historical events, and figures within the historical scenic area as nodes. Multiple types of edges are established to connect related nodes, such as temporal, spatial, and causal relationship edges. Attribute features are assigned to each node, including key information from multi-dimensional feature data. Graph neural network technology is used to process information transmission and aggregation between nodes. This layer provides a structured representation of historical knowledge. Then, a fault transition layer is constructed. This layer specifically addresses the knowledge connection problem in areas of cultural heritage discontinuity. Special transition nodes are set in each fault region. As a bridge between knowledge on both sides of the fault line, this layer determines the knowledge generation strategy for transition nodes based on semantic completion coefficients. Regions with high completion coefficients use large language models to generate content, while regions with low completion coefficients employ multi-source verification and probabilistic reasoning methods. A smooth knowledge transition mechanism is designed to ensure natural connection between knowledge on both sides of the fault line. This layer addresses the knowledge discontinuity problem in historical narratives. Next, a narrative coherence layer is constructed. This layer ensures the logical coherence of historical narratives across time and space. Based on contextual coherence coefficients, a smooth transition mechanism for narrative paths is designed. Regions with high coherence coefficients use complex narrative structures, such as flashbacks and interjections, while regions with low coherence coefficients maintain simple linear narratives, establishing a time anchor system. The system helps tourists maintain a stable sense of time across different periods. It designs a narrative pacing control mechanism to adjust the level of detail in different historical stages, ensuring the smoothness and comprehensibility of the historical narrative. Next, a content organization layer is constructed. This layer organizes the main and sub-narratives according to information allocation ratios, designs a main content framework containing necessary historical skeletal information to ensure narrative integrity, and builds a sub-content library containing rich historical details, anecdotes, and extended knowledge. A dynamic content selection algorithm is designed to adjust the ratio of main to sub-narratives based on tourist interests and interactive behaviors. This layer ensures the historical narrative has both a clear main storyline and rich content. Finally, the system... The interpretation generation layer provides multi-faceted interpretations of historical events. Utilizing the reasoning capabilities of a large language model, it offers multiple possible perspectives for interpreting historical events. It designs a historical viewpoint balancing mechanism to ensure objectivity and comprehensiveness when presenting different historical viewpoints. It constructs an evidence support system to provide historical data for important historical viewpoints, thereby enhancing credibility. It also designs a reflection-promoting mechanism to encourage visitors to form their own thoughts on historical interpretations. This layer enriches the depth and intellectual content of historical narratives. The five layers work together to form a complete multi-path reasoning model for cross-temporal and spatial historical narratives. This model can not only handle the knowledge connection problem of cultural discontinuities but also provide a rich, coherent, and personalized historical narrative experience.
[0099] In this embodiment of the invention, the method for presenting a multimodal fusion interactive scene includes:
[0100] Based on the spatial characteristics of historical scenic spots and the distribution of cultural relics and historical sites, a three-dimensional digital model of the physical space is constructed.
[0101] Based on the media types and presentation requirements of multimodal content streams, a multi-channel perception and interaction mechanism is designed.
[0102] Based on the need for the integration of physical space and virtual content, determine the application strategy of augmented reality technology;
[0103] Based on the visitor's location and line of sight, the spatial mapping relationship of the content presentation is dynamically adjusted to achieve adaptive spatial narrative.
[0104] In this embodiment, a precise 3D spatial model of the historical scenic area is first constructed using technologies such as 3D laser scanning and photogrammetry to obtain high-precision spatial data. A 3D digital model incorporating elements such as terrain, architecture, and cultural relics is then built. The location and attributes of key cultural relics and historical sites are marked within the digital model, and a coordinate reference system for the physical space is established. This digital model forms the basis for the virtual content space mapping. Next, a multi-channel perception interaction mechanism is designed. Based on the characteristics of different media types in the multimodal content stream, corresponding perception channels are designed: a visual channel (including AR glasses, smart projectors, holographic displays, etc.) for presenting visual content such as images, videos, and 3D models; an auditory channel (including directional sound fields, spatial audio, etc.) for presenting auditory content such as narration, sound effects, and music); a tactile channel (including haptic feedback devices, interactive physical models, etc.) to enhance the physical perception experience; and an olfactory channel (deploying an environmental odor system in specific areas to recreate the odor environment of historical scenes). A collaborative mechanism between channels is designed to ensure the harmonious unity of multi-channel perception. This multi-channel design enriches the visitor's perceptual experience. Finally, the application strategy of augmented reality technology is determined, and the characteristics of cultural relics and historical sites in each area are analyzed. To meet the needs of safety and protection, suitable augmented reality technologies are selected, such as landmark AR, landmarkless AR, and spatial AR. A visually fusion effect is designed to ensure the natural integration of virtual content with the physical environment. An interactive gesture and voice command system is designed to enable natural interaction between visitors and virtual content. A dynamic loading mechanism for augmented reality content is established, adjusting AR content in real time based on visitor behavior and system response. This strategy allows virtual content to seamlessly integrate into the physical environment. Finally, based on real-time location data and gaze trajectory data of visitors, a dynamic spatial mapping system is established to map content units in the multimodal content stream to specific locations in the physical space. A perspective adaptation mechanism is designed to adjust the presentation of content according to the visitor's gaze direction. A distance adaptation mechanism is designed to adjust the level of detail of content based on the distance between the visitor and the artifact. A spatial narrative sequence is constructed, allowing the narrative content to unfold naturally with the visitor's spatial movement. Special attention is paid to the spatial treatment of areas with breaks in the cultural heritage, strengthening the spatial continuity of virtual content in these areas. Through this adaptive spatial narrative mechanism, the system can provide visitors with an immersive and personalized historical and cultural experience, naturally integrating the transmission of historical knowledge with the spatial experience.
[0105] In this embodiment of the invention, the method for tracking the dynamic attention distribution map of tourist points of interest includes:
[0106] Based on real-time location data and line-of-sight trajectory data, an initial distribution model of tourist attention is established.
[0107] Analyze the sensor data streams from tourists' mobile devices to identify tourists' behavioral patterns and attention characteristics;
[0108] Update the dynamic attention distribution map of tourist points of interest based on behavioral patterns and attention characteristics;
[0109] Based on the updated distribution map, the trend of tourist interest migration can be predicted, enabling dynamic tracking of attention distribution.
[0110] In this embodiment, firstly, based on the collected real-time location data and gaze focus trajectory data, an initial distribution model of tourist attention is constructed. Spatial clustering algorithms (such as DBSCAN) are used to identify hotspots in the location data, which typically represent areas of interest to tourists. Gaze point density analysis is used to identify visual attention hotspots, which represent areas where tourists' visual attention is concentrated. The location hotspots and visual hotspots are fused to establish the initial attention distribution model, which provides the basic spatial distribution of tourist interests. Then, the rich sensor data streams provided by the tourist's mobile devices are analyzed, including data from accelerometers, gyroscopes, and orientation sensors, to identify typical tourist behavior patterns, such as fast walking, slow browsing, pausing to observe, and bending over to look. The device orientation and posture changes are analyzed to infer the direction and intensity of tourist attention. Combined with voice input and touch operation data, the tourist's focus of interest is further confirmed. Through these behavioral and attentional characteristics, a more comprehensive understanding of the tourist's interest state can be achieved. Then, based on the identified behavioral patterns and attentional characteristics, the dynamic... The system dynamically updates the distribution map of interest points, designs an attention decay model to reflect the natural decline of tourist interest over time, and an attention enhancement model to reflect the reinforcing effect of repeated attention or in-depth interaction on interest. Using a Bayesian update method, newly observed behavioral evidence is combined with the previous distribution to form an updated distribution. A time window mechanism is applied to ensure that the attention distribution reflects the current state of interest rather than historical accumulation. Through this dynamic update mechanism, the system can maintain a distribution map that reflects changes in tourist interest in real time. Finally, based on the updated attention distribution map, the system predicts possible future trends in tourist interest, analyzes the historical path characteristics of attention point migration, identifies typical interest evolution patterns, and uses sequence prediction models (such as Markov models, RNNs, etc.) to predict the next possible focus of attention. Considering the relevance of scenic area content, the system predicts related interests that may be triggered by the current interest. Through this prediction mechanism, the system can prepare potentially needed content in advance, achieving dynamic tracking and prediction of the distribution of interest point attention, laying the foundation for providing forward-looking personalized services.
[0111] This invention addresses the issue of connecting knowledge gaps in historical sites by leveraging the semantic completion and contextual coherence capabilities of a large language model in areas where cultural heritage is fragmented. It achieves a balance between personalized historical narratives and in-depth knowledge by combining a dynamic distribution map of visitor interest points with a multi-path reasoning model. Furthermore, it enhances the system's adaptability by dynamically adjusting the narrative rhythm and information density of interactive scenarios through real-time monitoring of the semantic features and emotional tendencies of visitor voice input. Finally, it enables immersive delivery and personalized interpretation of historical and cultural knowledge through the construction of multimodal fusion interactive scenarios, significantly improving the quality of the visitor experience and the effectiveness of cultural dissemination in historical sites.
[0112] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0113] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
Claims
1. A multimodal interactive system for historical scenic spots based on a large language model, characterized in that: include: The data acquisition module is used to acquire real-time location data and line-of-sight trajectory data of tourists in the historical scenic area, and to collect sensor data streams from tourists' mobile devices; The interest analysis module is used to construct a dynamic distribution map of tourist interest points based on real-time location data and gaze trajectory data. The feature extraction module is used to extract multi-dimensional feature data of various cultural relics and historical sites within the historical scenic area, including architectural structural features, historical document records, and cultural symbol features. The fault identification module is used to identify fault areas in the cultural heritage lineage based on the historical correlation gradient between adjacent cultural relics and historical sites in multi-dimensional feature data, including: The historical correlation value of each cultural relic and historical site in the multi-dimensional feature data is subjected to time-series weighting to obtain a weighted historical correlation value; Calculate the difference in the weighted historical correlation values between adjacent cultural relics and historical sites, and use it as the historical correlation gradient between adjacent cultural relics and historical sites; Adjacent cultural relics and historical sites whose historical correlation gradient is greater than a preset gradient threshold are recorded as candidate fault interfaces. A similarity analysis is performed on the cultural symbol features of cultural relics and historical sites on both sides of the candidate fault interface to obtain the cultural continuity index corresponding to the candidate fault interface. Candidate fault interfaces with a cultural continuity index less than a preset continuity threshold are marked as fault regions of the cultural heritage lineage. The coefficient calculation module is used to calculate the semantic completion coefficient and contextual coherence coefficient of the large language model in the fault regions based on the abrupt changes in knowledge density in the cultural heritage fault regions. This includes: The knowledge density values of cultural relics and historical sites on both sides of the fault area of the cultural heritage are obtained and recorded as the first knowledge density value and the second knowledge density value, respectively. Calculate the knowledge density mutation coefficient of the cultural heritage fault region based on the ratio of the first knowledge density value to the second knowledge density value; Based on the knowledge density mutation coefficient and the semantic understanding depth of the large language model, the semantic completion coefficient of the large language model in the fault region is calculated. Based on the product of the semantic completion coefficient and the knowledge density mutation coefficient, and combined with the context window length of the large language model, the context coherence coefficient of the large language model in the fault region is calculated. The reasoning construction module is used to build a multi-path reasoning model for cross-temporal and spatial historical narratives based on semantic completion coefficients and contextual coherence coefficients. The scene generation module is used to generate multimodal fusion interactive scenes based on the multipath reasoning model and the dynamic attention distribution map of tourists' points of interest. The pattern recognition module is used to identify the characteristic interaction patterns generated by the coupling of multimodal fusion interaction scenarios and tourists' cognitive needs by monitoring the semantic features and emotional tendencies of tourists' voice input signals in real time. The dynamic adjustment module is used to dynamically adjust the narrative rhythm and information density of multimodal fusion interaction scenarios based on the semantic offset and emotional response changes of the feature interaction patterns. The immersive delivery module is used to control interactive devices based on the adjusted multimodal fusion interactive scene parameters, so as to realize the immersive delivery and personalized interpretation of historical and cultural knowledge.
2. The multimodal interactive system for historical scenic spots based on a large language model according to claim 1, characterized in that, The multi-path reasoning model for cross-temporal and spatial historical narratives, based on semantic completion coefficients and contextual coherence coefficients, includes: Based on the multi-dimensional feature data, a spatiotemporal knowledge graph of historical scenic spots is constructed; Mark the location of the cultural heritage discontinuity region in the spatiotemporal knowledge graph, and generate a historical narrative path map containing the discontinuity region; Based on the historical narrative path map, the information allocation ratio of the main narrative path and the branch narrative path in each fault region of the large language model is calculated. Based on the information allocation ratio, the semantic completion coefficient, and the context coherence coefficient, a multi-path reasoning model for cross-temporal and spatial historical narrative is constructed. The multi-path reasoning model includes the causal relationship distribution of historical events and the superimposed distribution of time clues.
3. The multimodal interactive system for historical scenic spots based on a large language model according to claim 1, characterized in that, The process of generating a multimodal fusion interaction scenario based on a multi-path inference model and a dynamic attention distribution map of tourist interest points includes: Based on the dynamic attention distribution map of tourist interest points, the interactive content demand weights for each area within the historical scenic area are determined. Based on the multi-path reasoning model, calculate the initial presentation sequence and media combination of multimodal content that meet the interactive content requirement weights; The initial presentation sequence is temporally arranged to generate a temporally arranged multimodal content stream; The multimodal content stream, after being arranged in a specific time sequence, is mapped onto the physical space of the historical scenic area to generate the multimodal fusion interactive scene. The information transmission continuity of the multimodal fusion interactive scene is enhanced in the area where the cultural heritage is interrupted.
4. The multimodal interactive system for historical scenic spots based on a large language model according to claim 1, characterized in that, The method involves real-time monitoring of the semantic features and emotional tendencies of tourists' voice input signals to identify characteristic interaction patterns resulting from the coupling of multimodal fusion interaction scenarios and tourists' cognitive needs, including: Semantic parsing is performed on the tourist's voice input signal to extract the core semantic feature vector of the voice input signal; The frequency distribution of keywords related to history and culture is identified in the core semantic feature vector and denoted as candidate interaction intentions. Calculate the semantic similarity between the candidate interaction intent and the current narrative theme of the multimodal fusion interaction scene, and determine the semantic similarity using a cosine similarity algorithm; Candidate interaction intentions with semantic similarity greater than a preset similarity threshold are marked as the feature interaction patterns.
5. The multimodal interactive system for historical scenic spots based on a large language model according to claim 1, characterized in that, The method of dynamically adjusting the narrative rhythm and information density of multimodal fusion interaction scenarios based on semantic offsets and changes in emotional responses according to feature interaction patterns includes: The deviation between the semantic center of the feature interaction mode and the preset narrative thread of the multimodal fusion interaction scene is calculated and denoted as the semantic offset. The rate of polarity change of emotional words in the aforementioned characteristic interaction pattern within a preset time window is statistically analyzed and denoted as the emotional response change. Based on the semantic offset and the change in emotional response, determine the narrative rhythm adjustment amount and information density adjustment amount of the multimodal fusion interaction scene; Based on the narrative rhythm adjustment amount and the information density adjustment amount, update the content presentation speed and knowledge point distribution density of the multimodal fusion interactive scene.