Traditional culture digital output and interaction method and system based on virtual reality technology
By constructing a multimodal traditional culture knowledge graph and processing multi-sensor data, the system dynamically identifies user intentions and emotions, generates adaptive cultural contexts, and outputs multi-sensory feedback. This solves the problem of insufficient data integration and feedback in virtual reality technology, and enhances the realism and naturalness of user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-07
AI Technical Summary
Existing virtual reality technology suffers from insufficient multimodal data integration, poor user interaction intelligence, and low realism of multisensory feedback in the digital output and interaction of traditional culture. This results in fragmented cultural elements, unnatural interaction, high response latency, and monotonous feedback, making it difficult to adapt to personalized needs.
By constructing a multimodal traditional culture knowledge graph, combining it with user interaction data collected from multiple sensors for noise reduction and normalization, the system dynamically identifies user intentions and emotions, generates adaptive cultural contexts, and outputs multi-sensory feedback content, thereby achieving personalized adjustments and physical simulation.
It has improved the unified representation and accessibility of cultural elements, enhanced the realism and immersion of the user experience, reduced the cost of manually customized content, and realized the intelligent and automated transmission of culture.
Smart Images

Figure CN121807152A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of virtual reality interaction, and in particular to a digital output, interaction method and system for traditional culture based on virtual reality technology. Background Technology
[0002] With the rapid development of virtual reality technology and the increasing demand for digital preservation of traditional culture, virtual reality systems are gradually becoming more widely used in museums, education, and cultural heritage displays. These systems aim to enhance users' understanding and participation in traditional culture through immersive experiences, such as showcasing the details of cultural relics through 3D modeling or simulating historical scenes using interactive devices.
[0003] However, existing technologies have significant shortcomings in multimodal data integration, intelligent user interaction, and the realism of multisensory feedback. Traditional virtual reality systems often rely on a single data source (such as static 3D models or simple text descriptions), lacking effective integration of 3D data of cultural relics, historical documents, and intangible cultural heritage process sequences. This results in fragmented cultural elements, making it difficult to support deep semantic connections and dynamic context generation. Furthermore, user interaction processing is often based on preset rules or simple sensor inputs, failing to fully utilize real-time data collected by multiple sensors (such as inertial measurement units, eye trackers, and electromyography sensors). Insufficient denoising and normalization processing leads to large intent recognition errors and high response delays, affecting the naturalness and fluency of the interaction. Regarding feedback output, existing systems primarily focus on visual rendering, neglecting multisensory synchronization such as touch and hearing. Low physical simulation accuracy fails to realistically reproduce the material characteristics or behavioral patterns of cultural elements, resulting in a lack of immersion and realism in the user experience. Simultaneously, these systems typically rely on manual content configuration, which is inefficient and costly, making it difficult to adapt to personalized needs. Therefore, existing virtual reality technology still has considerable room for improvement in the digital output and interaction of traditional culture, and there is an urgent need for an innovative solution that can achieve intelligent fusion of multimodal data, accurate intent recognition, and adaptive multisensory feedback. Summary of the Invention
[0004] To address the aforementioned technical issues, this application provides a method and system for digital output and interaction of traditional culture based on virtual reality technology.
[0005] The above-mentioned objective of this application is achieved through the following technical solution:
[0006] A method for digital output and interaction of traditional culture based on virtual reality technology, the method comprising the following steps:
[0007] Acquire 3D data of cultural relics, historical documents, and sequences of intangible cultural heritage processes to construct a multimodal knowledge graph of traditional culture;
[0008] User interaction data is collected by multiple sensors, and the user interaction data is denoised and normalized based on a noise index optimization algorithm to generate processed user interaction data.
[0009] Based on the user interaction data processed by the multimodal traditional culture knowledge graph, the user's cultural operation intentions and emotional tendencies are dynamically identified.
[0010] An adaptive cultural context is generated based on the user's cultural operational intentions and emotional tendencies, and multi-sensory feedback content is output based on the adaptive cultural context.
[0011] By adopting the above technical solutions, the digital output and interaction methods of traditional culture based on virtual reality technology construct a multimodal traditional culture knowledge graph by acquiring 3D data of cultural relics, historical documents, and intangible cultural heritage process sequences. This solves the problems of data heterogeneity and semantic fragmentation in the digital presentation of traditional culture, and improves the unified representation and accessibility of cultural elements. By collecting user interaction data from multiple sensors and performing noise reduction and normalization processing, sensor noise and temporal misalignment are eliminated, enhancing the reliability and consistency of the data and providing clean input for intent recognition. The knowledge graph analyzes user interaction data to dynamically identify the user's cultural operation intent and emotional tendency, achieving real-time adaptation of the interaction process and avoiding the problems of sluggish response or misjudgment in traditional virtual reality systems. Adaptive cultural contexts are generated based on user intent and emotion, and multi-sensory feedback content is output. Through personalized adjustments and physical simulation, the realism and immersion of the user experience are significantly improved, while reducing the cost of manually customized content, realizing the intelligent and automated transmission of culture.
[0012] In a preferred embodiment, this application can be further configured as follows: the acquisition of three-dimensional data of cultural relics, historical document texts, and intangible cultural heritage process sequences, and the construction of a multimodal traditional cultural knowledge graph, specifically includes:
[0013] Collect high-precision three-dimensional point cloud data and multi-angle texture images of cultural relics, and generate geometric and visual feature vectors of cultural relics based on the three-dimensional point cloud data and multi-angle texture images;
[0014] Cultural elements are extracted from historical documents to describe text, and semantic feature vectors are generated based on the cultural element description text.
[0015] Based on the geometric features, visual feature vectors, and semantic feature vectors of the cultural relics, a fusion network is trained using a cross-modal alignment loss function to generate a unified multimodal traditional cultural knowledge graph.
[0016] Each cultural element in the knowledge graph is labeled with a dynamic contextual tag, including applicable scenarios, user cognitive levels, and interaction trigger conditions.
[0017] By adopting the above technical solutions, the specific steps for constructing a multimodal traditional culture knowledge graph are further defined. These steps include collecting high-precision 3D point cloud data and multi-angle texture images of cultural relics to generate geometric and visual feature vectors, extracting text from historical documents to generate semantic feature vectors, and training a fusion network through a cross-modal alignment loss function. This solves the problem of difficulty in aligning multimodal data (such as geometric, visual, and textual data), improving the semantic consistency and completeness of the knowledge graph. Furthermore, dynamic contextual labels (such as applicable scenarios and interaction triggering conditions) are added to cultural elements in the knowledge graph, enhancing the graph's contextual awareness capabilities. This allows virtual reality systems to dynamically adjust their output according to different scenarios (such as education or entertainment), avoiding content rigidity and thus improving the flexibility and practicality of cultural interaction, supporting more precise personalized user needs.
[0018] In a preferred embodiment, this application can be further configured as follows: the step of collecting user interaction data based on multiple sensors, and performing denoising and normalization processing on the user interaction data based on a noise index optimization algorithm to generate processed user interaction data specifically includes:
[0019] The user's hand gesture trajectory is collected by an inertial measurement unit and an infrared camera, and the Pearson correlation between acceleration and angular velocity is calculated to eliminate jitter noise.
[0020] An eye tracker is used to record the coordinates of the user's gaze point, and the attention distribution matrix is calculated by combining the positions of virtual objects in the cultural context.
[0021] The hand muscle signals are acquired by electromyography (EMG) sensors, the fine manipulation force is analyzed, and dynamic time warping is performed and matched with a standard cultural movement library.
[0022] The multi-source user interaction data was time-stamped and spatial coordinate system unified to generate processed user interaction data.
[0023] By adopting the above technical solutions, the system collects hand gesture trajectories using an inertial measurement unit and an infrared camera, calculates the Pearson correlation between acceleration and angular velocity to eliminate jitter noise, records the gaze point coordinates using an eye tracker to calculate the attention distribution matrix, and analyzes the fine-grained operational force using an electromyography sensor and matches it with a standard action library. This solves the problems of high noise and synchronization difficulties in multi-sensor data, thus improving data quality. Furthermore, the system synchronizes timestamps and unifies spatial coordinates for multi-source data, ensuring data consistency and integrability, avoiding error accumulation in interactive recognition, thereby improving the accuracy and real-time performance of user intent analysis, reducing the risk of system misjudgment, and enhancing the smoothness and reliability of virtual reality interaction.
[0024] In a preferred embodiment, this application can be further configured such that: the step of dynamically identifying the user's cultural operational intentions and emotional tendencies based on the user interaction data parsed and processed from the multimodal traditional culture knowledge graph also includes:
[0025] User behavior features are extracted based on the processed user interaction data, and temporal action patterns are extracted based on the user behavior features. The similarity between these patterns and standard action templates in the cultural knowledge graph is then calculated.
[0026] Multimodal intent features are generated by weighted fusion of semantic labels of the user's eye-tracking focus area, gesture operation object identifiers, and voice keywords through an attention mechanism.
[0027] By using a graph neural network to traverse the cultural knowledge graph, cultural connotation nodes associated with intent features are retrieved, and user intent classification is output.
[0028] By adopting the above technical solution, temporal action patterns are extracted based on user behavior features and their similarity to standard templates in the knowledge graph is calculated. Multimodal intent features are generated by weighted fusion of eye-tracking focus semantic tags, gesture object identifiers, and voice keywords through an attention mechanism. Furthermore, a graph neural network is used to traverse the knowledge graph to retrieve cultural connotation nodes. This solves the problems of single-modal data limitations and insufficient semantic depth in traditional intent recognition, improving the comprehensiveness and accuracy of recognition. Through multimodal fusion and graph query, fine-grained classification of user intent (such as learning or entertainment intent) is achieved, avoiding misjudgments caused by simple rules, thereby enhancing the system's adaptability and supporting a more natural and intelligent cultural interaction experience.
[0029] In a preferred embodiment, this application can be further configured as follows: generating an adaptive cultural context based on the user's cultural operational intentions and emotional tendencies, and outputting multi-sensory feedback content based on the adaptive cultural context, specifically including:
[0030] Based on the user's cultural operational intent, suitable narrative elements are retrieved from the cultural knowledge graph, and a personalized narration script is generated based on the narrative elements.
[0031] The depth of the output content is adjusted based on the user's cognitive level, and the structural details of the cultural relics are dynamically displayed through the 3D model decomposition module.
[0032] It integrates a physics engine to simulate real-life behaviors of cultural elements, maps surface friction data through the vibration frequency of the haptic handle, and generates multi-sensory synchronous feedback in conjunction with environmental sound effects.
[0033] By adopting the above technical solution, a personalized narration script is generated by retrieving narrative elements from a knowledge graph based on the user's cultural operational intentions. The depth of the output content is adjusted based on the user's cognitive level, and the structural details of cultural relics are dynamically displayed through a 3D model decomposition module. The physical engine is integrated to simulate the real behavior of cultural elements, and the vibration frequency of the haptic handle is mapped to the surface friction data in conjunction with environmental sound effects. This solves the problems of monotonous content and lack of realism in virtual reality feedback, and improves the coordination and realism of multi-sensory synchronous output. Through adaptive adjustment and physical simulation, dynamic matching between feedback content and user state is achieved, avoiding the experience gap caused by oversimplification or overcomplication. This improves user participation and satisfaction, reduces the labor cost of content production, and supports large-scale personalized applications.
[0034] In a preferred embodiment, this application can be further configured as follows: Based on the geometric features, visual feature vectors, and semantic feature vectors of the cultural relics, a fusion network is trained using a cross-modal alignment loss function to generate a unified multimodal traditional cultural knowledge graph, specifically including:
[0035] Geometric features and text features are normalized separately, and the inter-modal alignment loss is calculated using cosine similarity.
[0036] By introducing an adversarial training strategy, the fusion network can distinguish between correct and incorrect associations of cultural elements, enhance semantic consistency, add time dimension labels to the knowledge graph, support the diachronic evolution display of cultural elements, and generate a unified multimodal traditional cultural knowledge graph.
[0037] By adopting the above technical solutions, geometric and textual features are normalized separately, and the intermodal alignment loss is calculated using cosine similarity. An adversarial training strategy is introduced to enhance semantic consistency, and a time dimension label is added to the knowledge graph. This solves the problems of inconsistent feature scales and association errors, and improves the robustness and timeliness of the graph. It also supports the display of the diachronic evolution of cultural elements, enabling the system to dynamically reflect cultural changes and avoid the limitations of static knowledge bases. This enhances the historical accuracy and educational value of cultural output and supports long-term cultural research and public education applications.
[0038] In a preferred embodiment, this application can be further configured such that the fusion physics engine simulates the realistic behavior of cultural elements, specifically including:
[0039] Collect material light reflection data and surface friction tactile data of intangible cultural heritage objects to generate a physical property mapping table;
[0040] The interaction between ambient light and objects is simulated using a high dynamic range lighting model, and the reflection intensity distribution is calculated in real time.
[0041] The spring-mass model is used to simulate the impact vibration response, and the corresponding sound effect is generated by combining it with the acoustic model.
[0042] The physical rendering results are comprehensively evaluated. If the material matching index is lower than the threshold, a high-precision model recalculation is triggered.
[0043] By adopting the above technical solutions, material light reflection data and surface friction and tactile data of intangible cultural heritage objects are collected to generate a physical property mapping table. A high dynamic range lighting model is used to simulate ambient light interaction and calculate the reflection intensity distribution in real time. A spring-mass model is used to simulate the impact vibration response and combined with an acoustic model to generate sound effects. The physical rendering results are comprehensively evaluated and recalculated, which solves the problems of unrealistic physical simulation and resource waste in virtual reality, and improves the accuracy and efficiency of simulation. Through high-quality rendering and dynamic optimization, the realism and consistency of multi-sensory feedback are ensured, avoiding the sense of incongruity in the user experience, thereby improving the credibility and attractiveness of the virtual reality system and supporting high-end cultural display and training scenarios.
[0044] Secondly, the above-mentioned inventive objective of this application is achieved through the following technical solutions:
[0045] A digital output and interaction system for traditional culture based on virtual reality technology, the system comprising:
[0046] The cultural knowledge graph construction module is used to acquire three-dimensional data of cultural relics, historical documents and intangible cultural heritage process sequences to construct a multimodal traditional cultural knowledge graph.
[0047] The user interaction data processing module is used to collect user interaction data based on multiple sensors, perform noise reduction and normalization processing on the user interaction data based on a noise index optimization algorithm, and generate processed user interaction data.
[0048] The user intent recognition module is used to dynamically identify the user's cultural operation intent and emotional tendency based on the user interaction data parsed and processed by the multimodal traditional culture knowledge graph.
[0049] The interactive content output module is used to generate an adaptive cultural context based on the user's cultural operational intentions and emotional tendencies, and output multi-sensory feedback content based on the adaptive cultural context.
[0050] By adopting the above technical solution,
[0051] Thirdly, the above-mentioned objectives of this application are achieved through the following technical solutions:
[0052] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described traditional cultural digital output and interaction method based on virtual reality technology.
[0053] Fourthly, the above-mentioned objectives of this application are achieved through the following technical solutions:
[0054] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method for digital output and interaction of traditional culture based on virtual reality technology.
[0055] In summary, this application includes at least one of the following beneficial technical effects:
[0056] 1. A digital output and interaction method for traditional culture based on virtual reality technology constructs a multimodal traditional culture knowledge graph by acquiring 3D data of cultural relics, historical documents, and intangible cultural heritage process sequences. This solves the problems of data heterogeneity and semantic fragmentation in the digital presentation of traditional culture, improving the unified representation and accessibility of cultural elements. By collecting user interaction data from multiple sensors and performing denoising and normalization processing, sensor noise and temporal misalignment are eliminated, enhancing data reliability and consistency and providing clean input for intent recognition. The knowledge graph analyzes user interaction data to dynamically identify users' cultural operation intentions and emotional tendencies, achieving real-time adaptation of the interaction process and avoiding the problems of sluggish response or misjudgment in traditional virtual reality systems. Adaptive cultural contexts are generated based on user intentions and emotions, and multi-sensory feedback content is output. Through personalized adjustments and physical simulation, the realism and immersion of the user experience are significantly improved, while reducing the cost of manually customized content, realizing the intelligent and automated transmission of culture.
[0057] 2. By collecting gesture trajectories through an inertial measurement unit and an infrared camera, and calculating the Pearson correlation between acceleration and angular velocity to eliminate jitter noise, and using an eye tracker to record the coordinates of the gaze point to calculate the attention distribution matrix, and analyzing the fine-grained operational force through an electromyography sensor and matching it with a standard action library, the problems of high noise and synchronization difficulties in multi-sensor data are solved, thus improving data quality. Time stamp synchronization and spatial coordinate system unification of multi-source data ensure data consistency and integrability, avoiding error accumulation in interactive recognition, thereby improving the accuracy and real-time performance of user intent analysis, reducing the risk of system misjudgment, and enhancing the smoothness and reliability of virtual reality interaction.
[0058] 3. Based on user behavior features, temporal action patterns are extracted and their similarity to standard templates in the knowledge graph is calculated. Multimodal intent features are generated by weighted fusion of eye-tracking focus semantic labels, gesture object identifiers, and voice keywords through an attention mechanism. Furthermore, a graph neural network is used to traverse the knowledge graph to retrieve cultural connotation nodes, solving the problems of single-modal data limitations and insufficient semantic depth in traditional intent recognition, thus improving the comprehensiveness and accuracy of recognition. Through multimodal fusion and graph query, fine-grained classification of user intent (such as learning or entertainment intent) is achieved, avoiding misjudgments caused by simple rules, thereby enhancing the system's adaptability and supporting a more natural and intelligent cultural interaction experience.
[0059] 4. Based on the user's cultural operational intent, narrative elements are retrieved from the knowledge graph to generate personalized narration scripts. The depth of the output content is adjusted based on the user's cognitive level, and the structural details of cultural relics are dynamically displayed through a 3D model decomposition module. The physical engine is integrated to simulate the real behavior of cultural elements, and the vibration frequency of the haptic handle is mapped to the surface friction data in conjunction with environmental sound effects. This solves the problems of monotonous content and lack of realism in virtual reality feedback, and improves the coordination and realism of multi-sensory synchronous output. Through adaptive adjustment and physical simulation, dynamic matching between feedback content and user state is achieved, avoiding the experience gap caused by oversimplification or overcomplication. This improves user participation and satisfaction, reduces the labor cost of content production, and supports large-scale personalized applications. Attached Figure Description
[0060] Figure 1 This is a flowchart of a traditional culture digital output and interaction method based on virtual reality technology in one embodiment of this application;
[0061] Figure 2 This is a flowchart illustrating the implementation of step S10 in a traditional culture digital output and interaction method based on virtual reality technology in one embodiment of this application.
[0062] Figure 3 This is a flowchart illustrating the implementation of step S20 in a traditional culture digital output and interaction method based on virtual reality technology in one embodiment of this application.
[0063] Figure 4 This is a flowchart illustrating the implementation of step S30 in a traditional culture digital output and interaction method based on virtual reality technology in one embodiment of this application.
[0064] Figure 5 This is a flowchart illustrating the implementation of step S40 in a traditional culture digital output and interaction method based on virtual reality technology in one embodiment of this application.
[0065] Figure 6This is a flowchart illustrating the implementation of step S13 in a traditional culture digital output and interaction method based on virtual reality technology in one embodiment of this application.
[0066] Figure 7 This is a flowchart illustrating the implementation of step S43 in a traditional culture digital output and interaction method based on virtual reality technology in one embodiment of this application.
[0067] Figure 8 This is a principle block diagram of a traditional culture digital output and interaction system based on virtual reality technology in one embodiment of this application;
[0068] Figure 9 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0069] The present application will be further described in detail below with reference to the accompanying drawings.
[0070] In one embodiment, such as Figure 1 As shown, this application discloses a method for digital output and interaction of traditional culture based on virtual reality technology, which specifically includes the following steps:
[0071] S10: Acquire 3D data of cultural relics, historical documents and intangible cultural heritage process sequences, and construct a multimodal traditional cultural knowledge graph.
[0072] In this embodiment, constructing a multimodal traditional culture knowledge graph is a fundamental step in the method. It aims to integrate multi-source data, including cultural relics, documents, and intangible cultural heritage crafts, to form a structured knowledge base to support subsequent virtual reality interactions. The construction of the knowledge graph involves multimodal data fusion, which can uniformly represent the geometric, visual, and semantic features of cultural elements, providing data support for dynamic context generation. Multimodal data includes 3D point clouds, images, text, and sequence data. These data sources are heterogeneous, requiring cross-modal alignment to ensure consistency.
[0073] Specifically, 3D data of cultural relics is acquired using high-precision 3D scanning equipment (such as laser scanners) to generate point cloud data of the relics' surfaces; historical document texts are extracted from digital archives or databases, involving natural language processing techniques to analyze cultural element descriptions; and intangible cultural heritage process sequences are recorded using motion capture systems to form time-series data. After data acquisition, preprocessing is performed: denoising and registration of point cloud data, word segmentation and entity recognition of text data, and standardization of sequence data. Then, feature extraction algorithms (such as convolutional neural networks for image features and word embedding models for text features) are used to generate feature vectors. Cross-modal fusion is achieved by training a fusion network that uses an alignment loss function (such as cosine similarity) to minimize the differences between features from different modalities, ultimately generating a unified knowledge graph. Each node in the knowledge graph represents a cultural element (such as a cultural relic or a process step) and is labeled with dynamic contextual tags, such as applicable scenarios (such as education or entertainment), user cognitive levels (such as beginners or experts), and interaction triggering conditions (such as user gestures triggering explanations). Through this construction method, knowledge graphs can dynamically adapt to different interaction scenarios, enhancing the personalization and realism of virtual reality experiences.
[0074] S20: Based on the collection of user interaction data by multiple sensors, the user interaction data is denoised and normalized based on the noise index optimization algorithm to generate processed user interaction data.
[0075] In this embodiment, user interaction data includes multi-source information such as gestures, eye movements, and electromyography signals. A noise index optimization algorithm is used to eliminate random errors in the data, such as sensor jitter or environmental interference, to ensure data quality. Denoising and normalization processing improves the reliability and consistency of the data, providing clean input for intent recognition.
[0076] Specifically, the multi-sensor system includes an inertial measurement unit (IMU) for gesture tracking, an eye tracker for gaze recording, and an electromyography (EMG) sensor for muscle signal acquisition. During data acquisition, real-time synchronization is performed: timestamps are used to align different sensor streams, avoiding timing misalignments. Noise optimization involves calculating Pearson correlations (such as the correlation between acceleration and angular velocity) to identify jitter noise, and then removing it using filtering algorithms (such as Kalman filtering). Normalization scales the data to a standard range (such as 0-1) for easier subsequent processing. For example, after denoising, gesture trajectory data is matched with a standard action library using dynamic time warping (DTW) to calculate similarity. Through multi-sensor fusion and signal processing, the accuracy and usability of user interaction data are enhanced, supporting precise intent analysis.
[0077] S30: Based on the user interaction data after parsing and processing the multimodal traditional culture knowledge graph, dynamically identify the user's cultural operation intentions and emotional tendencies.
[0078] In this embodiment, intent recognition is the core component. By analyzing the correlation between user interaction data (such as gestures and eye movements) and the knowledge graph, the system infers the cultural operation the user intends to perform (such as rotating an artifact or querying history). Sentiment analysis assesses the user's emotions (such as curiosity or confusion) and quantifies them based on behavioral characteristics (such as operation speed or gaze duration). Dynamic recognition allows the system to adapt to changes in interaction in real time.
[0079] Specifically, the parsing process first extracts features from the processed interaction data: for example, extracting temporal patterns (such as speed changes) from gesture trajectories and gaze duration from eye-tracking data. Then, machine learning models (such as recurrent neural networks, RNNs) are used to analyze these features, calculating similarity to standard action templates in the knowledge graph (e.g., using cosine similarity). An attention mechanism weightedly fuses multimodal features (such as eye-tracking focus semantic labels and gesture object identifiers) to generate an intent feature vector. Finally, a graph neural network (GNN) traverses the knowledge graph, retrieving nodes related to the intent features and outputting an intent classification (e.g., "learning intent" or "entertainment intent") and a sentiment score (e.g., positive or negative). This step, through multimodal fusion and graph querying, achieves high-precision user state recognition and supports personalized feedback.
[0080] S40: Generate an adaptive cultural context based on the user's cultural operational intentions and emotional tendencies, and output multi-sensory feedback content based on the adaptive cultural context.
[0081] In this embodiment, the adaptive cultural context is a dynamically generated virtual environment setting based on user intent and emotions, such as adjusting the narration content or scene difficulty; multi-sensory feedback includes visual, auditory, and tactile outputs, aiming to provide an immersive experience. Context generation ensures that the content matches the user's state, enhancing the naturalness of the interaction.
[0082] Specifically, the generation process first retrieves narrative elements (such as the story of an artifact or the steps of its production) from a knowledge graph, and then selects appropriate content based on intent (such as providing detailed explanations for learning intentions). Next, the output depth is adjusted according to emotional state: for example, interactive elements are added for positive emotions, while content is simplified for negative emotions. A 3D model decomposition module dynamically displays the artifact's structure (e.g., layered display), and a physics engine simulates realistic behaviors (e.g., object collision effects). Haptic handles map surface friction data through vibration frequency (e.g., rough textures correspond to high-frequency vibrations), and environmental sound effects are generated based on the scene (e.g., ancient music background). Feedback content is rendered in real-time to ensure multi-sensory synchronization. This step, through dynamic adjustment and physical simulation, achieves highly adaptive virtual reality output, enhancing the realism and engagement of the user experience.
[0083] In this embodiment, the traditional culture digital output and interaction method based on virtual reality technology constructs a multimodal traditional culture knowledge graph by acquiring 3D data of cultural relics, historical documents, and intangible cultural heritage process sequences. This solves the problems of data heterogeneity and semantic fragmentation in the digital presentation of traditional culture, and improves the unified representation and accessibility of cultural elements. By collecting user interaction data from multiple sensors and performing noise reduction and normalization processing, sensor noise and temporal misalignment are eliminated, enhancing data reliability and consistency and providing clean input for intent recognition. The method dynamically identifies users' cultural operation intentions and emotional tendencies based on the knowledge graph's analysis of user interaction data, achieving real-time adaptation of the interaction process and avoiding the problems of sluggish response or misjudgment in traditional virtual reality systems. Adaptive cultural contexts are generated based on user intentions and emotions, and multi-sensory feedback content is output. Through personalized adjustments and physical simulation, the realism and immersion of the user experience are significantly improved, while reducing the cost of manually customized content, thus realizing the intelligent and automated transmission of culture.
[0084] In one embodiment, such as Figure 2 As shown, in step S10, which involves acquiring 3D data of cultural relics, historical documents, and sequences of intangible cultural heritage processes, and constructing a multimodal traditional cultural knowledge graph, the specific steps include:
[0085] S11: Collect high-precision three-dimensional point cloud data and multi-angle texture images of cultural relics, and generate geometric features and visual feature vectors of cultural relics based on the three-dimensional point cloud data and multi-angle texture images.
[0086] In this embodiment, 3D point cloud data refers to the set of 3D coordinates of points on the surface of an artifact obtained through laser scanning or photogrammetry, which can accurately represent the geometric shape of the artifact; multi-angle texture images are images of the artifact's surface taken from different perspectives, used to capture color and texture details. Geometric feature vectors quantify the shape attributes of the artifact (such as curvature and volume), while visual feature vectors represent appearance attributes (such as color distribution and texture patterns). These features form the basis for knowledge graph construction.
[0087] Specifically, high-resolution 3D scanners (such as structured light scanners) are used to acquire point cloud data of cultural relics, ensuring full coverage of the relics and avoiding occlusion during the scanning process. The point cloud data undergoes preprocessing, including denoising (using statistical filtering to remove outliers) and registration (aligning multi-view point clouds to a unified coordinate system). Then, geometric features are extracted based on the point cloud data: for example, principal component analysis (PCA) is used to calculate the main axes and dimensions of the relics, or shape descriptors (such as spherical harmonic coefficients) are used to generate geometric feature vectors. Simultaneously, multi-angle texture images are acquired using a digital camera. After image calibration and distortion correction, visual feature vectors are extracted using feature extraction algorithms (such as SIFT or deep learning models such as ResNet). Visual feature vectors may include local texture descriptors or global color histograms. Finally, the geometric and visual feature vectors are normalized and concatenated to form a multimodal feature representation of the relics, used for subsequent knowledge graph fusion. Automated acquisition and feature extraction reduce manual intervention and improve the efficiency and accuracy of data processing.
[0088] S12: Extract cultural elements from historical documents to describe the text, and generate a semantic feature vector based on the cultural elements describing the text.
[0089] In this embodiment, historical document texts include cultural descriptions in ancient books, archives, or research papers. Semantic feature vectors refer to text converted into numerical vectors using natural language processing technology, which can capture the semantic meaning of cultural elements, such as the historical background of cultural relics or the symbolic meaning of craftsmanship. Semantic feature vectors are easy to align with geometric and visual features, enabling multimodal knowledge fusion.
[0090] Specifically, historical documents are obtained from digital databases (such as museum databases or academic resources), and the text data may contain unstructured descriptions. First, text preprocessing is performed, including word segmentation, stop word removal, and entity recognition (such as identifying key entities like artifact names and dates). Then, semantic models are used to generate feature vectors: for example, word embedding techniques (such as Word2Vec or BERT) are used to map the text into a low-dimensional vector space. For long texts, attention mechanisms are used to focus on key phrases (such as "bronze ornamentation"), generating condensed semantic feature vectors. After vector generation, normalization is performed to ensure consistency with the scale of other modal features. Through automated text analysis, the semantic richness of the knowledge graph is enhanced, supporting contextual understanding in subsequent intent recognition.
[0091] S13: Based on the geometric features, visual feature vectors, and semantic feature vectors of the cultural relics, a fusion network is trained using a cross-modal alignment loss function to generate a unified multimodal traditional cultural knowledge graph.
[0092] In this embodiment, the cross-modal alignment loss function is an optimization objective used to reduce the differences between features from different modalities, ensuring the consistency of geometric, visual, and semantic information in the knowledge graph. The fusion network is typically a deep learning model, such as a multimodal transformer, that learns to map heterogeneous features to a unified space, generating a structured knowledge graph.
[0093] Specifically, the training process of the fusion network includes: First, geometric, visual, and semantic feature vectors are input into different branches of the network, and each branch performs feature encoding (e.g., using fully connected layers). Then, cross-modal alignment loss is calculated, for example, by comparing the similarity between geometric and textual features using cosine similarity; the loss value is used for backpropagation to update network parameters. Adversarial training strategies may be introduced to enable the network to distinguish between correct and incorrect associations, enhancing semantic consistency. After training, the network outputs a graph-structured knowledge graph, where nodes represent cultural elements and edges represent relationships between elements (e.g., the association between cultural relics and crafts). The knowledge graph also adds a time dimension label to support the diachronic evolution of cultural elements (e.g., changes in cultural relics across different dynasties). This step, through end-to-end training, achieves seamless integration of multimodal data, improving the accuracy and dynamism of the knowledge graph.
[0094] S14: Label each cultural element in the knowledge graph with dynamic contextual tags, including applicable scenarios, user cognitive levels, and interaction triggering conditions.
[0095] In this embodiment, dynamic context tags are metadata used to describe the usage context of cultural elements in virtual reality, such as applicable scenarios (e.g., exhibitions or games), user cognitive level (e.g., novice users need to simplify content), and interaction triggering conditions (e.g., voice commands trigger animations).
[0096] Specifically, tagging is based on rule engines or machine learning models: for example, clustering algorithms are used to automatically assign cognitive level tags based on user behavior data. The tags are stored in graph nodes in the form of key-value pairs and can be retrieved in real time during interaction. Through context awareness, the responsiveness and personalization of the virtual reality system are improved.
[0097] In one embodiment, such as Figure 3 As shown, in step S20, user interaction data is collected based on multiple sensors, and the user interaction data is denoised and normalized based on a noise index optimization algorithm to generate processed user interaction data. Specifically, this includes:
[0098] S21: The user's hand gesture trajectory is collected by the inertial measurement unit and infrared camera, the Pearson correlation between acceleration and angular velocity is calculated, and jitter noise is eliminated.
[0099] In this embodiment, the inertial measurement unit (IMU) provides acceleration and angular velocity data, the infrared camera captures gesture images, and Pearson correlation is used to quantify the linear relationship in motion. High correlation indicates noise (such as hand tremors) and needs to be removed to clean up the data.
[0100] Specifically, IMU and camera data are transmitted in real time via Bluetooth or wired connection, and the Pearson correlation coefficient is calculated: r = cov(acceleration, angular velocity) / (σ acc * σ ang If |r| > the threshold (e.g., 0.8), it is identified as a jitter segment and removed using median filtering. This statistical method improves the smoothness of the gesture data.
[0101] S22: Use an eye tracker to record the coordinates of the user's gaze point, and combine this with the position of virtual objects in the cultural scene to calculate the attention distribution matrix.
[0102] In this embodiment, the gaze point coordinates represent the user's visual focus, and the attention distribution matrix is a probability matrix that describes the degree of user attention to different objects in the virtual scene, used to infer user interests.
[0103] Specifically, after eye-tracking data calibration, the (x, y) coordinates of the gaze point are obtained; the position of the virtual object is extracted from the scene model. An attention matrix is calculated: Euclidean distance or cosine similarity is used to compare the gaze point with the object position, generating attention weights. Through visual analysis, the understanding of user behavior is enhanced.
[0104] S23: Acquire hand muscle signals through electromyography (EMG) sensors, analyze fine manipulation force, and perform dynamic time warping matching with a standard cultural movement library.
[0105] In this embodiment, electromyography (EMG) signals reflect muscle electrical activity, and the intensity of the analysis can identify the fineness of the user's operation (such as light touch or heavy pressure); Dynamic Time Warping (DTW) is a sequence alignment algorithm used to compare the similarity between the user's actions and a standard template.
[0106] Specifically, after amplification and filtering, electromyographic signals are used to extract features (such as root mean square value) to represent force; the DTW algorithm calculates the minimum path distance between the user's action sequence and the standard library template, matches similar actions, and improves the accuracy of operation recognition through time series analysis.
[0107] S24: The multi-source user interaction data was time-stamped and spatial coordinate system unified to generate processed user interaction data.
[0108] In this embodiment, timestamp synchronization ensures that all sensor data are aligned in time, and the spatial coordinate system transforms different sensor data to the same reference system (such as the world coordinate system) to avoid data conflicts.
[0109] Specifically, synchronization is achieved using the Network Time Protocol (NTP), and coordinate transformation is implemented using a homogeneous matrix. The processed data is stored in a structured format (such as JSON) for easy subsequent use. This step, through system integration, ensures data consistency.
[0110] In one embodiment, such as Figure 4 As shown, in step S30, which involves dynamically identifying the user's cultural operational intentions and emotional tendencies based on the user interaction data processed by the multimodal traditional culture knowledge graph, the specific steps include:
[0111] S31: Extract user behavior features based on the processed user interaction data, extract temporal action patterns based on the user behavior features, and calculate the similarity between the patterns and standard action templates in the cultural knowledge graph.
[0112] In this embodiment, behavioral features include motion trajectories, gaze sequences, etc., temporal action patterns capture dynamic changes in actions, and similarity calculations (such as DTW or Euclidean distance) are used to quantify the degree of matching between user actions and standard templates.
[0113] Specifically, feature extraction uses a sliding window to segment time-series data and extract statistical features (such as mean and variance); similarity calculation uses distance metrics and thresholds to determine intent categories, and pattern recognition improves the objectivity of intent analysis.
[0114] S32: Multimodal intent features are generated by weighted fusion of semantic labels of the user's eye-tracking focus area, gesture operation object identifiers, and voice keywords through an attention mechanism.
[0115] In this embodiment, the attention mechanism is a deep learning technique that automatically assigns weights to different modal features to highlight important information; the multimodal intent feature is a comprehensive vector that represents a fusion representation of the user's intent.
[0116] Specifically, attention weights are calculated using a function and then weighted and summed to generate intent features. For example, if a user is looking at a cultural relic, the semantic label of that cultural relic will have a higher weight. Through feature optimization, the robustness of intent recognition is enhanced.
[0117] S33: Use graph neural networks to traverse the cultural knowledge graph, retrieve cultural connotation nodes associated with intent features, and output user intent classification.
[0118] In this embodiment, the graph neural network can process graph structure data and traverse the knowledge graph to find relevant nodes; cultural connotation nodes represent deeper cultural meanings, and the intention classification output is such as "exploration intention" or "creation intention".
[0119] Specifically, the graph neural network uses a message-passing mechanism to update node representations and retrieve nearest neighbor nodes; classification outputs a probability distribution through a softmax layer. This detail, through graph analysis, enhances the semantic depth of the recognition.
[0120] In one embodiment, such as Figure 5 As shown, in step S40, an adaptive cultural context is generated based on the user's cultural operational intentions and emotional tendencies. Based on the adaptive cultural context, multi-sensory feedback content is output, specifically including:
[0121] S41: Based on the user's cultural operation intention, retrieve suitable narrative elements from the cultural knowledge graph, and generate a personalized narration script based on the narrative elements.
[0122] In this embodiment, narrative elements are story nodes in a knowledge graph, such as historical events or legends; personalized narration scripts are text or voice content customized based on user profiles (such as age or interests).
[0123] Specifically, retrieval uses query language to extract relevant elements from the graph; script generation is achieved through template filling or NLG models. This detail enhances the appeal of the content through personalization technology.
[0124] S42: Adjust the depth of the output content based on the user's cognitive level, and dynamically display the structural details of cultural relics through the 3D model decomposition module.
[0125] In this embodiment, the cognitive level represents the user's knowledge level, and the content depth is adjusted, such as a simplified version for beginners and a detailed version for experts; the 3D model decomposition allows users to interactively explore the internal structure of cultural relics.
[0126] Specifically, the decomposition module uses a geometric segmentation algorithm to dynamically load detailed level models; depth adjustment automatically switches based on preset rules. This step, through adaptive rendering, meets the needs of different users.
[0127] S43: It integrates a physics engine to simulate the real behavior of cultural elements, maps the vibration frequency of the haptic handle to the surface friction data, and generates multi-sensory synchronous feedback in conjunction with environmental sound effects.
[0128] In this embodiment, the physics engine simulates object dynamics, tactile vibrations provide tactile feedback, and ambient sound effects enhance auditory immersion; multi-sensory synchronization ensures a consistent experience.
[0129] Specifically, the physics simulation is based on rigid body dynamics and collision detection; vibration frequency is linearly mapped to friction force data; and spatial audio technology is used for sound effects. This detail, through hardware integration, enhances the realism of the interaction.
[0130] In one embodiment, such as Figure 6 As shown, in step S13, based on the geometric features, visual feature vectors, and semantic feature vectors of the cultural relics, a fusion network is trained using a cross-modal alignment loss function to generate a unified multimodal traditional cultural knowledge graph, specifically including:
[0131] S131: Normalize the geometric features and text features respectively, and calculate the inter-modal alignment loss using cosine similarity.
[0132] In this embodiment, normalization is used to scale the feature vectors to the same scale (e.g., the range of 0-1) to avoid certain modalities dominating the training process; cosine similarity is used to measure the angular similarity between vectors, and alignment loss minimizes the consistency of features between modalities.
[0133] Specifically, normalization uses min-max scaling or Z-score standardization to ensure comparability of all features. Then, cosine similarity is calculated: for a geometric feature vector G and a text feature vector T, the similarity score is cos(G,T) = (G·T) / (||G|| * ||T||), and the alignment loss is defined as 1 - cos(G,T) to optimize the fusion network. This detail, through mathematical optimization, enhances the stability of multimodal fusion.
[0134] S132: Introducing an adversarial training strategy enables the fusion network to distinguish between correct and incorrect associations of cultural elements, enhances semantic consistency, adds time dimension labels to the knowledge graph, supports the diachronic evolution display of cultural elements, and generates a unified multimodal traditional cultural knowledge graph.
[0135] In this embodiment, adversarial training, through the game between the generator and the discriminator, enables the network to learn to distinguish between real and false cultural associations, thereby improving the reliability of the knowledge graph; the time dimension label allows the graph to dynamically display cultural evolution, such as the changes in the style of cultural relics over time.
[0136] Specifically, in adversarial training, the generator attempts to generate seemingly plausible feature associations, while the discriminator judges the authenticity of these associations; the loss function combines alignment loss and adversarial loss for joint optimization. Time tags are added based on a historical timeline, such as using timestamps to mark the creation or modification time of nodes. This step, through advanced training techniques, ensures the timeliness and accuracy of the knowledge graph.
[0137] In one embodiment, such as Figure 7 As shown, in step S43, which involves integrating the physics engine to simulate the realistic behavior of cultural elements, the specific steps include:
[0138] S431: Collect material light reflection data and surface friction tactile data of intangible cultural heritage objects, and generate a physical property mapping table.
[0139] In this embodiment, light reflection data describes the material's response to light, and friction tactile data represents surface roughness; the physical property mapping table is a lookup table that associates material properties with virtual effects.
[0140] Specifically, data acquisition utilizes a spectrometer and a tribometer; the mapping table is stored as key-value pairs for real-time lookup. This step, through physical measurement, ensures the accuracy of the simulation.
[0141] S432: Simulates the interaction between ambient light and objects using a high dynamic range lighting model, and calculates the reflection intensity distribution in real time.
[0142] In this embodiment, high dynamic range (HDR) lighting simulates real light variations, and the reflection intensity distribution describes the scattering of light on the object's surface, used to generate realistic visual effects.
[0143] Specifically, the HDR model uses a radiosity algorithm to calculate in real time based on scene lighting and material properties. This detail enhances visual quality through graphics technology.
[0144] S433: Uses a spring-mass model to simulate the impact vibration response, and combines it with an acoustic model to generate the corresponding sound effects.
[0145] In this embodiment, the spring-mass model is a physical model that simulates the vibration of an object; the acoustic model generates sound effects based on the vibration frequency and provides synchronous feedback.
[0146] Specifically, the model solves differential equations to calculate the vibration response; sound effects are achieved through waveform synthesis or sampling. This step enhances multi-sensory consistency through physical acoustic integration.
[0147] S434: Performs a comprehensive evaluation of the physically rendered results. If the material matching index is below the threshold, it triggers a high-precision model recalculation.
[0148] In this embodiment, the material matching index quantifies the similarity between the simulation and the real material, and the threshold determines whether optimization is needed; recalculation uses a more refined model to improve accuracy.
[0149] Specifically, the evaluation involves comparing the rendered results with reference data; recalculation dynamically switches to a high-resolution mesh. This detail, through quality monitoring, ensures the reliability of the output.
[0150] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0151] In one embodiment, a traditional culture digital output and interaction system based on virtual reality technology is provided. This system corresponds one-to-one with the traditional culture digital output and interaction methods based on virtual reality technology described in the above embodiments. Figure 8 As shown, this digital output and interaction system for traditional culture based on virtual reality technology includes a cultural knowledge graph construction module, a user interaction data processing module, a user intent recognition module, and an interactive content output module. Detailed descriptions of each functional module are as follows:
[0152] The cultural knowledge graph construction module is used to acquire three-dimensional data of cultural relics, historical documents and intangible cultural heritage process sequences to construct a multimodal traditional cultural knowledge graph.
[0153] The user interaction data processing module is used to collect user interaction data based on multiple sensors, perform noise reduction and normalization processing on the user interaction data based on a noise index optimization algorithm, and generate processed user interaction data.
[0154] The user intent recognition module is used to dynamically identify the user's cultural operation intent and emotional tendency based on the user interaction data parsed and processed by the multimodal traditional culture knowledge graph.
[0155] The interactive content output module is used to generate an adaptive cultural context based on the user's cultural operational intentions and emotional tendencies, and output multi-sensory feedback content based on the adaptive cultural context.
[0156] Specific limitations regarding the digital output and interaction system of traditional culture based on virtual reality technology can be found in the above section on the limitations of the digital output and interaction methods of traditional culture based on virtual reality technology, and will not be repeated here. Each module in the aforementioned digital output and interaction system of traditional culture based on virtual reality technology can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the electronic device in hardware form or independent of it, or stored in the memory of the electronic device in software form, so that the processor can call and execute the corresponding operations of each module.
[0157] In one embodiment, an electronic device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9As shown, this electronic device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The database stores user interaction data and a traditional cultural knowledge graph. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a traditional cultural digital output and interaction method based on virtual reality technology.
[0158] In one embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0159] Acquire 3D data of cultural relics, historical documents, and sequences of intangible cultural heritage processes to construct a multimodal knowledge graph of traditional culture;
[0160] User interaction data is collected by multiple sensors, and the user interaction data is denoised and normalized based on a noise index optimization algorithm to generate processed user interaction data.
[0161] Based on the user interaction data processed by the multimodal traditional culture knowledge graph, the user's cultural operation intentions and emotional tendencies are dynamically identified.
[0162] An adaptive cultural context is generated based on the user's cultural operational intentions and emotional tendencies, and multi-sensory feedback content is output based on the adaptive cultural context.
[0163] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0164] Acquire 3D data of cultural relics, historical documents, and sequences of intangible cultural heritage processes to construct a multimodal knowledge graph of traditional culture;
[0165] User interaction data is collected by multiple sensors, and the user interaction data is denoised and normalized based on a noise index optimization algorithm to generate processed user interaction data.
[0166] Based on the user interaction data processed by the multimodal traditional culture knowledge graph, the user's cultural operation intentions and emotional tendencies are dynamically identified.
[0167] An adaptive cultural context is generated based on the user's cultural operational intentions and emotional tendencies, and multi-sensory feedback content is output based on the adaptive cultural context.
[0168] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0169] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0170] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for digital output and interaction of traditional culture based on virtual reality technology, characterized in that, The method for digital output and interaction of traditional culture based on virtual reality technology includes the following steps: Acquire 3D data of cultural relics, historical documents, and sequences of intangible cultural heritage processes to construct a multimodal knowledge graph of traditional culture; User interaction data is collected by multiple sensors, and the user interaction data is denoised and normalized based on a noise index optimization algorithm to generate processed user interaction data. Based on the user interaction data processed by the multimodal traditional culture knowledge graph, the user's cultural operation intentions and emotional tendencies are dynamically identified. An adaptive cultural context is generated based on the user's cultural operational intentions and emotional tendencies, and multi-sensory feedback content is output based on the adaptive cultural context.
2. The method for digital output and interaction of traditional culture based on virtual reality technology according to claim 1, characterized in that, The acquisition of 3D data of cultural relics, historical documents, and sequences of intangible cultural heritage craft processes, and the construction of a multimodal traditional cultural knowledge graph, specifically includes: Collect high-precision three-dimensional point cloud data and multi-angle texture images of cultural relics, and generate geometric and visual feature vectors of cultural relics based on the three-dimensional point cloud data and multi-angle texture images; Cultural elements are extracted from historical documents to describe text, and semantic feature vectors are generated based on the cultural element description text. Based on the geometric features, visual feature vectors, and semantic feature vectors of the cultural relics, a fusion network is trained using a cross-modal alignment loss function to generate a unified multimodal traditional cultural knowledge graph. Each cultural element in the knowledge graph is labeled with a dynamic contextual tag, including applicable scenarios, user cognitive levels, and interaction trigger conditions.
3. The method for digital output and interaction of traditional culture based on virtual reality technology according to claim 1, characterized in that, The process of collecting user interaction data from multiple sensors and then performing denoising and normalization processing on the user interaction data using a noise index optimization algorithm to generate processed user interaction data specifically includes: The user's hand gesture trajectory is collected by an inertial measurement unit and an infrared camera, and the Pearson correlation between acceleration and angular velocity is calculated to eliminate jitter noise. An eye tracker is used to record the coordinates of the user's gaze point, and the attention distribution matrix is calculated by combining the positions of virtual objects in the cultural context. The hand muscle signals are acquired by electromyography (EMG) sensors, the fine manipulation force is analyzed, and dynamic time warping is performed and matched with a standard cultural movement library. The multi-source user interaction data was time-stamped and spatial coordinate system unified to generate processed user interaction data.
4. The method for digital output and interaction of traditional culture based on virtual reality technology according to claim 1, characterized in that, The step of dynamically identifying users' cultural operational intentions and emotional tendencies based on the user interaction data parsed and processed from the multimodal traditional culture knowledge graph also includes: User behavior features are extracted based on the processed user interaction data, and temporal action patterns are extracted based on the user behavior features. The similarity between these patterns and standard action templates in the cultural knowledge graph is then calculated. Multimodal intent features are generated by weighted fusion of semantic labels of the user's eye-tracking focus area, gesture operation object identifiers, and voice keywords through an attention mechanism. By using a graph neural network to traverse the cultural knowledge graph, cultural connotation nodes associated with intent features are retrieved, and user intent classification is output.
5. The method for digital output and interaction of traditional culture based on virtual reality technology according to claim 1, characterized in that, The step of generating an adaptive cultural context based on the user's cultural operational intentions and emotional tendencies, and outputting multi-sensory feedback content based on the adaptive cultural context, specifically includes: Based on the user's cultural operational intent, suitable narrative elements are retrieved from the cultural knowledge graph, and a personalized narration script is generated based on the narrative elements. The depth of the output content is adjusted based on the user's cognitive level, and the structural details of the cultural relics are dynamically displayed through the 3D model decomposition module. It integrates a physics engine to simulate real-life behaviors of cultural elements, maps surface friction data through the vibration frequency of the haptic handle, and generates multi-sensory synchronous feedback in conjunction with environmental sound effects.
6. The method for digital output and interaction of traditional culture based on virtual reality technology according to claim 2, characterized in that, The process of generating a unified multimodal traditional cultural knowledge graph by training a fusion network using a cross-modal alignment loss function based on the geometric features, visual feature vectors, and semantic feature vectors of the cultural relics, specifically includes: Geometric features and text features are normalized separately, and the inter-modal alignment loss is calculated using cosine similarity. By introducing an adversarial training strategy, the fusion network can distinguish between correct and incorrect associations of cultural elements, enhance semantic consistency, add time dimension labels to the knowledge graph, support the diachronic evolution display of cultural elements, and generate a unified multimodal traditional cultural knowledge graph.
7. The method for digital output and interaction of traditional culture based on virtual reality technology according to claim 5, characterized in that, The fusion physics engine simulates the real-world behavior of cultural elements, specifically including: Collect material light reflection data and surface friction tactile data of intangible cultural heritage objects to generate a physical property mapping table; The interaction between ambient light and objects is simulated using a high dynamic range lighting model, and the reflection intensity distribution is calculated in real time. The spring-mass model is used to simulate the impact vibration response, and the corresponding sound effect is generated by combining it with the acoustic model. The physical rendering results are comprehensively evaluated. If the material matching index is lower than the threshold, a high-precision model recalculation is triggered.
8. A digital output and interactive system for traditional culture based on virtual reality technology, characterized in that: The traditional culture digital output and interaction system based on virtual reality technology includes: The cultural knowledge graph construction module is used to acquire three-dimensional data of cultural relics, historical documents and intangible cultural heritage process sequences to construct a multimodal traditional cultural knowledge graph. The user interaction data processing module is used to collect user interaction data based on multiple sensors, perform noise reduction and normalization processing on the user interaction data based on a noise index optimization algorithm, and generate processed user interaction data. The user intent recognition module is used to dynamically identify the user's cultural operation intent and emotional tendency based on the user interaction data parsed and processed by the multimodal traditional culture knowledge graph. The interactive content output module is used to generate an adaptive cultural context based on the user's cultural operational intentions and emotional tendencies, and output multi-sensory feedback content based on the adaptive cultural context.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the traditional culture digital output and interaction method based on virtual reality technology as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the traditional culture digital output and interaction method based on virtual reality technology as described in any one of claims 1 to 7.