Non-perpetual culture immersive interactive experience device and system based on virtual digital human
Through technologies such as multimodal perception, knowledge base, digital human modeling, immersive interaction, adaptive learning and cross-platform deployment, the problems of multimodal fusion, immersive interaction and security in the digitization of intangible cultural heritage have been solved, and efficient and safe dissemination and inheritance of intangible cultural heritage have been achieved.
Patent Information
- Application Number
- CN202510839627.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the process of digitizing intangible cultural heritage, existing technologies have problems such as insufficient multimodal data fusion, poor immersive interactive experience, difficulty in achieving personalized learning, complex cross-platform deployment, and insufficient security and privacy protection. These problems make it difficult to meet the needs of widespread dissemination and in-depth inheritance of intangible cultural heritage.
A comprehensive technical solution adopts multimodal perception module, intangible cultural heritage knowledge base module, digital human modeling engine, immersive interaction engine, adaptive learning module, cross-platform deployment module and cultural security protection module. Through depth cameras, microphone arrays, tactile sensors, knowledge graphs, generative adversarial networks, WebGL, WebXR, blockchain and other technologies, it realizes multi-sensory collaboration, personalized learning, cross-platform operation and security protection.
It has achieved multi-dimensional dynamic modeling of intangible cultural heritage, personalized immersive interactive experience, smooth cross-platform operation and all-round security protection, which has improved the efficiency and attractiveness of the dissemination of intangible cultural heritage, lowered the learning threshold, and ensured the reliability and security of cultural inheritance.
Smart Images

Figure CN120743104A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent interaction technology for intangible cultural heritage, and in particular to an immersive interactive experience device and system for intangible cultural heritage based on a virtual digital human. Background Art
[0002] As an important carrier of human civilization, intangible cultural heritage embodies unique cultural memories and the essence of craftsmanship. However, in the process of modernization, it faces practical challenges such as discontinuities in inheritance, a single form of dissemination, and limited experience scenarios. Traditional intangible cultural heritage preservation relies heavily on physical displays, documentary records, and oral transmission, which are difficult to break through the limitations of time and space, and are particularly lacking in appeal among young people. With the rapid development of digital technology, virtual reality, artificial intelligence, big data, and other technologies have provided new avenues for the living transmission of intangible cultural heritage. However, existing technologies still have shortcomings in multimodal data fusion, immersive interactive experiences, and personalized learning adaptation. For example, traditional intangible cultural heritage digitization projects often remain at the level of simple image recording or static modeling, lacking the ability to accurately capture and restore the dynamic characteristics of intangible cultural heritage skills (such as gesture rhythm, voice intonation, and operating force). At the interactive level, the lack of synergy between multiple sensory channels can easily lead to a sense of fragmented user experience. At the same time, existing systems struggle to achieve adaptive content adjustment and experience transfer to meet the personalized learning needs of different user groups, limiting the widespread dissemination and in-depth inheritance of intangible cultural heritage.
[0003] In the field of cross-platform applications and cultural security, existing technologies also have obvious shortcomings. The use threshold of intangible cultural heritage digitization systems is often high due to equipment compatibility issues. The operating efficiency of complex algorithms on different terminals varies significantly, and the application of technologies such as edge computing and lightweight rendering has not yet formed a mature system. In addition, issues such as copyright protection of intangible cultural heritage elements, compliance review of interactive content, and user privacy security are becoming increasingly prominent in the process of digital communication. Traditional copyright management methods and data security technologies are difficult to meet the needs of the industrialization development of intangible cultural heritage. How to build a comprehensive system that integrates multimodal perception, intelligent modeling, immersive interaction, adaptive learning, cross-platform deployment and security protection has become a key issue that needs to be urgently addressed in the current field of intangible cultural heritage digitization. Based on the above technical pain points, the present invention proposes an immersive interactive experience system for intangible cultural heritage based on virtual digital humans through the integration of multidisciplinary technologies. It aims to break through the limitations of traditional protection models and provide more vital technical solutions for the inheritance, dissemination and innovation of intangible cultural heritage. Summary of the Invention
[0004] The present invention proposes an immersive interactive experience device and system for intangible cultural heritage based on virtual digital humans to solve the problems mentioned in the above-mentioned prior art.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] An immersive interactive experience system for intangible cultural heritage based on virtual digital humans, including the following modules:
[0007] Multimodal perception module: Depth cameras, microphone arrays, and tactile sensors distributed throughout the interactive space collect user body movements, voice commands, and touch feedback data in real time to construct user behavior vectors in three-dimensional space.
[0008] Intangible Cultural Heritage Knowledge Base Module: This module includes a structured intangible cultural heritage knowledge graph and a dynamic case library. The knowledge graph uses a graph database to store the semantic associations between intangible cultural heritage elements, while the case library stores dynamic demonstration videos of intangible cultural heritage skills based on spatiotemporal indexing.
[0009] Digital human modeling engine: This uses a multimodal fusion algorithm based on a generative adversarial network to fuse the intangible cultural heritage inheritor's feature vector with user interaction data in real time to generate a virtual digital human. The algorithm satisfies the following equation: H(v) = α·H(t) + (1-α)·H(u), where H(v) is the virtual digital human's feature vector, H(t) is the inheritor's feature vector, H(u) is the user's feature vector, and α is the cultural feature retention coefficient.
[0010] Immersive Interaction Engine: This engine builds a cross-platform rendering environment based on WebGL and WebXR standards. It uses spatial audio rendering algorithms and tactile feedback mapping algorithms to achieve an interactive experience that integrates all five senses. The spatial audio rendering algorithm meets the following requirements: Where θ is the horizontal azimuth angle, is the vertical azimuth, is the spatial sound pressure distribution perceived by the user, is the binaural transfer function, A(d) is the distance attenuation function, and S(f) is the sound source characteristic function;
[0011] Adaptive learning module: A two-layer LSTM network is used to analyze user interaction sequences in real time. The knowledge distillation algorithm is used to incorporate user preferences into the digital human behavior decision model. The algorithm meets the following requirements: Where L(θ) is the model loss function, which is used to measure the difference between the model prediction results and the actual results, and θ is the model parameter; is the expected log-likelihood term, which reflects the accuracy of the model in predicting the output y under a given input x; λ is the weight coefficient, D kl (Q||P) is the KL divergence.
[0012] Furthermore, it also includes a cultural communication effect evaluation module, which constructs a user attention heat map through eye tracking data and calculates the distribution of gaze duration of intangible cultural heritage elements; analyzes user facial micro-expressions based on the emotional computing model and extracts cultural resonance indicators; constructs a communication influence evaluation formula: I = β1·A+β2·E+β3·S, where I is the communication influence index, which is a quantitative indicator for comprehensively evaluating the effect of intangible cultural heritage communication; β1, β2, and β3 are the weight coefficients of attention distribution entropy A, emotional arousal E, and social sharing frequency S, respectively, which are set according to different evaluation needs.
[0013] Furthermore, it also includes a cultural innovation generation module, which uses conditional variational autoencoders to build a space for the generation of intangible cultural heritage elements; realizes the integration and innovation of traditional elements and modern design through cultural gene algorithms; and constructs an innovation evaluation function: C = D (K t ,K m )·(1-ρ(K t ,K m ), where C is the innovation index, which quantitatively evaluates the degree of innovation and integration of intangible cultural heritage elements; D(K t ,K m ) is the cultural distance metric, and the set of traditional intangible cultural elements K is calculated. t With modern design elements collection K m The degree of difference between t ,K m ) is the style similarity coefficient, which measures the similarity between traditional and modern elements in style.
[0014] Furthermore, the multimodal perception module also includes: an action recognition unit based on a spatiotemporal pyramid, configured to extract the spatiotemporal feature vectors of intangible cultural heritage gestures; a dialect speech recognition unit, which uses a multi-task learning framework to simultaneously process speech content and dialect features; and a tactile feedback encoding unit, which maps the intangible cultural heritage skill operation process into a vibration frequency sequence.
[0015] Furthermore, the intangible cultural heritage knowledge base module also includes: a dynamic knowledge update unit, configured to continuously expand the knowledge graph through web crawlers and expert annotations; a cultural semantic analysis unit, which uses a pre-trained language model to analyze the deep semantic relationships of intangible cultural heritage texts; and a case index optimization unit, which realizes the retrieval of intangible cultural heritage demonstration videos based on the spatiotemporal cube index structure.
[0016] Furthermore, the digital human modeling engine also includes: a multimodal feature fusion unit, which uses an attention mechanism to weight the integration of visual, speech and text features; a cultural feature inheritance unit, which quantifies the degree of cultural feature retention through feature perturbation experiments; and a real-time action generation unit, which synthesizes intangible cultural heritage actions that conform to physical laws based on a generative adversarial network based on kinematic constraints.
[0017] Furthermore, the immersive interaction engine also includes: a multi-channel rendering synchronization unit, which uses a frame prediction algorithm to compensate for the delay differences between different sensory channels; a tactile feedback mapping unit, which maps the force curve of intangible cultural heritage skills operation into a vibration intensity sequence; and a spatial audio rendering unit, which realizes 3D stereo field simulation based on the HRTF database.
[0018] Furthermore, the adaptive learning module also includes: a user preference modeling unit, which uses a Bayesian personalized ranking algorithm to construct a user interest graph; an interaction strategy optimization unit, which dynamically adjusts the interaction difficulty and content depth based on reinforcement learning; and a knowledge transfer unit, which realizes the transfer of cross-user interaction experience through meta-learning.
[0019] Furthermore, it also includes a cross-platform deployment module, a performance optimization framework based on WebAssembly, which enables cross-platform operation of the algorithm; an adaptive resolution rendering unit, which dynamically adjusts the rendering quality according to the performance of the terminal device; and an edge computing collaboration unit, which offloads computing-intensive tasks to the edge server.
[0020] Furthermore, it also includes a cultural security protection module, an intangible cultural heritage copyright protection unit, which uses blockchain technology to record the usage trajectory of cultural elements; a content review unit, which identifies inappropriate interactive content based on a multimodal sentiment analysis algorithm; and a privacy protection unit, which optimizes the model while protecting user privacy through a federated learning mechanism.
[0021] Compared with the existing technology, the beneficial effects of the present invention are:
[0022] The multimodal perception module leverages depth cameras, microphone arrays, and tactile sensors to accurately capture user body movements, voice commands, and touch feedback in real time. Combined with technologies like spatiotemporal pyramid motion recognition and dialect speech multi-task learning, it fully preserves the spatiotemporal rhythms of intangible cultural heritage gestures, the regional characteristics of dialect speech, and the dynamic texture of skillful manipulation, providing a multidimensional, dynamic data foundation for the digital modeling of intangible cultural heritage. The intangible cultural heritage knowledge base module, through the collaboration of a dynamic knowledge update unit and a cultural semantic analysis unit, continuously expands the knowledge graph and mines deep semantic relationships. Furthermore, case library retrieval technology based on spatiotemporal cube indexing significantly improves the efficiency of accessing intangible cultural heritage demonstration videos, building a dynamically growing and clearly structured intangible cultural heritage knowledge system.
[0023] The digital human modeling engine, through the combination of a generative adversarial network and an attention mechanism, achieves an intelligent fusion of the characteristics of intangible cultural heritage inheritors and user interaction data, ensuring the preservation of cultural characteristics while dynamically adjusting digital human behavior based on user preferences. For example, the cultural characteristic inheritance unit, through quantitative assessment through feature perturbation experiments, can prevent the loss of core elements of intangible cultural heritage during the digitization process. The real-time action generation unit, based on an adversarial network with kinematic constraints, can synthesize intangible cultural heritage movements that conform to the laws of physics, making virtual digital human presentations more realistic and professional. The immersive interaction engine builds a cross-platform rendering environment based on WebGL and WebXR standards. It eliminates sensory latency differences through multi-channel rendering synchronization technology. Combined with 3D sound field simulation and tactile feedback mapping algorithms based on the HRTF database, it creates an immersive experience with highly coordinated auditory, tactile, and visual senses. Users can deeply participate in the learning and experience of intangible cultural heritage skills through physical interaction and voice commands, enhancing cultural resonance.
[0024] The adaptive learning module, leveraging a Bayesian personalized ranking algorithm and reinforcement learning, accurately captures user interests and preferences and dynamically adjusts interaction difficulty, enabling personalized learning paths tailored to each user. The knowledge transfer unit, leveraging meta-learning, rapidly reuses cross-user interaction experiences, lowering the learning barrier for new users and improving the efficiency and universality of intangible cultural heritage knowledge dissemination. The cross-platform deployment module, leveraging WebAssembly performance optimization, adaptive resolution rendering, and edge computing collaboration, ensures smooth operation across diverse devices, overcoming device performance limitations. The cultural security protection module, leveraging blockchain copyright protection, multimodal content moderation, and federated learning privacy protection, builds a comprehensive protection system encompassing copyright management, content security, and data privacy, providing reliable assurance for the industrialized dissemination of intangible cultural heritage. Overall, through technological integration and innovation, this system has achieved a significant leap from static recording to dynamic interaction, and from one-way dissemination to two-way immersion, providing an efficient, secure, and highly engaging technical paradigm for the dynamic inheritance and modern innovation of intangible cultural heritage. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A schematic block diagram of the present invention;
[0026] Figure 2 A survey of 1,000 users was conducted for this system, and a satisfaction score bar chart (1-5 points) was created by age and functional module.
[0027] Figure 3 A line chart comparing the duration of experience of different intangible cultural heritage projects in the traditional system and this system;
[0028] Figure 4 This is a radar chart comparing the fusion effects of digital human modeling features between the traditional fusion method and this method. DETAILED DESCRIPTION
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0030] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.
[0031] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined. In addition, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be a connection between the two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. The present invention will be further described in detail below with reference to the accompanying drawings.
[0032] Reference Figures 1 to 4 : An immersive interactive experience system for intangible cultural heritage based on virtual digital humans, including the following modules:
[0033] Multimodal Perception Module: Eight Azure Kinect DK depth cameras are deployed in an equilateral pattern at the top and four corners of the interactive space. These cameras integrate RGB cameras and depth sensors, achieving a resolution of 1920×1080 and a frame rate of 60fps, enabling full, all-around, stereoscopic visual capture of the interactive area. The Openpose algorithm processes image data captured by the cameras, extracting the coordinates of 25 key points on the human body, such as the head and joints, in real time. Combined with the Spatiotemporal Pyramid Network (STPN), this algorithm accurately identifies intangible cultural heritage gestures, such as Tai Chi and opera movements, by analyzing the temporal trajectory and spatial relative position of key points.
[0034] The dialect speech recognition unit, built on the Transformer architecture, collects a large-scale dataset covering 10 typical Chinese dialects, including Wu, Cantonese, and Minnan, totaling over 100,000 hours of speech content. During model training, a multi-head attention mechanism enables the model to simultaneously focus on different features and contextual information of the speech signal, optimizing its ability to understand complex dialects.
[0035] The tactile feedback encoding unit, designed for Suzhou embroidery, utilizes a high-precision pressure sensor integrated into a custom embroidery tool to capture pressure curves during the embroidery process in real time. A pre-established pressure-vibration frequency mapping model converts pressure data into a corresponding vibration frequency sequence. A PID controller precisely controls the tactile feedback glove, achieving a force feedback accuracy of ±0.5N. This allows users to experience the tactile feedback brought about by varying force applied to different stitches during Suzhou embroidery.
[0036] Intangible Cultural Heritage Knowledge Base Module: This module uses the JanusGraph distributed graph database as its storage medium, building a massive and detailed intangible cultural heritage knowledge graph. The graph includes 1,200 intangible cultural heritage item nodes, covering various intangible cultural heritage categories, including traditional skills, folk customs, and folk literature. 8,500 relationship edges are also established, clearly illustrating semantic relationships between intangible cultural heritage items, such as inheritance, regional connections, and cultural similarities. For example, these edges clearly illustrate the similarities and differences in embroidery techniques between Suzhou and Hunan embroidery, as well as their respective positions within the embroidery cultural system.
[0037] The case library is dedicated to storing 3,000 hours of dynamic demonstration videos of intangible cultural heritage techniques. To enable fast and efficient retrieval, Space-Time Cube Index technology is used to segment and label video content along a timeline, and index it using spatial scene information. When a user searches for the manifestations of a specific intangible cultural heritage technique across different historical periods, the system responds within milliseconds, quickly locating the relevant video clips.
[0038] The Cultural Semantic Parsing Unit, redeveloped based on Baidu's ERNIE pre-trained model, is fine-tuned for specialized texts in the field of intangible cultural heritage, such as proposals for intangible cultural heritage projects and biographies of inheritors. This unit accurately understands the deep semantic relationships within texts and uncovers the underlying meaning of intangible cultural heritage, providing strong support for knowledge graph development and user interaction. The Dynamic Knowledge Update Unit features automated crawling tasks, regularly acquiring new intangible cultural heritage information monthly from authoritative sources such as the official website of the National Intangible Cultural Heritage Protection Center and local cultural heritage databases. Content is then reviewed through a combination of expert annotation and machine learning.
[0039] Digital Human Modeling Engine: This engine, deeply customized and developed based on the StyleGAN3 model, generates highly realistic digital human avatars. During the generation process, extensive feature perturbation experiments determined that the cultural feature retention coefficient α is 0.75. This means that the generated digital human retains 75% of the characteristics of the intangible cultural heritage inheritor while incorporating 25% of the user's personalized features, achieving an organic combination of cultural heritage and user interaction.
[0040] The multimodal feature fusion unit employs a multi-head attention mechanism to perform a weighted integration of three types of features: visual, speech, and text. Visual features extract key information such as the digital human's facial expressions and body movements; speech features include intonation, speaking speed, and dialect characteristics; and text features primarily draw on intangible cultural heritage content and user interaction statements. This mechanism ensures that the digital human can deliver a consistent performance based on different interaction scenarios and information input.
[0041] The real-time motion generation unit is built using a kinematically constrained generative adversarial network (GAN). This unit incorporates specialized knowledge of intangible cultural heritage movements, such as the stylized norms of operatic movements and the rhythmic rhythms of traditional dance, to establish corresponding constraints for motion generation. During the motion generation process, the generator creates the motion sequence, while the discriminator evaluates the authenticity and rationality of the generated motions. Through continuous iterative optimization, the resulting intangible cultural heritage motions achieve a Naturalness Score (MOS) of 4.2 (out of a maximum of 5), enabling the virtual digital human to showcase the charm of intangible cultural heritage skills with smooth and natural movements. The generation algorithm satisfies the following equation: H(v) = α·H(t) + (1-α)·H(u), where H(v) is the feature vector of the virtual digital human, which comprehensively reflects the digital human's appearance, movements, language and other characteristics; H(t) is the inheritor's feature vector, which includes the unique cultural characteristics of the intangible cultural heritage inheritor, such as appearance, skills, movements, and language style; H(u) is the user's feature vector, which reflects the user's movement habits, voice characteristics and other personalized characteristics; α is the cultural feature retention coefficient, which ranges from 0 to 1 and is used to adjust the proportion of the inheritor's cultural characteristics in the virtual digital human.
[0042] Immersive Interaction Engine: Based on WebGL and WebXR standards, the immersive interaction engine builds a cross-platform immersive rendering environment that can run on multiple terminals such as PC, mobile devices, VR / AR / MR, etc., providing users with a consistent high-quality experience. In terms of spatial audio rendering, the HRIR (Head-Related Impulse Response) database, such as CIPICHRTFDatabase, is used in combination with spatial audio rendering algorithms. in is the spatial sound pressure distribution perceived by the user, describing the user's different orientations in three-dimensional space (θ is the horizontal azimuth angle, is the vertical azimuth) perceived sound pressure; is the binaural transfer function, which simulates the transmission characteristics of sound from the sound source to the human ears due to the influence of the head, torso, etc.; A(d) is the distance attenuation function, which reflects the law that sound intensity attenuates with increasing propagation distance d; S(f) is the sound source characteristic function, which includes the frequency f, timbre, loudness and other characteristics of the sound source; through this algorithm, a spatial positioning accuracy of ±5° is achieved, allowing users to accurately judge the direction and distance of the sound source through hearing, enhancing the sense of immersion.
[0043] The tactile feedback mapping unit targets the operation of intangible cultural heritage skills. Taking Suzhou embroidery as an example, it pre-collects the force curve data of professional embroiderers under different needle techniques and establishes a detailed force-vibration intensity mapping table. During the user experience, the force changes of the user's actual operation are mapped in real time to the vibration intensity sequence of the tactile feedback gloves, and the vibration frequency error is controlled within ±3Hz, allowing users to feel the operational details of the intangible cultural heritage skills through touch. The multi-channel rendering synchronization unit uses a frame prediction algorithm to perform real-time analysis and prediction of data from different sensory channels such as vision, hearing, and touch. By dynamically adjusting the rendering frame rate and data transmission order, the audio-visual delay is reduced to 12ms, effectively avoiding the experience fragmentation problem caused by the lack of synchronization of multi-channel data, and realizing an immersive interactive experience that integrates the five senses.
[0044] Adaptive Learning Module: Utilizing a two-layer LSTM (Long Short-Term Memory) architecture, this module conducts in-depth analysis of user-system interaction sequences. During training, the batch size is set to 64, and the learning rate is 0.001. Model parameters are continuously optimized through iteration to capture both long-term dependencies and short-term trends in user interactions. In the knowledge distillation algorithm, the weight coefficient λ is set to 0.3 to balance knowledge transfer between the original model and the student model, efficiently integrating preference knowledge accumulated from historical user interactions into the digital human behavioral decision-making model.
[0045] The user preference modeling unit is based on the Bayesian Personalized Ranking (BPR) algorithm. It builds a personalized user interest graph based on the user's interaction history data, such as the intangible cultural heritage items they have browsed, the interactive activities they have participated in, and the evaluation feedback they have given. The NDCG@10 index reaches 0.82, which can accurately predict the user's interest in different intangible cultural heritage content. The interaction strategy optimization unit is based on the Proximal Policy Optimization (PPO) reinforcement learning algorithm. It uses the user's interaction feedback as a reward signal to dynamically adjust the difficulty level, content depth, and methods of interaction between the digital human and the user. The algorithm meets the following requirements: Where L(θ) is the model loss function, which is used to measure the difference between the model prediction results and the actual results, and θ is the model parameter; is the expected log-likelihood term, which reflects the accuracy of the model in predicting the output y under a given input x; λ is the weight coefficient, which is used to adjust the relative importance of the two losses; D kl (Q||P) is the KL divergence (Kullback-Leibler divergence), which is used to measure the difference between the user's historical interaction distribution Q and the current interaction prediction distribution P.
[0046] The present invention also includes a cultural communication effectiveness evaluation module. This module uses high-precision eye tracking equipment to collect real-time user eye movement data during interaction and constructs an attention heat map using a specialized algorithm. In a large-scale test involving 1,000 people, an analysis of areas displaying core intangible cultural heritage content revealed an average gaze duration of 8.3 seconds. The calculated attention distribution entropy A was 1.23, a low value indicating a high degree of user focus on the core displayed content.
[0047] The emotional arousal E was assessed using the OpenFace algorithm, which analyzes real-time micro-expressions on the user's face, such as raised corners of the mouth and frowns. Statistical analysis revealed that the average emotional arousal during the test was 0.67 (maximum 1), indicating a strong emotional response during the user interaction. The social sharing frequency S, measured through the system's built-in sharing function, measures the number of times users proactively share intangible cultural heritage content. Using the communication influence evaluation formula I = β1·A + β2·E + β3·S, with β1 = 0.4, β2 = 0.3, and β3 = 0.3, the final calculated communication influence index I was 0.78, demonstrating the system's effectiveness in promoting intangible cultural heritage. Among them, I is the communication influence index, which is a quantitative indicator for comprehensively evaluating the communication effect of intangible cultural heritage; β1, β2, and β3 are the weight coefficients of attention distribution entropy A, emotional arousal E, and social sharing frequency S, respectively, which are set according to different evaluation requirements; A is the attention distribution entropy, which reflects the degree of concentration of users' attention on intangible cultural heritage elements. The lower the entropy value, the more concentrated the attention; E is the emotional arousal, which is obtained by analyzing users' facial micro-expressions through an emotional calculation model. The higher the value, the stronger the user's emotional response; S is the social sharing frequency, which counts the number of times users share intangible cultural heritage content.
[0048] The present invention also includes a cultural innovation generation module, which uses a conditional variational autoencoder (CVAE) to construct an intangible cultural heritage element generation space, sets the potential variable dimension to 64 dimensions, and enables the model to generate diverse and innovative intangible cultural heritage elements in the latent space through learning and encoding a large amount of intangible cultural heritage element data. The cultural gene algorithm is based on traditional intangible cultural heritage patterns, combined with modern design concepts and popular elements, and realizes the fusion and innovation of traditional elements and modern design through operations such as gene recombination and mutation. In terms of innovation evaluation, the evaluation function C=D(K t ,K m )·(1-ρ(K t ,K m ), where C is the innovation index, which quantitatively evaluates the degree of innovation and integration of intangible cultural heritage elements; D(K t ,K m ) is the cultural distance metric, by calculating the set of traditional intangible cultural elements K tWith modern design elements collection K m The difference between them measures the degree of innovation breakthrough; ρ(K t ,K m ) is the style similarity coefficient, ranging from 0 to 1, which measures the degree of similarity between traditional and modern styles. Calculations yielded a cultural distance metric, D, of 0.68, a style similarity coefficient, ρ, of 0.32, and a final innovation index, C, of 0.46.
[0049] In the present invention, the multimodal perception module also includes: an action recognition unit based on the spatiotemporal pyramid, which uses a hierarchical aggregation strategy to fuse the spatial posture and time series of intangible cultural heritage gestures into a model, and accurately extracts the spatiotemporal feature vector containing dynamic rhythm and spatial trajectory; the dialect speech recognition unit adopts a multi-task learning framework to construct a dual-channel model that processes speech content and dialect acoustic features in parallel, retaining regional language characteristics while recognizing semantic information; the tactile feedback encoding unit quantifies and deconstructs the operating process of intangible cultural heritage skills, and converts operation dimensions such as strength and speed into interactive vibration frequency sequences through a dynamic parameter mapping algorithm, thereby realizing digital tactile expression of skill operations.
[0050] In the present invention, the intangible cultural heritage knowledge base module also includes: a dynamic knowledge update unit combined with web crawler technology to capture news, research results and other data related to intangible cultural heritage on the entire network in real time, and introduce an expert annotation mechanism to continuously enrich and improve the knowledge graph through manual verification and collaboration with intelligent algorithms; a cultural semantic analysis unit uses advanced pre-trained language models to deeply analyze intangible cultural heritage texts, which can not only identify surface information, but also explore deep semantic associations between texts to achieve structured knowledge sorting; a case index optimization unit uses a spatiotemporal cube index structure to accurately divide and mark the spatiotemporal dimensions of intangible cultural heritage demonstration videos, significantly improving retrieval efficiency, and users can quickly locate the required intangible cultural heritage cases.
[0051] In the present invention, the digital human modeling engine also includes: a multimodal feature fusion unit constructs a feature aggregation network based on the self-attention mechanism, realizes the deep fusion of visual, voice and text features through a dynamic weight allocation strategy, and effectively captures the spatiotemporal correlation and semantic consistency in the intangible cultural heritage expression; the cultural feature inheritance unit designs a feature perturbation experimental framework, and through systematic adjustment of the digital human feature parameters, quantitatively evaluates the retention of cultural features, and ensures the integrity of the core elements of intangible cultural heritage during the digital conversion process; the real-time action generation unit adopts a generative adversarial network architecture based on kinematic constraints, embeds the laws of physics into the action generation process, realizes the natural synthesis of intangible cultural heritage actions that conform to the laws of human movement, and supports real-time interactive response and personalized action editing.
[0052] In the present invention, the immersive interaction engine also includes: a multi-channel rendering synchronization unit uses advanced frame prediction algorithms to accurately capture and compensate for the delay differences between different sensory channels such as vision, hearing, and touch, ensuring the seamless connection and synchronous presentation of multimodal information; a tactile feedback mapping unit performs a detailed analysis of the operation force curve of the intangible cultural heritage skills, and through a complex mathematical mapping model, converts it into a layered vibration intensity sequence, allowing users to truly perceive the force changes of the skills during the operation process; the spatial audio rendering unit relies on a huge HRTF (head-related transfer function) database to simulate the human ear's perception characteristics of sounds in different directions, achieve a realistic 3D stereo field, and create an immersive intangible cultural atmosphere for users.
[0053] In the present invention, the adaptive learning module also includes: the user preference modeling unit uses the Bayesian personalized ranking algorithm to deeply mine the historical behavior data of users in the process of intangible cultural heritage learning, dynamically constructs an accurate user interest map, and accurately captures individual learning tendencies; the interaction strategy optimization unit is driven by reinforcement learning, and dynamically adjusts the interaction difficulty level and content depth based on the user's real-time feedback and learning status, ensuring the fluency of learning while maintaining the challenge; the knowledge transfer unit uses the meta-learning mechanism to break through the limitations of the traditional learning model, efficiently realize the rapid transfer of cross-user interaction experience, accelerate the learning process of new users, and make the transmission of intangible cultural heritage knowledge more efficient and universal.
[0054] The present invention also includes a cross-platform deployment module, which is configured as: a performance optimization framework based on WebAssembly, which eliminates compatibility barriers between different operating systems and browsers by compiling complex algorithms into underlying bytecodes, and significantly improves the execution efficiency of algorithms on multiple platforms; an adaptive resolution rendering unit equipped with an intelligent monitoring system to detect hardware parameters such as the GPU performance and memory capacity of terminal devices in real time, and dynamically adjust the rendering resolution and picture details to ensure smooth operation on low-configuration devices and ultimate image quality on high-performance devices; an edge computing collaboration unit builds an intelligent scheduling network for cloud and edge devices, intelligently offloading computing-intensive tasks such as rendering and data processing to the nearest edge server, greatly reducing latency, improving system response speed, and bringing users a seamless cross-platform usage experience.
[0055] The present invention also includes a cultural security protection module, which is configured as follows: the intangible cultural heritage element copyright protection unit uses the distributed ledger characteristics of blockchain technology to generate a unique digital identity for each intangible cultural heritage element, and records the entire life cycle trajectory of its creation, authorization, and use in an encrypted manner to ensure that copyright ownership is traceable and cannot be tampered with; the content review unit is equipped with a multimodal sentiment analysis algorithm, which can simultaneously analyze potential risks in text semantics, voice emotions and visual images, accurately identify inappropriate interactive content, and effectively maintain the purity of intangible cultural heritage communication; the privacy protection unit adopts a federated learning mechanism to enable collaborative model training without leaving the local data, which not only ensures user privacy security, but also integrates multi-party data to continuously optimize system performance, achieving a dual improvement in security and efficiency.
[0056] The above are only preferred specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. An immersive interactive experience system for intangible cultural heritage based on virtual digital humans, characterized by: Includes the following modules: Multimodal perception module: Depth cameras, microphone arrays, and tactile sensors distributed throughout the interactive space collect user body movements, voice commands, and touch feedback data in real time to construct user behavior vectors in three-dimensional space. Intangible Cultural Heritage Knowledge Base Module: This module includes a structured intangible cultural heritage knowledge graph and a dynamic case library. The knowledge graph uses a graph database to store the semantic associations between intangible cultural heritage elements, while the case library stores dynamic demonstration videos of intangible cultural heritage skills based on spatiotemporal indexing. Digital human modeling engine: This uses a multimodal fusion algorithm based on a generative adversarial network to fuse the intangible cultural heritage inheritor's feature vector with user interaction data in real time to generate a virtual digital human. The algorithm satisfies the following equation: H(v) = α·H(t) + (1-α)·H(u), where H(v) is the virtual digital human's feature vector, H(t) is the inheritor's feature vector, H(u) is the user's feature vector, and α is the cultural feature retention coefficient. Immersive Interaction Engine: This engine builds a cross-platform rendering environment based on WebGL and WebXR standards. It uses spatial audio rendering algorithms and tactile feedback mapping algorithms to achieve an interactive experience that integrates all five senses. The spatial audio rendering algorithm meets the following requirements: Where θ is the horizontal azimuth angle, is the vertical azimuth, is the spatial sound pressure distribution perceived by the user, is the binaural transfer function, A(d) is the distance attenuation function, and S(f) is the sound source characteristic function; Adaptive learning module: A two-layer LSTM network is used to analyze user interaction sequences in real time. The knowledge distillation algorithm is used to incorporate user preferences into the digital human behavior decision model. The algorithm meets the following requirements: Where L(θ) is the model loss function, which is used to measure the difference between the model prediction results and the actual results, and θ is the model parameter; is the expected log-likelihood term, which reflects the accuracy of the model in predicting the output y under a given input x; λ is the weight coefficient, D kl (Q||P) is the KL divergence.
2. The intangible cultural heritage immersive interactive experience system based on virtual digital humans according to claim 1 is characterized in that: It also includes a cultural communication effect evaluation module, which constructs a user attention heat map through eye tracking data and calculates the distribution of gaze duration of intangible cultural heritage elements; analyzes user facial micro-expressions based on the emotional calculation model and extracts cultural resonance indicators; constructs a communication influence evaluation formula: I = β1·A+β2·E+β3·S, where I is the communication influence index, which is a quantitative indicator for comprehensively evaluating the effect of intangible cultural heritage communication; β1, β2, and β3 are the weight coefficients of attention distribution entropy A, emotional arousal E, and social sharing frequency S, respectively, which are set according to different evaluation needs.
3. The intangible cultural heritage immersive interactive experience system based on virtual digital humans according to claim 1 is characterized in that: It also includes a cultural innovation generation module, which uses conditional variational autoencoders to build a space for the generation of intangible cultural heritage elements; realizes the integration and innovation of traditional elements and modern design through cultural gene algorithms; and constructs an innovation evaluation function: C = D (K t ,K m )·(1-ρ(Kt,K m ), where C is the innovation index, which quantitatively evaluates the degree of innovation and integration of intangible cultural heritage elements; D(K t ,K m ) is the cultural distance metric, and the set of traditional intangible cultural elements K is calculated. t With modern design elements collection K m The degree of difference between t ,K m ) is the style similarity coefficient, which measures the similarity between traditional and modern elements in style.
4. The intangible cultural heritage immersive interactive experience system based on virtual digital humans according to claim 1 is characterized in that: The multimodal perception module also includes: an action recognition unit based on a spatiotemporal pyramid, which extracts the spatiotemporal feature vectors of intangible cultural heritage gestures; a dialect speech recognition unit, which uses a multi-task learning framework to simultaneously process speech content and dialect features; and a tactile feedback encoding unit, which maps the intangible cultural heritage skill operation process into a vibration frequency sequence.
5. The intangible cultural heritage immersive interactive experience system based on virtual digital humans according to claim 1 is characterized in that: The intangible cultural heritage knowledge base module also includes: a dynamic knowledge update unit, which is configured to continuously expand the knowledge graph through web crawlers and expert annotation; a cultural semantic analysis unit, which uses a pre-trained language model to analyze the deep semantic relationships of intangible cultural heritage texts; and a case index optimization unit, which realizes the retrieval of intangible cultural heritage demonstration videos based on the spatiotemporal cube index structure.
6. The intangible cultural heritage immersive interactive experience system based on virtual digital humans according to claim 1 is characterized in that: The digital human modeling engine also includes: a multimodal feature fusion unit, which uses an attention mechanism to weight the integration of visual, speech and text features; a cultural feature inheritance unit, which quantifies the degree of cultural feature retention through feature perturbation experiments; and a real-time action generation unit, which synthesizes non-legacy actions that conform to physical laws based on a generative adversarial network based on kinematic constraints.
7. The intangible cultural heritage immersive interactive experience system based on virtual digital humans according to claim 1 is characterized in that: The immersive interaction engine also includes: a multi-channel rendering synchronization unit, which uses a frame prediction algorithm to compensate for the delay differences between different sensory channels; a tactile feedback mapping unit, which maps the force curve of intangible cultural heritage skills operation into a vibration intensity sequence; and a spatial audio rendering unit, which realizes 3D stereo field simulation based on the HRTF database.
8. The intangible cultural heritage immersive interactive experience system based on virtual digital humans according to claim 1 is characterized in that: The adaptive learning module also includes: a user preference modeling unit, which uses a Bayesian personalized ranking algorithm to build a user interest graph; an interaction strategy optimization unit, which dynamically adjusts the interaction difficulty and content depth based on reinforcement learning; and a knowledge transfer unit, which realizes the transfer of cross-user interaction experience through meta-learning.
9. The intangible cultural heritage immersive interactive experience system based on virtual digital humans according to claim 1 is characterized in that: It also includes a cross-platform deployment module and a performance optimization framework based on WebAssembly to enable cross-platform operation of algorithms; Adaptive resolution rendering unit, dynamically adjusting rendering quality based on terminal device performance; Edge computing collaboration unit offloads computationally intensive tasks to edge servers.
10. The intangible cultural heritage immersive interactive experience system based on virtual digital humans according to claim 1 is characterized in that: It also includes a cultural security protection module, an intangible cultural heritage copyright protection unit, which uses blockchain technology to record the usage of cultural elements; a content review unit that identifies inappropriate interactive content based on a multimodal sentiment analysis algorithm; and a privacy protection unit that optimizes the model while protecting user privacy through a federated learning mechanism.
Citation Information
Cited By
Digital display system and method for non-perpetual cultural heritage
CN121187451A
Video AR virtual reality interactive communication method and system based on AI model
CN121788771A
Mask art multi-dimensional visual state display method and mask art multi-dimensional visual state display system
CN121837502A
Food safety interactive science popularization method and system based on AI digital human
CN121918702A
Food safety interactive science popularization method and system based on AI digital person
CN121918702B