Dynamic activation system and method for oral history intangible cultural heritage based on multi-modal fusion and digital life
By systematically combining multimodal data collection with AI-generated content, digital human modeling, VR/AR technology, and large language models, the problem of recording and transmitting oral history intangible cultural heritage has been solved, achieving high-fidelity, interactive, and immersive digital transmission and forming a sustainable cultural heritage ecosystem.
Patent Information
- Application Number
- CN202511595428.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-02-27
AI Technical Summary
Existing technologies cannot effectively capture and reproduce the "cultural space" and "interactive context" of oral history intangible cultural heritage. The aging and passing away of inheritors lead to cultural discontinuity. Digital archives lack intelligent re-creation and personalized narrative capabilities, resulting in a fragile inheritance chain.
By systematically combining multimodal data acquisition, AIGC (Artificial Intelligence Generated Content), high-fidelity digital human modeling, virtual reality (VR) and augmented reality (AR) technologies, and large language models (LLM), an interactive digital life form is constructed to achieve a holographic recording and dynamically evolving cultural narrative experience.
It achieves high-fidelity preservation of the appearance, movements, voice, and cultural knowledge of inheritors. Users can have natural conversations with digital inheritors, obtain personalized answers, and enjoy an immersive cultural experience. Furthermore, the knowledge base can be continuously updated with new research, forming a sustainable digital ecosystem for intangible cultural heritage.
Smart Images

Figure CN121578879A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of digital humanities and cultural heritage computing, and specifically relates to a system and method for dynamic recording, reconstruction and activation of oral history type intangible cultural heritage by comprehensively using multi-modal data collection, artificial intelligence generated content (AIGC), high-fidelity digital human modeling, virtual reality technology (VR) and augmented reality technology (AR), and large language model technology (LLM). BACKGROUND
[0002] (1) The protection and inheritance of oral history type intangible cultural heritage (such as epic, story, and ballad) is faced with the core contradiction between its living state and the static and fragmented nature of existing technical means. Existing technical solutions are mostly the application of a single technology, which has inherent defects: 1) Limitations of archival records: Traditional recording and video recording can only form a "digital specimen", and cannot capture and reproduce the "cultural space" and "interactive context" of the narration, separating recording from inheritance; 2) Lack of integration of technology: Although some solutions attempt to introduce three-dimensional scanning or virtual reality technology (VR), they mostly stop at scene reproduction and fail to model and activate the inheritor as a core digital asset, lacking semantic association and dynamic driving capability between data; 3) Shallow level of intelligence: Existing digital archives are mostly "dead data" that cannot respond to user queries, let alone intelligent content recreation and personalized storytelling, resulting in passive user experience; 4) Weakness of inheritance chain: The aging and death of the inheritor means the permanent discontinuity of the culture, and existing technology cannot effectively extend the identity and function of the "narrator".
[0003] (2) Therefore, there is an urgent need in the field for a comprehensive technical solution that can systematically solve the entire process from "recording" to "reconstruction" and achieve the "digital immortality" of the inheritor. SUMMARY
[0004] (1) Invention purpose: The present application aims to overcome the series of bottlenecks of existing technology and achieve the following purposes through a system and method of deep integration of multiple technologies: 1) Achieve holographic and structured recording of oral history intangible cultural heritage; 2) Build a high-fidelity "digital life form" of the inheritor with interactive ability; 3) Create an immersive interactive and dynamically evolving cultural narrative experience; 4) Form a sustainable and increasingly active digital ecosystem of intangible cultural heritage.
[0005] (ii) Technical Solution: The core of the invention lies in the systematic combination and logical connection of multiple cutting-edge technologies, forming a closed loop of technology that enhances each other.
[0006] 1) Refer to Figure 1 As mentioned above, the operation principle and mode of the entire system follow the closed-loop logic of "collection - modeling - injection - interaction": 1.1 Data infusion: Through data collection and preprocessing layer, the information of the inheritors of the physical world is fully digitized, forming an original data pool; 1.2 Model casting: In the artificial intelligence generated content (AIGC) enhancement and digital modeling layer, the original data is refined, enhanced, and cast into two core models: "digital human model" and "knowledge graph model"; 1.3 Life injection: In the digital life engine layer, the knowledge graph is used as "memory and knowledge", the large language model (LLM) is used as "thinking ability", and the voice and action driving model is used as "expression ability", which are injected into the digital human model together, thus creating a "digital life" with interactive ability; 1.4 Context interaction: In the application and interaction layer, users enter a specific context through virtual reality technology (VR) and augmented reality technology (AR) devices and interact with digital life. The user's question is processed by the dialogue agent to generate an answer that fits the context, and the digital human expresses it with the voice and expression of the inheritor, completing an immersive cultural dialogue.
[0007] 2) Refer to Figure 2 As mentioned above, the overall technical architecture level of the system includes the following four logical levels:
[0008] 2.1 Refer to Figure 3 As mentioned above, the data collection and preprocessing layer is composed of multiple 4K / 8K high-definition cameras, 360° panoramic cameras, high-fidelity microphone arrays, inertial motion capture suits / optical motion capture systems, and three-dimensional laser scanners arranged in specific spaces. It is responsible for synchronously collecting the appearance texture, micro-expression, stereo acoustic voice, full-body skeletal action, and three-dimensional geometric information of the environment of the inheritors;
[0009] 2.2 Refer to Figure 4 As mentioned above, the artificial intelligence generated content (AIGC) enhancement and digital modeling layer is the intelligent processing core of the invention, including: 2.2.1 Multi-modal data synchronization and processing module: Using synchronization algorithms based on hardware triggers or software timestamps to ensure that all data streams are strictly aligned; 2.2.2 Artificial intelligence generated content (AIGC) data enhancement module: voice processing submodule: an end-to-end automatic speech recognition model (ASR) is used to convert audio to text, and a large language model (LLM) is introduced to deeply clean the original transcription text, segment the semantics, mark the emotions, and extract the key entities (characters, places, objects, and events); video processing submodule: a super-resolution model (SR) based on a generative adversarial network (GAN) is used to enhance the quality of the video, and a neural radiance field technology (NeRF) is used to reconstruct a high-precision three-dimensional digital model (3D) from multi-angle videos; 2.2.3 Knowledge construction submodule: a large language model (LLM) is used to automatically construct a non-heritage project knowledge graph from structured text, and the characters, events, places, and cultural symbols are associated; 2.2.4 Digital human construction and driving module: image-driven, based on the reconstructed three-dimensional digital model (3D), combined with expression binding technology and motion capture data, a special expression and motion driving model is trained; voice-driven, using voice cloning technology, a personalized acoustic model is trained using a small amount of inheritor voice data, and combined with a high-quality text-to-speech engine, a synthesized voice is generated that is highly similar to the inheritor's tone and intonation;
[0010] 2.3 Reference Figure 5 Said, digital life engine layer: this is the control center and innovation soul of the invention. Based on the synergy mechanism, it integrates: 2.3.1 Digital human rendering engine: real-time rendering drives the image and motion of the digital human; 2.3.2 Knowledge graph database: stores and manages non-heritage knowledge; 2.3.3 Dialog agent: based on a large language model (LLM) and using retrieval augmented generation technology (RAG), the knowledge graph and the original transcription text are used as external knowledge bases to ensure that the digital human's answers are both logical and strictly faithful to the non-heritage ontology culture; 2.3.4 Narrative logic controller: dynamically schedules the behavior, voice, and scene changes of the digital human according to user interaction and pre-set scenarios;
[0011] 2.4 Reference Figure 6 Said, application and interaction layer: the output end for end users, including: 2.4.1 Virtual reality technology (VR) immersive narrative system: users wear virtual reality technology (VR) headsets and enter a 360° cultural scene generated by neural radiance field technology (NeRF) or artificial intelligence generated content (AIGC) to interact with the digital human inheritor face-to-face; 2.4.2 Augmented Reality (AR) interactive experience system: through augmented reality (AR) glasses or mobile devices, digital people are summoned to explain and perform in real spaces such as museums and cultural sites; 2.4.3 Interactive digital archives: web (Web) or client platform, providing semantic retrieval based on knowledge graph, clicking on any text can jump to the corresponding audio and video clips and digital explanation;
[0012] (Three) reference Figure 7 The characteristics and effects of the invention
[0013] 1) Combination of innovation characteristics: 1.1 Closed loop of recording and reconstruction: data that is "passively recorded" is converted into "actively output" digital life through artificial intelligence content generation (AIGC) and digital human technology, forming a complete closed loop from the physical world to the digital world, and then back to the physical world, experienced through virtual reality (VR) and augmented reality (AR); 1.2 Semantic-level fusion of multi-modal data: Instead of simply bundling data, knowledge graphs and large language models (LLM) are used to deeply integrate audio, video, text, and motion data at the semantic level, creating intelligent chemical reactions between data; 1.3 Digital human construction with "shape, sound, and spirit": combining neural radiance field technology (NeRF) (shape), voice cloning (sound), and large language models (LLM) + retrieval augmented generation technology (RAG) (spirit), a digital heritage subject that not only looks like the original but also conveys the essence and meaning is created, surpassing traditional virtual avatars;
[0014] 2) Technical integration effects: 2.1 Fidelity: Achieving ultra-high fidelity in preserving the appearance, movements, voice, and cultural knowledge of the heritage person; 2.2 Interactivity: Users can have natural language conversations with digital heritage people and receive personalized answers; 2.3 Immersion: Virtual reality (VR) and augmented reality (AR) provide a deep experience of being in a cultural context; 2.4 Evolution: The knowledge graph in the system's background can be updated with new research, and the digital person's knowledge base can also grow; 2.5 Universality: This methodology can be applied to most oral history-based intangible cultural heritage projects, making it highly valuable for widespread promotion. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 The entire system operation principle and mode flowchart of the invention; Figure 2The four-level diagram of the overall technical architecture of the system of the present application; Figure 3 The structure diagram of the data collection and preprocessing layer of the present application; Figure 4 The structure diagram of the artificial intelligence generated content (AIGC) enhancement and digital modeling layer of the present application; Figure 5 The structure diagram of the digital life engine layer of the present application; Figure 6 The structure diagram of the application and interaction layer of the present application; Figure 7 The structure diagram of the features and effects of the present application; Figure 8 The flowchart of the specific implementation scheme of the present application. DETAILED DESCRIPTION
[0016] Reference Figure 8 With a specific implementation scheme - "Digital activation of Mongolian epic 'The Legend of King Gesar' performance art" as an example, the operation of the present application is described in detail:
[0017] (I) Data collection: 1) In the inheritor's home or in the arranged cultural scene, set up 8K camera multi-position, microphone array, and ask the inheritor to wear the inertial motion capture suit; 2) Use a three-dimensional scanner to finely scan the inheritor and the instruments (such as the Tobu Show) and costumes used by the inheritor; 3) Record the inheritor's performance of "The Legend of King Gesar" for at least 2 hours, including normal speech and different emotional expressions.
[0018] (II) Artificial intelligence generated content (AIGC) enhancement and modeling: 1) Voice and text: Automatic Speech Recognition Model (ASR) converts performance audio, Large Language Model (LLM) (such as ChatGPT API or similar models deployed locally) corrects text (overcomes archaic language, proper nouns), and marks entities such as "Gesar", "Zhum", "Demon King", and events such as "warfare" and "celebration"; 2) Video and three-dimensional: Use Neural Radiance Field technology (NeRF) to generate high-precision three-dimensional digital assets of the inheritor from multi-angle videos. Action data is used to drive body movements; 3) Knowledge graph: Large Language Model (LLM) automatically constructs a knowledge graph centered on "King Gesar", connecting heroes, concubines, mounts, enemies, locations, and treasures; 4) Digital human driving: Use 5-minute inheritor voice data to train voice cloning model. Bind motion capture data with three-dimensional model to train action-driven model.
[0019] (Three) Scene construction and engine integration: 1) Use artificial intelligence generated content (AIGC) text to 3D tool (such as Masterpiece X), input "Inner Mongolia grassland, snow mountain, felt house", generate virtual reality technology (VR) scene basic model, and then refine by artist; 2) In Unity / Unreal engine, integrate digital human model, knowledge graph database and large language model (LLM) dialogue interface (such as access OpenAI API, and preset role as "Gesar King rapper").
[0020] (Four) User experience: 1) Virtual reality technology (VR) mode: students wear virtual reality technology (VR) headsets and find themselves in the vast grassland, with the digital inheritor sitting by the campfire. Students can approach to listen to the rap and ask, "What is special about the red horse of King Gesar?" The digital human will pause the rap and turn to the student to answer in the voice and manner of the inheritor: "Young man, you asked a good question! This red horse is the incarnation of the god, can run 10,000 miles a day, and walk on clouds…" At the same time, relevant knowledge cards will appear beside; 2) Augmented reality technology (AR) mode: in front of the "King Gesar" exhibit in the museum, visitors use tablet computers to scan the exhibits, and virtual digital inheritor augmented reality technology (AR) image appears and begins to explain the epic story behind the exhibits.
Claims
1. A dynamic revitalization system and method for oral history intangible cultural heritage based on multimodal fusion and digital life, characterized in that, The method includes the following steps: Step 1, synchronously collecting multimodal data of the inheritor through multi-source heterogeneous sensors, including high-definition video, audio, motion capture data, and environmental 3D information; Step 2, using generative artificial intelligence (AIGC) technology to enhance and structure the multimodal data, specifically including: generating high-quality, semantically annotated text using automatic speech recognition (ASR) and large language model (LLM); reconstructing a high-precision 3D model of the inheritor from the video using neural radiation field (NeRF) technology; and automatically extracting entities and relationships from the text based on the large language model (LLM) to construct an intangible cultural heritage knowledge graph; Step 3, based on the data obtained in Step 2... The method involves using 3D models, motion data, and audio data to drive the construction of a high-fidelity digital human model of the inheritor, integrating voice cloning and text-to-speech synthesis technologies to enable the digital human to broadcast using a similar timbre and tone to the inheritor; step four, utilizing the knowledge graph and generative artificial intelligence (AIGC) scene generation technology, creating a virtual reality or augmented reality immersive experience environment, and placing the digital human model constructed in step three into this environment; step five, within the immersive experience environment, by integrating a large-scale language model and retrieval-enhanced generation technology, enabling the user to interact with the digital human in natural language, and allowing the digital human to dynamically generate responses and narratives consistent with its cultural identity based on the knowledge graph. According to claim 1, the method is characterized in that, in step two, the processing of text using a large-scale language model includes: performing at least one of the following operations on the original text generated by automatic speech recognition: error correction, segmentation, sentiment analysis, key entity recognition, and relation extraction, to generate structured data that can be used to construct a knowledge graph.
2. The method according to claim 1, characterized in that, In step three, the driving of the high-fidelity digital human model includes: using a deep learning-based facial expression driving model to map audio features into corresponding lip movements, eye movements, and facial muscle movement sequences; and using motion capture data to drive full-body skeletal animation to achieve coordination and synchronization of voice, facial expressions, and body movements.
3. The method according to claim 1, characterized in that, In step four, the creation of the immersive experience environment using virtual reality (VR) and augmented reality (AR) technologies includes using 3D modeling technology to automatically or semi-automatically generate 3D virtual scene models based on scene descriptions in the intangible cultural heritage narrative.
4. The method according to claim 1, characterized in that, In step five, the natural language interaction specifically involves the user's voice input being converted into text via speech recognition and then input into a large-scale language model that incorporates retrieval-enhanced generation technology. This model first retrieves relevant background information from the knowledge graph and the original transcribed text, then generates an accurate and context-appropriate text response, and finally drives the digital human to read the response and display corresponding expressions and actions through the speech synthesis technology.
5. A system for implementing the method of any one of claims 1-5, characterized in that, The system includes: a data acquisition module, consisting of a camera, microphone, motion capture equipment, and 3D scanner, used to execute step one; a generative artificial intelligence (AIGC) processing and modeling server, equipped with graphics processing unit (GPU) computing power, used to run algorithms for automatic speech recognition (ASR), large language model (LLM), neural radiation field (NeRF) reconstruction, speech cloning, and knowledge graph construction, used to execute steps two and three; a digital life form engine, integrated into a game engine or dedicated rendering platform, including digital human rendering, knowledge graph management, dialogue proxy, and narrative logic control functions, used to execute the core logic of step five; and a virtual reality (VR) and augmented reality (AR) interactive terminal, including a virtual reality (VR) head-mounted display device, augmented reality (AR) glasses, or mobile computing device, used to render and present the final immersive experience, used to execute user interaction in steps four and five.