Virtual-real combined script tour immersion experience scene construction and interaction method

By integrating high-precision virtual and real scenes, AI NPC interaction, and user co-creation, the problems of high labor costs, rigid virtual NPCs, and slow content updates in traditional scripted games have been solved. This has enabled highly immersive cultural experiences and commercial value-added services, and supports the promotion of cultural tourism projects that are adaptable to multiple scenarios and cost-effective.

CN122488935APending Publication Date: 2026-07-31HEFEI SHIWEI DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEFEI SHIWEI DIGITAL TECH CO LTD
Filing Date
2026-04-29
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Traditional scripted games rely on real-life NPCs for guidance, which results in high labor costs, poor character performance stability, long plot update cycles, stiff virtual NPC interactions, low integration of virtual and real scenes, difficulty in meeting personalized experience needs, superficial cultural transmission, slow content updates, poor device adaptability, and difficulty in achieving cost-effective intelligent promotion.

Method used

It uses laser scanning technology to construct a 1:1 three-dimensional spatial model, achieves high-precision coordinate alignment between virtual elements and real space through edge computing, integrates AI NPCs and multimodal interaction, supports user-generated content creation, dynamic lighting adaptation, fault self-healing mechanism, constructs multi-NPC collaborative logic and interaction intent fusion, optimizes immersion evaluation and response priority, and embeds commercial nodes.

Benefits of technology

It achieves a highly immersive virtual-real experience, enhances cultural awareness and participation, improves content update efficiency, reduces operating costs, supports multi-scenario adaptability, and promotes the differentiated development and commercial value of the cultural tourism industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122488935A_ABST
    Figure CN122488935A_ABST
Patent Text Reader

Abstract

This invention discloses a method for constructing and interacting with immersive script-based games that combines virtual and real elements. It relates to the fields of cultural tourism and virtual-real interaction technology, and includes: extracting cultural elements from scenic spots and determining scene themes based on user profiles; modeling real-world spaces using laser scanning, deploying multiple types of presentation terminals to achieve precise alignment of virtual and real elements; creating AI NPCs adapted to cultural settings based on technologies such as LLM; designing multimodal interaction rules and multi-user collaborative logic, including voice and gestures; pushing storyline tasks during gameplay, with the AI ​​NPCs dynamically adjusting the storyline, and collecting data to fine-tune the model and optimize the experience. This invention enhances the immersiveness and cultural transmission effect of script-based games, provides flexible AI NPC interaction adapted to multiple scenarios, supports user co-creation of rich content to enhance user engagement, links the industry chain to reduce costs, balances cultural experience with commercial value-added, and contributes to the upgrading of cultural tourism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cultural tourism and virtual-real interaction technology, and in particular to the construction and interaction methods of immersive experience scenarios for scripted games that combine virtual and real elements. Background Technology

[0002] Under the trend of digital transformation in the cultural tourism industry, scripted tours, with their strong interactivity and immersive experience, have become an important form for enriching scenic spot formats and enhancing tourist participation. Traditional scripted tours mostly rely on live NPCs to guide the plot, which not only faces problems such as high labor costs, poor character performance stability, and long plot update cycles, but is also limited by fixed scenes and linear storylines, making it difficult to meet tourists' needs for personalized and diversified experiences. At the same time, most cultural tourism projects suffer from insufficient exploration of cultural connotations, conveying cultural information only through simple landscape displays or text descriptions, lacking storytelling and scene-based presentation methods. This results in tourists' perception of culture remaining superficial, with insufficient project fun and memorable elements, making it difficult to form a differentiated competitive advantage.

[0003] With the development of AI and virtual-real fusion technologies, some cultural tourism projects have begun to introduce virtual NPCs and AR / VR devices. However, existing solutions still have significant technical shortcomings. On the one hand, virtual NPCs often use fixed dialogue scripts, lacking natural language interaction capabilities and dynamic plot adjustment mechanisms. When tourists ask questions beyond the script's scope or perform unconventional actions, the NPCs cannot respond effectively, easily disrupting the immersive experience. On the other hand, the integration of virtual and real scenes is low, and the alignment accuracy between virtual elements and real space is insufficient. In outdoor environments with strong light or complex terrain, virtual images are prone to stuttering, blurring, or shifting. Furthermore, the presentation terminals have limited functionality, making it difficult to adapt to the experience needs of different scenarios and failing to achieve a seamless immersive effect. In addition, existing solutions lack user co-creation mechanisms. Script content is unilaterally output by the operator, making it difficult for tourists to participate in content creation. This results in slow content updates, rapid decline in attractiveness, and an inability to maintain long-term tourist engagement.

[0004] From an industry collaboration perspective, the current construction and operation of cultural tourism script-based tour projects heavily rely on single teams, with weak integration capabilities of upstream and downstream resources. The lack of an efficient collaboration platform among content creators, technology providers, scenic area operators, and businesses results in long project construction cycles, high costs, and a clumsy integration of commercial elements with script-based tour scenarios, making it difficult to achieve the synergistic effect of "cultural experience + commercial value-added." Furthermore, most projects lack standardized technical support systems, resulting in poor compatibility between different equipment and systems, and significant difficulties in subsequent maintenance and upgrades. Especially in outdoor settings, the equipment's anti-interference capabilities and environmental adaptability are insufficient, further limiting the large-scale promotion and sustainable development of script-based tour projects, and failing to meet the cultural tourism industry's demand for intelligent, lightweight, and cost-effective solutions. Summary of the Invention

[0005] The present invention proposes a method for constructing and interacting with immersive experience scenarios in scripted games that combines virtual and real elements, in order to solve the problems mentioned in the prior art.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: A method for constructing and interacting with immersive scenarios in script-based games that combine virtual and real elements, comprising the following steps: Scenario requirements analysis and cultural element mapping steps: For the target scenic spot / venue, extract regional cultural elements and construct a cultural element knowledge graph; Combine user profiles to determine the scenario theme and output a requirements document; The steps for constructing a virtual-real scene fusion are as follows: Laser scanning technology is used to model the real space and generate a 1:1 three-dimensional spatial model; based on the model, presentation terminals are deployed with a terminal spacing of ≤5m to ensure continuous interaction; the coordinates of virtual elements and real space are aligned through edge computing hosts with an anchor point positioning error of ≤3cm to generate a virtual-real fusion scene; Steps for creating and configuring AI NPC agents: Build AI NPC language logic based on the LLM model; integrate ASR, TTS, and action generation modules; configure NPC appearance according to the scene theme, set dialogue style, and bind cultural knowledge base to make NPC responses conform to cultural settings; Multimodal interaction rule design steps: Define user interaction methods, including voice commands, gesture operations, and AR gesture grasping; set up virtual and real interaction logic, that is, after the user touches the real plant, AR triggers a virtual plant information pop-up window and the AI ​​NPC explains in sync; formulate multi-user collaboration rules to support 2-8 people in a team, with the synchronization deviation of team users' actions ≤100ms and shared plot progress. Experience operation and dynamic optimization steps: After the user enters the scene, the edge computing host collects interaction data in real time, including user dwell time, number of interactions, and voice command recognition rate; push story tasks through handheld terminals, and the AI ​​NPC adjusts the story branches according to the user's response. When the user answers the plant characteristics correctly, a hidden story is triggered, and when the answer is wrong, clues are given; after the experience ends, user satisfaction scores are collected. Every 100 scores are accumulated, the AI ​​NPC dialogue model is fine-tuned, and the presentation position of virtual elements is optimized.

[0007] Furthermore, it also includes a real-time scene immersion assessment step, using a formula. Calculate scene immersion, where I represents immersion. As a weight for scene integration, The score is given for the alignment accuracy of the coordinates of the virtual and real elements. As a weight for interaction smoothness, Score for interaction response latency. Weighting for cultural compatibility, The system scores the accuracy of AI NPC cultural representation. When I < 80, the system automatically adjusts the terminal position, reduces interaction latency, and updates the NPC cultural knowledge base.

[0008] Furthermore, it also includes AI NPC interaction response priority scheduling steps, through formulas. Calculate the response priority, where P is the response priority. C represents the relevance weight of the plot, and C represents the relevance of the user command to the current plot. As a weighted average of waiting time, where T is the user's waiting time for a response. U represents the user's role weight, where U is the importance of the user's role in the team.

[0009] Furthermore, the virtual-real scene fusion construction steps also include dynamic lighting adaptation: collecting real-world spatial lighting data and adjusting the lighting parameters of virtual elements through edge computing hosts; for outdoor scenes, using light sensors to monitor lighting changes in real time, and automatically increasing the contrast of the water curtain projection and reducing the transparency of virtual elements when the light intensity is >800 lux.

[0010] Furthermore, the creation and configuration steps of the AI ​​NPC intelligent agent also include multi-role collaboration logic: when the scene contains multiple AI NPCs, a dialogue model between NPCs is constructed. This model is based on dialogue state tracking technology, and the dialogue state is updated at a frequency of 1 second / time. NPC collaboration rules are defined, with academic NPCs prioritizing the answering of professional questions and story-based NPCs prioritizing the advancement of the plot. When a user asks a question to multiple NPCs, the optimal responding NPC is selected based on the role matching degree. The formula for calculating the role matching degree is: M = overlap between the NPC knowledge base and the question × 0.8 + fit between the NPC role setting and the question × 0.2.

[0011] Furthermore, the multimodal interaction rule design process also includes interaction intent fusion judgment: using a weighted fusion algorithm to determine the intent of voice commands. Gestures and intentions Facial expressions and intentions Integration as the ultimate intention When the confidence of a single interaction intent is insufficient, it can be supplemented by combining other modalities.

[0012] Furthermore, the experience operation and dynamic optimization steps also include the integration of user-generated content (UGC): providing script fragment editing tools, allowing users to submit their creations to the platform; the platform adopts AI initial review + manual review, which is completed within 24 hours, and the approved content is marked with the "user co-creation" tag and included in the scene plot library; creators who use the content ≥ 100 times will be given a revenue sharing reward, with the reward amount = number of times the content is used × 0.1 yuan, and the content will be adapted to different display terminals to enrich the scene content.

[0013] Furthermore, it also includes steps for embedding commercial nodes and driving traffic: setting commercial trigger points in the storyline, through formulas. Calculate merchant exposure, where E is the exposure value. P represents the weight of the trigger point in the plot, where P is the importance of the trigger point in the plot. V represents the proportion of users who actually visit the merchant after the trigger is activated; the trigger point position is adjusted according to E.

[0014] Furthermore, it also includes a self-healing step for virtual-real interaction faults: the edge computing host monitors the terminal status in real time, and automatically switches to the backup transparent LED screen when the holographic cabin communication is interrupted; when the AR anchor point error is greater than 5cm, visual repositioning is initiated; the fault type and repair measures are recorded, and the fault diagnosis model is updated every 50 fault data points accumulated to reduce manual intervention.

[0015] Furthermore, the scenario requirements analysis and cultural element mapping steps also include a cultural element suitability assessment: three assessment indicators are constructed: cultural authenticity, user awareness, and interaction feasibility. The analytic hierarchy process (AHP) is used to determine the weights of each indicator, with cultural authenticity weighted at 0.5, user awareness at 0.3, and interaction feasibility at 0.2. Each cultural element is scored across the three indicator dimensions, ranging from 1 to 10, and the suitability score is calculated. To score for authenticity, To score awareness, For feasibility score, only elements with A≥8 are retained.

[0016] Compared with existing technologies, the beneficial effects of this invention are: This invention breaks through the bottlenecks of existing solutions from multiple dimensions, including cultural tourism experience, technological integration, and industrial collaboration, providing comprehensive optimization for the construction and interaction of scenario-based games that combine virtual and real elements. In terms of enhancing cultural experience and immersion, this invention deeply explores regional cultural elements, constructs a cultural knowledge graph, and integrates it with AI NPCs and plot design, allowing tourists to naturally perceive cultural connotations through interaction, thus solving the problem of superficial cultural transmission in traditional projects. Simultaneously, by leveraging high-precision virtual-real coordinate alignment and diverse presentation terminals, it achieves seamless integration of virtual elements and real space. Combined with multimodal interaction, it significantly enhances the immersion and participation of the experience, transforming tourists from "passive visitors" to "active participants in the plot," significantly improving the project's attractiveness and memorability.

[0017] In terms of technical adaptability and user experience flexibility, this invention optimizes technical solutions for different scenarios. Through dynamic lighting adaptation and fault self-repair mechanisms, it improves the stability of the equipment in environments such as strong light and complex terrain, avoiding problems such as virtual screen stuttering and offset. The AI ​​NPC has the ability to interact with natural language and dynamically adjust the plot. It can flexibly switch plot branches according to the visitor's response, breaking the limitation of fixed scripts. At the same time, it supports multi-NPC collaborative response to avoid interaction conflicts. The user co-creation mechanism allows visitors to participate in script creation, which is included in the content library after review. This greatly improves the efficiency of content updates, maintains the freshness of the project in the long term, and solves the pain points of traditional solutions such as monotonous content and slow updates.

[0018] In terms of industrial synergy and commercial value, the technical support system and collaboration platform constructed by this invention enable efficient linkage between content creation, technology supply, scenic area operation, and businesses, shortening project construction cycles and reducing operating costs, providing scenic areas with lightweight and cost-effective solutions. By rationally designing commercial trigger points, it naturally integrates commercial needs such as cultural and creative product sales and merchant traffic generation with the scripted game plot, avoiding the forced insertion of commercial elements. While enhancing the tourist experience, it creates additional revenue for scenic areas and businesses. In addition, the standardized technical architecture and adaptation mechanism facilitate subsequent project maintenance, upgrades, and large-scale promotion. It not only helps scenic areas create differentiated cultural tourism IPs but also drives the coordinated development of upstream and downstream industrial chains, injecting new vitality into the cultural tourism industry and possessing both cultural dissemination value and economic value-added benefits. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the method for constructing and interacting with immersive script-based games that combines virtual and real elements, as proposed in this invention. Figure 2 A bar chart comparing the immersion of virtual and real-world scenario-based game scripts in different settings; Figure 3 Radar chart comparing content updates and user engagement in scripted games; Figure 4 A bar chart comparing the technical performance of different presentation terminals. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0022] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified. Furthermore, the terms "installed," "connected," and "linked" should be interpreted broadly; for example, they may refer to a fixed connection, a detachable connection, or an integral connection; they may refer to a mechanical connection or an electrical connection; they may refer to a direct connection or an indirect connection through an intermediate medium; and they may refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. The invention will now be described in further detail with reference to the accompanying drawings.

[0023] Reference Figures 1 to 4 A method for constructing and interacting with immersive scenarios in script-based games that combine virtual and real elements, comprising the following steps: Scenario Requirements Analysis and Cultural Element Mapping Steps: For target scenic spots / venues, including historical sites, botanical gardens, etc., extract regional cultural elements, covering historical events, intangible cultural heritage skills, stories of famous people, plant characteristics, etc., and construct a cultural element knowledge graph containing 500+ cultural entities and their relationships; combined with user profiles, the user age range is 10-60 years old, and divided into three categories: study tours, families, and young tourists, determine the scenario theme, such as "Exploring Shennong's Herbal Medicine" and "Dialogue with Quantum Scientists", and output a requirements document, which includes information such as cultural element weights, user experience goals, and interaction frequency requirements.

[0024] The steps for constructing a virtual-real scene fusion are as follows: Laser scanning technology is used to model the real space with an accuracy of ±2mm. The real space includes areas such as scenic walkways and exhibition halls, generating a 1:1 three-dimensional spatial model. Based on this model, presentation terminals are deployed, including a holographic capsule all-in-one machine, a transparent LED screen, and a water curtain projection. The holographic capsule all-in-one machine has a 4K resolution and a 120° field of view; the transparent LED screen has a brightness of 800 nits and a response time of 1ms; the water curtain projection is 30m wide × 15m high with a projection frame rate of 60fps. The terminal spacing is ≤5m to ensure interactive continuity. An edge computing host is used to align the coordinates of virtual elements with the real space. The edge computing host has a computing power of 6TOPS and a latency of ≤50ms. Virtual elements include AI NPCs and plot prop models, with an anchor point positioning error of ≤3cm. Finally, a virtual-real fusion scene is generated.

[0025] Steps for creating and configuring AI NPC agents: First, construct the AI ​​NPC language logic based on an LLM model, such as LLaMA2-7B, whose fine-tuning dataset contains 100,000 cultural dialogue samples. Second, integrate ASR, TTS, and action generation modules. The ASR module has an accuracy rate ≥95% and supports dialects such as Cantonese and Sichuanese. The TTS module has a naturalness score of ≥4.8 / 5. The action generation module has an action frame rate of 30fps and supports 200+ basic actions. Third, configure the NPC image according to the scene theme, such as Shennong Yandi or Einstein, and set the dialogue style, including ancient style and academic style. Bind the NPC to a cultural knowledge base, such as the content of the "Shennong Bencao Jing" (Shennong's Classic of Materia Medica) and the basic theories of quantum science, so that the NPC's response conforms to the cultural settings.

[0026] Multimodal interaction rule design steps: Define user interaction methods, including voice commands, gesture operations, and AR gesture grasping. An example of a voice command is "Shennong, what plant is this?" Gesture operations include waving to wake up an NPC and pointing to select an item. The YOLOv8 gesture recognition response time is ≤300ms, and the AR gesture grasping accuracy for virtual items is ±5cm. Set up virtual-real interaction logic. Specifically, after a user touches a real plant, AR triggers a pop-up window for virtual plant information, and the AI ​​NPC provides a simultaneous explanation. Develop multi-user collaboration rules that support teams of 2-8 people, with a synchronization deviation of ≤100ms for team members' actions, and shared story progress.

[0027] Experience Operation and Dynamic Optimization Steps: After the user enters the scene, the edge computing host collects interaction data in real time, including user dwell time, number of interactions, and voice command recognition rate; pushes storyline tasks through handheld terminals (such as AR glasses), such as "find 3 medicinal plants and have the NPC identify them"; the AI ​​NPC adjusts the storyline branches based on the user's response. If the user correctly answers the plant characteristics, a hidden storyline is triggered; if the answer is incorrect, clues are given; after the experience ends, user satisfaction scores are collected (score range 1-5 points). Every 100 scores are accumulated, the AI ​​NPC dialogue model is fine-tuned (iteration 5 rounds, learning rate set to 1e-5), while the presentation position of virtual elements is optimized to improve immersion.

[0028] This invention also includes a real-time scene immersion assessment step, using a formula. Calculate the scene immersion level, where I is the immersion level, with a value range of 0-100. When I ≥ 80, it is considered to be high immersion. The scene integration weight is fixed at 0.4. The score is given for the alignment accuracy of the virtual and real element coordinates, with a value range of 0-100. A positioning error of ≤3cm earns 100 points, and 10 points are deducted for every 1cm increase. The interaction smoothness weight is fixed at 0.3. The score is for the interaction response delay, with a value range of 0-100. A delay of ≤50ms earns 100 points, and 5 points are deducted for every 10ms increase. The weight for cultural fit is fixed at 0.3. The accuracy score for AI NPC cultural representation is calculated, ranging from 0 to 100. A perfect score of 100 is awarded for representations that accurately reflect the cultural context, and 15 points are deducted for each deviation. When I < 80, the system automatically adjusts the terminal position to optimize scene integration, reduces interaction latency to improve interaction smoothness, and updates the NPC cultural knowledge base to enhance cultural fit. For example... =80、 =90、 When I = 85, I = 0.4 × 80 + 0.3 × 90 + 0.3 × 85 = 32 + 27 + 25.5 = 84.5 ≥ 80, which is considered a high level of immersion.

[0029] This invention also includes an AI NPC interaction response priority scheduling step, which is achieved through a formula. Calculate the response priority, where P is the response priority, with a value ranging from 0 to 10, and the larger the value, the higher the priority; The plot relevance weight is fixed at 0.5, and C is the relevance between the user command and the current plot, with a value range of 0-1, where 1 is obtained for complete relevance and 0 is obtained for no relevance. The waiting time weight is fixed at 0.3, and T is the user's waiting time for response (unit: s, value: 0.5-10). Let U be the user role weight (fixed at 0.2), and U be the importance of the user's role in the team, ranging from 0 to 1, with the team leader receiving 1 and ordinary members receiving 0.5. For example, if a user (team leader) issues the command "Trigger Key Plot" (C=1) and waits for T=1s, then P=0.5×1+0.3×(1 / 1)+0.2×1=0.5+0.3+0.2=10, and the AI ​​NPC responds first; if an ordinary user issues an irrelevant command (C=0) and waits for T=5s, then P=0.5×0+0.3×(1 / 5)+0.2×0.5=0+0.06+0.1=0.16, and the response is delayed.

[0030] In this invention, the virtual-real scene fusion construction step also includes dynamic lighting adaptation: collecting real-world spatial lighting data, with a brightness range of 100-1000 lux and a color temperature range of 3000-6500K, and adjusting the lighting parameters of virtual elements through an edge computing host. Specifically, the adjustment rule is that the virtual brightness is equal to the real brightness multiplied by 1.2, and the virtual color temperature is equal to the real color temperature ±200K. For outdoor scenes (such as botanical gardens), a light sensor is used to monitor changes in lighting in real time. The sensor sampling frequency is 1 second / time. When the light intensity is >800 lux, the system automatically increases the contrast of the water curtain projection (from 500:1 to 1000:1) and reduces the transparency of virtual elements (from 70% to 50%) to avoid blurring of virtual elements due to strong light and ensure consistency between virtual and real visuals.

[0031] In this invention, the AI ​​NPC intelligent agent creation and configuration steps also include multi-role collaboration logic: when the scene contains multiple AI NPCs (such as Einstein and Deng Jiaxian in the "Quantum History Museum"), a dialogue model between NPCs is constructed. This model is based on Dialogue State Tracking (DST) technology, and the dialogue state update frequency is 1 second / time. NPC collaboration rules are defined, with academic NPCs prioritizing answering professional questions and story-based NPCs prioritizing advancing the plot. When a user asks a question to multiple NPCs, the optimal responding NPC is selected based on the role matching degree. The role matching degree calculation formula is M = overlap between the NPC's knowledge base and the question × 0.8 + fit between the NPC's role setting and the question × 0.2. For example, if a user asks "the development history of quantum mechanics", Einstein's role matching degree M = 0.9, so he responds first, while Deng Jiaxian's role matching degree M = 0.4, so he provides supplementary assistance, thereby avoiding NPC response conflicts.

[0032] In this invention, the multimodal interaction rule design step also includes interaction intent fusion judgment: using a weighted fusion algorithm to determine the intent of voice commands. (Recognition confidence level ≥ 0.8) and gesture intent (Recognition confidence level ≥ 0.7) and facial expression intent Based on facial key point recognition (with a confidence level ≥ 0.6), the data is fused into the final intent. When the confidence of a single interaction intent is insufficient (e.g., voice confidence of 0.75), it can be supplemented by other modalities, such as the user's voice "open the virtual manual" ( =0.75) while performing a "page-turning gesture" ( =0.9), then =0.6×0.75+0.3×0.9+0.1×0.8=0.45+0.27+0.08=0.8≥0.8, the intent is deemed valid, avoiding misjudgment based on a single interaction method.

[0033] In this invention, the experience operation and dynamic optimization steps also include user-generated content (UGC) access: a script fragment editing tool is provided, which supports text, virtual props, and NPC dialogue editing functions. After completing their creations, users submit their content to the platform. The platform adopts a review method that combines AI initial review and manual review. The AI ​​initial review is based on a cultural compliance model, with an accuracy rate of ≥98% in identifying illegal content. The manual review is completed within 24 hours. Content that passes the review is labeled with the "user co-creation" tag and included in the scene plot library. Creators of high-quality content (judged by user usage ≥100 times) are given a revenue-sharing reward. The reward amount is calculated as: reward amount = number of times content is used × 0.1 yuan. At the same time, high-quality content is adapted to different presentation terminals (such as holographic capsule version and AR glasses version) to enrich the scene content.

[0034] This invention also includes steps for embedding and driving traffic to commercial nodes: setting commercial trigger points in the storyline (such as "after completing plant identification, you can go to the cultural and creative store to redeem souvenirs"), and using formulas... Calculate the merchant's exposure, where E is the exposure (value from 0 to 100). P represents the weight of the plot position (fixed at 0.6), and P represents the importance of the trigger point in the plot (value from 0 to 100, with key plot points receiving 100 points). The user access weight is fixed at 0.4, and V is the proportion of users who actually visit the merchant after triggering the event (value from 0 to 100). The trigger point position is adjusted according to E. For example, when P=90 and V=60, E=0.6×90+0.4×60=54+24=78. When the trigger point is adjusted to after the climax of the story, P=100 and V=80, E=0.6×100+0.4×80=60+32=92, which improves the merchant's traffic generation effect.

[0035] This invention also includes a self-repairing step for virtual-real interaction faults: the edge computing host monitors the terminal status in real time (communication delay, screen lag, anchor point offset), and when a communication interruption of the holographic cabin is detected (delay > 200ms), it automatically switches to the backup transparent LED screen (response time ≤ 100ms); when the AR anchor point is offset (error > 5cm), visual repositioning is initiated (based on real-world spatial feature points, repositioning time ≤ 1s); the fault type (communication fault, positioning fault) and repair measures are recorded, and the fault diagnosis model is updated every 50 fault data points accumulated (fault feature parameters are added, improving the diagnostic accuracy by 5%-8%), reducing manual intervention.

[0036] In this invention, the scenario requirements analysis and cultural element mapping steps also include a cultural element suitability assessment: constructing assessment indicators (cultural authenticity, user awareness, and interaction feasibility), and using the Analytic Hierarchy Process (AHP) to determine weights (authenticity 0.5, awareness 0.3, feasibility 0.2); scoring each cultural element (1-10 points), and calculating the suitability score. ( To score for authenticity, To score awareness, For feasibility score); only elements with A≥8 are retained, such as the story of "Shennong tasting hundreds of herbs" ( =10、 =9、 =9, A=0.5×10+0.3×9+0.2×9=5+2.7+1.8=9.5≥8) are included in the scenario, "niche intangible cultural heritage skills" ( =5, A=6.1<8) are not included for the time being to ensure that cultural elements are both authentic and easily accepted by users.

[0037] The following two examples further illustrate the specific implementation of this system: Example 1: Historical and Cultural Scenic Area Scene (Tang Dynasty Chang'an West Market "Silk Road Merchants" Script Tour) This embodiment focuses on the Xi'an Tang Chang'an West Market Site (covering an area of ​​approximately 15,000 square meters, including museum exhibition halls, restored shops, and pedestrian streets), constructing a "Silk Road Merchants" themed scripted tour. Tourists can play the role of Tang Dynasty merchants to complete tasks such as trade transactions and cultural exchanges. Through all the technical solutions in this application, the problems of superficial cultural transmission, rigid NPC interaction, and poor integration of virtual and real elements in historical scenic spot scripted tours are solved, and the technical parameters and practical procedures of each step are refined.

[0038] I. Scenario Requirements Analysis and Cultural Element Mapping Cultural element extraction: Through historical research (studying "Tang Liudian" and "Liangjing Zaji") and expert interviews, core cultural elements of the Tang Dynasty's West Market were extracted, including 12 types of shops (such as silk shops and taverns run by foreign merchants), 8 types of transaction currencies (such as Kaiyuan Tongbao and Persian silver coins), 6 commercial systems (such as the management of the Maritime Trade Office and the standardization of weights and measures), and 10 historical figures (such as the magistrate of the West Market and Persian merchants). Based on these elements, a knowledge graph was constructed, containing 520 cultural entities with 860 relationships between them, such as the relationship "tavern run by foreign merchants - main products - grapes and fine wine".

[0039] User profiling and needs matching: For three core user groups (study tour participants aged 10-18, families aged 25-45, and young adults aged 18-30), a questionnaire survey (sample size 1000) was conducted to determine the needs of each group: the study tour group focused on "acquiring knowledge of trade history," the families needed "parent-child collaborative tasks," and the young adults preferred "story-driven competitions and social interaction." Based on the survey results, a requirements document was generated, which included weights for cultural elements (commercial system weight 0.3, historical figures weight 0.3, shop type weight 0.2, and currency weight 0.2) and interaction frequency requirements (requiring an average of one key interaction every 5 minutes).

[0040] Cultural Adaptability Assessment: The AHP assessment of the "Maritime Trade Office System" element received an A. t =10 (Authenticity), A c =8 (awareness), A f =9 (feasibility), calculate A = 0.5×10 + 0.3×8 + 0.2×9 = 5 + 2.4 + 1.8 = 9.2 ≥ 8, include in the scenario; for "Tang Dynasty clothing making techniques" (A) c =6), A=6.8<8, so they are not included for the time being to ensure that cultural elements are easily understood by users.

[0041] II. Construction of Virtual and Real Scene Integration 3D Modeling and Terminal Deployment: The West Market Site was scanned using a Faro laser scanner with an accuracy of ±2mm. The scan generated a 1:1 3D model containing 300+ feature points, such as shop column bases and floor tile textures. Holographic capsule all-in-one machines were deployed in the restored shops with parameters of 4K resolution, 120° field of view, and a terminal spacing of 4m. Transparent LED screens were set up in the pedestrian street area with parameters of 800 nits brightness, 1ms response time, and dimensions of 5m wide × 3m high. A water curtain projection was configured in the square area with projection parameters of 30m × 15m size and 60fps frame rate.

[0042] Virtual-to-real alignment and dynamic lighting: Virtual element anchoring is achieved through an edge computing host with a computing power of 6 TOPS and a latency of 45ms. The anchoring operation is completed using 300+ feature points, with a positioning error of ≤2.5cm. A light sensor is used to monitor the lighting conditions, with a sampling frequency of 1s / time. In the exhibition hall (lighting range 300-500 lux), the brightness of the virtual elements is adjusted according to the rule of "real brightness × 1.2". For example, when the real brightness is 400 lux, the brightness of the virtual elements is adjusted to 480 lux accordingly. The color temperature of the virtual elements is adjusted according to "real color temperature ±150K". In the outdoor pedestrian street (sunny day lighting range 800-1000 lux), the contrast ratio of the LED screen is automatically increased to 1000:1, while the transparency of the virtual elements is reduced to 45% to avoid blurring of the virtual elements due to strong light.

[0043] III. Creation and Configuration of AI NPC Agents Core NPC Design: Create three types of NPCs: "West Market Magistrate" (administrator), "Persian Merchant" (foreign merchant), and "Silk Merchant" (local merchant). Language model: Based on LLaMA2-7B fine-tuning (100,000 Tang Dynasty commercial dialogue samples, 800 iterations, learning rate 2e-5), the dialogue style of "West Market Magistrate" is rigorous (word accuracy 98%), and "Persian Merchant" has a foreign accent (TTS speech naturalness 4.9 / 5).

[0044] Interactive module: ASR supports Tang Dynasty title recognition (such as "Lingjun" and "Zhanggui", with an accuracy of 96%), and the AIR module contains 200+ Tang Dynasty etiquette actions (such as cupping hands and bowing, with a frame rate of 30fps).

[0045] Knowledge base binding: "Persian merchants" is linked to knowledge such as "Silk Road trade routes" and "characteristics of goods from the Western Regions", with a response accuracy rate of 97%.

[0046] Multi-NPC collaborative logic: When a user asks "the price difference between silk and Persian brocade", the "silk merchant" (M=0.9, knowledge base overlap 0.9) responds first, and the "Persian merchant" (M=0.8) supplements the procurement cost in the Western Regions. DST technology ensures that the dialogue state is synchronized (updated every 1 second) to avoid conflicts.

[0047] IV. Multimodal Interaction and Priority Scheduling Interaction rule design: Voiceover: "West Market Order, how to apply for a customs clearance document" (ASR recognition accuracy rate 97%).

[0048] Gestures: wave your hand to wake up the NPC (YOLOv8 recognition response 280ms), raise your thumb to confirm the transaction (accuracy ±3cm).

[0049] AR Grab: Grab a virtual "Kaiyuan Tongbao" using AR glasses (grab error ≤ 4cm) to complete the payment interaction.

[0050] Intent fusion judgment: User's voice "View Persian silver coins" (I v =0.82) + pointing to virtual currency (I g =0.9) + Curious expression (I e =0.7), calculate I f =0.6×0.82+0.3×0.9+0.1×0.7=0.492+0.27+0.07=0.832≥0.8, the intention is deemed valid.

[0051] Priority scheduling: The team leader issues the "Trigger the Maritime Trade Office inspection plot" (C=1, T=1.2s, U=1), calculates P=0.5×1+0.3×(1 / 1.2)+0.2×1≈0.5+0.25+0.2=0.95 (9.5 / 10 after normalization), and the "West Market Order" responds first; ordinary users ask irrelevant questions (C=0, T=4s, U=0.5), P=0.22, and the response is delayed.

[0052] V. Experience Operation and Optimization Storyline: Users form a team (3-5 people) to receive the task "Purchase grapes and wine from Persian merchants and sell them to wine shops after verification by the Western Market Magistrate". The AR glasses push the task node; when the purchase is completed, the animation "Western Regions Caravan Transportation" is displayed on the water screen projection, and the NPC explains the trade route in real time.

[0053] Immersion assessment: Real-time calculation ,in =95 (positioning error 2.5cm) =90 (delay 45ms) =98 (accurate cultural expression), so I=0.4×95+0.3×90+0.3×98=38+27+29.4=94.4≥80, which indicates high immersion.

[0054] UGC and fault repair: A user created a script fragment of "Tang Dynasty Tavern Bidding". After the platform's AI initial review (99% violation detection) and the review by historical experts, it was launched. It was used 150 times within 30 days, and the creator received a share of 15 yuan. When the holographic cabin communication was interrupted (delay of 220ms), the system switched to the backup LED screen within 100ms, recorded the fault and updated the diagnostic model.

[0055] VI. Performance Verification Form Table 1 Table 1 data is based on experience testing with 100 groups of tourists (3-5 people per group). Traditional solutions rely on live NPCs and static display boards, resulting in a cultural transmission accuracy rate of less than 70%, and alignment errors between virtual and real elements (such as AR markers) exceeding 10cm, disrupting the immersive experience. This invention, through cultural knowledge graphs and high-precision anchoring technology, improves the cultural transmission accuracy to over 92%, controlling alignment errors within 2.5cm, allowing tourists to naturally absorb historical knowledge through interaction. NPC response latency is reduced from 2 seconds to within 500ms, avoiding interruptions to the experience; user repeat experience intention increases to 65%, stemming from UGC content updates and dynamic storyline branches, addressing the pain points of traditional scripted games' "fixed content and poor repeatability," increasing cultural memory points to 6-8, significantly enhancing the dissemination of historical culture.

[0056] Example 2: Scene from a natural science popularization scenic area (Shennong Herbal Botanical Garden "Herbal Exploration" script tour) This embodiment focuses on the Shennong Herbal Botanical Garden (covering an area of ​​30,000 square meters, including 10 herbal areas, a science museum, and a greenhouse), and constructs a "Herbal Exploration" themed scripted tour. Visitors follow the AI ​​Shennong to learn herbal identification and processing, and complete the "curing the plague" storyline task. The key focus is on verifying technologies such as outdoor scene adaptation, multi-terminal collaboration, and fault self-repair, and refining special technical solutions for natural science popularization scenarios.

[0057] I. Scenario Requirements Analysis and Cultural Element Mapping Cultural and scientific elements extraction: Integrating classic texts such as "Shennong's Classic of Materia Medica" and "Compendium of Materia Medica" with modern botanical data, core cultural and scientific elements were extracted, including 30 kinds of medicinal plants (such as astragalus and angelica), 8 processing techniques (such as slicing and stir-frying), the legend of Shennong tasting hundreds of herbs, and 5 kinds of knowledge about the efficacy of herbs; based on these elements, a knowledge graph was constructed, which contains 350 entities and 620 relationships between entities, such as the relationship "Astragalus - efficacy - tonifying qi and consolidating the exterior".

[0058] User needs matching: For two core user groups (students on study tours, aged 10-15; and families with children aged 3-12 and their parents), the focus of needs is clearly defined as "herb identification", "fun interaction" and "parent-child collaboration". Based on the needs analysis results, a needs document is output, which includes the weight of cultural and popular science elements (herb characteristics weight 0.4, processing technology weight 0.3, and legends weight 0.3), as well as interaction requirements (requiring one plant touch interaction every 2 minutes).

[0059] Cultural compatibility assessment: Element A of "Astragalus identification features" t =10、A c =7、A f=9, A=0.5×10+0.3×7+0.2×9=5+2.1+1.8=8.9≥8, included in the scenario; "Complex Processing Chemical Principles" (A c =4), A=5.8<8, simplified to an animated demonstration and then included.

[0060] II. Construction of Virtual and Real Scene Integration 3D Modeling and Terminal Deployment: The botanical garden was modeled using UAV laser scanning technology with an accuracy of ±3mm. The scan generated a 3D model containing 500+ plant feature points. A waterproof holographic cabin with an IP65 protection rating and 4K resolution was deployed in the herb area. A transparent LED screen with anti-glare treatment was installed in the greenhouse area. A solar-powered water curtain projector with a projection size of 30m×15m and a battery life of 8 hours was installed on the outdoor walkway. All terminals were spaced 5m apart to ensure coverage without blind spots.

[0061] Dynamic lighting and environmental adaptation: The outdoor herb area experiences significant variations in lighting, ranging from 200 lux on cloudy days to 1200 lux on sunny days. The system adjusts equipment parameters in real time using a light sensor: when the lighting exceeds 800 lux, the contrast ratio of the water curtain projection is increased to 1200:1, while the brightness of the virtual herb model is increased by 20%; when it rains, the system automatically triggers the terminal's rainproof mode, specifically by automatically closing the windows of the holographic cabin and activating the defogging function of the LED screen, thereby ensuring the normal operation of all equipment.

[0062] III. Creation and Configuration of AI NPC Agents Core NPC Design: Create two types of NPCs: "AI Shennong" (main guide) and "Medicine Boy" (assistant in teaching). Language model: Based on LLaMA2-7B fine-tuning (80,000 herbal knowledge dialogue samples), "AI Shennong" explains in classical Chinese (e.g., "This herb is called Angelica sinensis, which nourishes and invigorates blood, just like a traveler returning home"), with a TTS speech naturalness score of 4.8 / 5.

[0063] Interactive modules: ASR supports plant alias recognition (such as "Huangqi" is also known as "Huangqi", with an accuracy of 95%), and the AIR module includes actions for picking and processing herbs (30fps frame rate).

[0064] Knowledge base binding: linking the original text of "Shennong's Classic of Materia Medica" with modern pharmacological research (such as "Ephedra - efficacy - diaphoretic and exterior-releasing, modernly used to relieve asthma").

[0065] Multi-NPC collaboration: When a user asks "How to slice Angelica sinensis", the "Medicine Boy" (M=0.92, skilled in processing) demonstrates the slicing action, and the "AI Shennong" (M=0.85) supplements the slice thickness requirements. The dialogue status is synchronized without conflict.

[0066] IV. Multimodal Interaction and Priority Scheduling Interaction rule design: Voice prompt: "Shennong, what is the name of this herb?" (pointing to Astragalus membranaceus, ASR recognition accuracy 96%)

[0067] Gesture: Make a "herb picking" gesture with both hands (YOLOv8 recognition response time 250ms) to trigger the virtual picking tutorial.

[0068] AR grasp: Scan real Astragalus membranaceus with AR glasses → generate virtual plant cross-section → grasp virtual roots with gestures (accuracy ±3cm) to view medicinal parts.

[0069] Intent fusion judgment: User's voice "stir-fry Atractylodes macrocephala" (I v =0.78) + make a stirring gesture (I g =0.92) + Focused expression (I e =0.8), I f =0.6×0.78+0.3×0.92+0.1×0.8=0.468+0.276+0.08=0.824≥0.8, therefore it is valid.

[0070] Priority scheduling: In the parent-child group, if a parent asks "What diseases can Astragalus treat?" (C=1, T=1s, U=1), P=0.5×1+0.3×1+0.2×1=1 (10 / 10), "AI Shennong" will answer first; if a child asks an irrelevant question (C=0, T=3s, U=0.5), P=0.21, "Yaotong" will respond later.

[0071] V. Experience Operation and Optimization Storyline: Users receive the "Plague Treatment" task and need to identify 3 kinds of herbs (Astragalus membranaceus, Isatis indigotica, and Lonicera japonica). They scan the real plants with AR glasses → AI Shennong explains the characteristics → virtual processing (stir-frying according to steps) → submit the prescription, and the greenhouse LED screen displays the "patient recovery" animation.

[0072] Immersion assessment: S1=90 (positioning error 3cm), S2=85 (delay 50ms), S3=96 (accurate herbal knowledge), I=0.4×90+0.3×85+0.3×96=36+25.5+28.8=89.3≥80, high immersion.

[0073] Commercial Embedding and Fault Repair: After completing the task, the "Exchange Herbal Sachet at Cultural and Creative Store" (commercial trigger point) is triggered. The calculation is E=0.6×P+0.4×V, P=100 (after the climax of the plot), V=85 (the proportion of people going there), and E=0.6×100+0.4×85=60+34=94, which has a significant effect on attracting traffic. The AR anchor point is offset (error 6cm), and it is repositioned within 1 second by plant feature points, restoring the accuracy to 2cm.

[0074] VI. Performance Verification Form Table 2 Table 2 data is based on experience tests conducted by 80 parent-child visitors. Traditional botanical gardens rely on signage and guides, resulting in a herbal knowledge acquisition rate of less than 45%, and outdoor AR devices are almost unusable in strong light. This invention, through virtual-real fusion and multimodal interaction, increases the knowledge acquisition rate to over 85%, and achieves 90% clarity for outdoor virtual elements, allowing children to "learn through play." The frequency of parent-child collaboration has increased from 4 times / hour to 15 times, enhancing the interactive experience; equipment failure recovery time has been reduced from minutes to within 2 seconds, avoiding experience interruptions; and the efficiency of updating popular science content has increased to once every 2 weeks, combined with user-created "fun herbal stories," providing a new paradigm for natural science education.

[0075] Reference Figure 2 This diagram visually demonstrates the immersive advantages of this invention in different scenarios, specifically addressing the core pain point of insufficient immersion in traditional scripted games. Traditional scripted games, limited by live-action NPCs and static scenes, achieve an immersion level of only around 3 points in complex outdoor and open-air environments, failing to meet the needs of tourists. This invention, through deep integration of cultural elements, high-precision virtual-real alignment, and environmental adaptation technology, enhances the immersion level of each scene to over 7.5 points, with indoor scenes even reaching 9 points. Especially in outdoor and open-air scenes, this invention, through anti-interference equipment and dynamic projection, compensates for the shortcomings of traditional solutions in complex environments, allowing tourists to deeply immerse themselves in historical interaction and nature exploration, significantly enhancing the memorability and appeal of the cultural tourism experience.

[0076] Reference Figure 3 This diagram highlights the breakthroughs of this invention in content updates and user engagement, addressing the pain points of traditional scripted games: "content stagnation and low user retention." Traditional solutions rely on one-way updates from the operator, occurring only once every 1-2 months, resulting in a user repeat experience rate of less than 25%, making it difficult to maintain long-term appeal. This invention, through a UGC co-creation mechanism (user creation + AI + manual review), increases the content update frequency to once every 2 weeks, allowing over 35% of high-quality user content to be uploaded, significantly enriching the storyline library. Continuous content innovation extends the single experience duration to 90 minutes, increases the user repeat experience rate to 70%, and achieves a recommendation intention (NPS) score of over 75, indicating that tourists are not only willing to participate multiple times but also actively recommend it to others, bringing a continuous flow of visitors to the scenic area and achieving long-term sustainable development.

[0077] Reference Figure 4This diagram illustrates the technological advantages of the present invention's terminal, addressing the problems of poor performance and weak adaptability of traditional terminals. Traditional terminals suffer from low resolution, narrow field of view, and lack waterproofing for outdoor use, resulting in a poor user experience in complex scenarios. The present invention's terminal, through upgrades such as 4K resolution, a 120° field of view, and IP65 protection, significantly improves display quality and environmental adaptability. For example, the holographic cabin makes virtual NPCs clearer and more three-dimensional, while the waterproof terminal ensures normal use in rainy weather. Although the present invention's terminal is slightly more expensive (1.5-2.2 times), the performance improvement far outweighs the cost increase, and it is adaptable to multiple indoor and outdoor scenarios, avoiding the extra expense of frequent terminal replacements required by traditional solutions. Simultaneously, high-precision positioning and rapid response provide hardware support for virtual-real fusion, ensuring that visitors receive a consistent and clear immersive experience in different scenarios.

[0078] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for constructing and interacting with immersive script-based games that combines virtual and real elements, characterized in that... Includes the following steps: Scenario requirements analysis and cultural element mapping steps: For the target scenic spot / venue, extract regional cultural elements and construct a cultural element knowledge graph; Combine user profiles to determine the scenario theme and output a requirements document; The steps for constructing a virtual-real scene fusion are as follows: Laser scanning technology is used to model the real space and generate a 1:1 three-dimensional spatial model; based on the model, presentation terminals are deployed with a terminal spacing of ≤5m to ensure continuous interaction; the coordinates of virtual elements and real space are aligned through edge computing hosts with an anchor point positioning error of ≤3cm to generate a virtual-real fusion scene; Steps for creating and configuring AI NPC agents: Build AI NPC language logic based on the LLM model; integrate ASR, TTS, and action generation modules; configure NPC appearance according to the scene theme, set dialogue style, and bind cultural knowledge base to make NPC responses conform to cultural settings; Multimodal interaction rule design steps: Define user interaction methods, including voice commands, gesture operations, and AR gesture grasping; set up virtual and real interaction logic, that is, after the user touches the real plant, AR triggers a virtual plant information pop-up window and the AI ​​NPC explains in sync; formulate multi-user collaboration rules to support 2-8 people in a team, with the synchronization deviation of team users' actions ≤100ms and shared plot progress. Experience operation and dynamic optimization steps: After the user enters the scene, the edge computing host collects interaction data in real time, including user dwell time, number of interactions, and voice command recognition rate; push story tasks through handheld terminals, and the AI ​​NPC adjusts the story branches according to the user's response. When the user answers the plant characteristics correctly, a hidden story is triggered, and when the answer is wrong, clues are given; after the experience ends, user satisfaction scores are collected. Every 100 scores are accumulated, the AI ​​NPC dialogue model is fine-tuned, and the presentation position of virtual elements is optimized.

2. The method for constructing and interacting with a script-based immersive experience scene combining virtual and real elements as described in claim 1, characterized in that, It also includes a real-time scene immersion assessment step, using a formula. Calculate scene immersion, where I represents immersion. As a weight for scene integration, The score is given for the alignment accuracy of the coordinates of the virtual and real elements. As a weight for interaction smoothness, Score for interaction response latency. Weighting for cultural compatibility, The system scores the accuracy of AI NPC cultural representation. When I < 80, the system automatically adjusts the terminal position, reduces interaction latency, and updates the NPC cultural knowledge base.

3. The method for constructing and interacting with a script-based immersive experience scene combining virtual and real elements as described in claim 1, characterized in that, It also includes AI NPC interaction response priority scheduling steps, through formulas. Calculate the response priority, where P is the response priority. C represents the relevance weight of the plot, and C represents the relevance of the user command to the current plot. As a weighted average of waiting time, where T is the user's waiting time for a response. U represents the user's role weight, where U is the importance of the user's role in the team.

4. The method for constructing and interacting with a script-based immersive experience scene combining virtual and real elements as described in claim 1, characterized in that, The virtual-real scene fusion construction process also includes dynamic lighting adaptation: collecting real-world lighting data and adjusting the lighting parameters of virtual elements through edge computing hosts; for outdoor scenes, using light sensors to monitor lighting changes in real time, and automatically increasing the contrast of the water curtain projection and reducing the transparency of virtual elements when the light intensity is >800 lux.

5. The method for constructing and interacting with a script-based immersive experience scene combining virtual and real elements as described in claim 1, characterized in that, The creation and configuration steps of AI NPC intelligent agents also include multi-role collaboration logic: when a scene contains multiple AI NPCs, a dialogue model between NPCs is built. This model is based on dialogue state tracking technology, and the dialogue state is updated every 1 second. NPC collaboration rules are defined, with academic NPCs taking priority in answering professional questions and story NPCs taking priority in advancing the plot. When a user asks a question to multiple NPCs, the NPC that responds best is selected based on the role matching degree. The formula for calculating the role matching degree is: M = overlap between the NPC's knowledge base and the question × 0.8 + fit between the NPC's role setting and the question × 0.

2.

6. The method for constructing and interacting with a script-based immersive experience scene combining virtual and real elements as described in claim 1, characterized in that, The multimodal interaction rule design process also includes interaction intent fusion and judgment: using a weighted fusion algorithm to determine the intent of voice commands. Gestures and intentions Facial expressions and intentions Integration as the ultimate intention When the confidence of a single interaction intent is insufficient, it can be supplemented by combining other modalities.

7. The method for constructing and interacting with immersive script-based game scenarios that combine virtual and real elements, as described in claim 1, is characterized in that... The experience operation and dynamic optimization process also includes the integration of user-generated content (UGC): providing script fragment editing tools, allowing users to submit their creations to the platform; the platform adopts AI initial review + manual review, which is completed within 24 hours, and the approved content is marked with the "user co-creation" tag and included in the scene plot library; creators who use the content ≥ 100 times will be given a revenue sharing reward, with the reward amount = number of times the content is used × 0.1 yuan, and the content will be adapted to different display terminals to enrich the scene content.

8. The method for constructing and interacting with a script-based immersive experience scene combining virtual and real elements as described in claim 1, characterized in that, It also includes steps for embedding commercial nodes and driving traffic: setting commercial trigger points in the storyline, through formulas Calculate merchant exposure, where E is the exposure value. P represents the weight of the trigger point in the plot, where P is the importance of the trigger point in the plot. V represents the proportion of users who actually visit the merchant after the trigger is activated; the trigger point position is adjusted according to E.

9. The method for constructing and interacting with a script-based immersive experience scene combining virtual and real elements as described in claim 1, characterized in that, It also includes a self-healing step for virtual-real interaction faults: the edge computing host monitors the terminal status in real time, and automatically switches to the backup transparent LED screen when the communication interruption of the holographic cabin is detected; when the AR anchor point error is greater than 5cm, visual repositioning is initiated; the fault type and repair measures are recorded, and the fault diagnosis model is updated every 50 fault data points accumulated to reduce manual intervention.

10. The method for constructing and interacting with a script-based immersive experience scene combining virtual and real elements as described in claim 1, characterized in that, The scenario requirements analysis and cultural element mapping steps also include a cultural element suitability assessment: Three assessment indicators are constructed: cultural authenticity, user awareness, and interaction feasibility. The analytic hierarchy process (AHP) is used to determine the weights of each indicator, with cultural authenticity weighted at 0.5, user awareness at 0.3, and interaction feasibility at 0.

2. Each cultural element is scored across the three indicator dimensions, ranging from 1 to 10, and the suitability score is calculated. To score for authenticity, To score awareness, For feasibility score, only elements with A≥8 are retained.