A script-based virtual-real fusion guide method and system based on Beidou positioning

By combining BeiDou positioning and AI music composition technology with virtual-real fusion rendering algorithms, the problems of positioning accuracy and interactive experience in existing virtual-real fusion tour guide technology have been solved. This has enabled high-precision synchronization of virtual content with the real scene and personalized interaction, thereby improving the visitor experience and the efficiency of scenic area operation.

CN122261391APending Publication Date: 2026-06-23NANJING NICEBRIDGE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610347316.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-20
Publication Date
2026-06-23

Smart Images

  • Figure CN122261391A_ABST
    Figure CN122261391A_ABST
Patent Text Reader

Abstract

The application provides a script-based virtual-real fusion guide method and system based on Beidou positioning, and relates to the technical field of intelligent travel guide, which comprises filtering and calibrating the user position collected through Beidou positioning and collecting real scene pictures; an interactive framework is constructed, and the node content is triggered when the deviation between the user position and the node position is not greater than the deviation threshold; the virtual character is called according to the node content to explain the real scene pictures, and the interactive task is published in combination with the behavior data of the user; the emotion recognition is performed on the node content, the real scene pictures, the behavior data and the interactive task to generate emotion features, which are input into an AI composition model to output scene music; the interactive interface of the virtual character and the interactive task is fused with the real scene pictures based on a fusion rendering algorithm, and the guide picture is output and the scene music is played synchronously, so that the script-based virtual-real fusion immersive guide based on Beidou high-precision positioning can be realized, and the tour experience of tourists and the operation efficiency of the scenic spot can be greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent cultural tourism guide technology, and in particular to a scripted virtual-real fusion guide method and system based on BeiDou positioning. Background Technology

[0002] With the digital upgrade of the cultural and tourism industry, traditional tour guide methods can no longer meet tourists' needs for immersive and interactive tour experiences. As a result, virtual and real integrated tour guide technology has emerged and is gradually being applied to various cultural and tourism scenarios.

[0003] However, existing virtual-real fusion tour guide technologies have several significant shortcomings. First, most existing technologies rely on standard GPS positioning, which has a large positioning error and cannot achieve precise linkage between the visitor's location and virtual content. This results in a disconnect between the virtual scene and the actual scenic area, leading to a clunky interactive experience. Second, existing tour guide systems are mostly one-way information transmission models, lacking a scripted interactive framework and personalized experience design. This results in low visitor participation and fails to guide visitors to an orderly tour and a deeper understanding of the scenic area's culture. Third, in terms of creating scene atmosphere, most rely on fixed background music, failing to dynamically adapt original music to different scenes within the scenic area. Furthermore, music copyright traceability is difficult, easily leading to copyright disputes. Finally, existing virtual-real fusion tour guide technologies based on digital twins focus primarily on model construction and tour guide decision optimization within exhibition halls. They do not integrate BeiDou high-precision positioning to achieve scripted interaction in outdoor cultural tourism scenes, nor do they integrate integrated functions such as AI music composition and copyright traceability. These shortcomings result in limited functionality, poor adaptability, and an inadequate user experience.

[0004] Therefore, it is necessary to provide a scripted virtual-real fusion tour guide method and system based on BeiDou positioning to solve the above-mentioned technical problems. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a scripted virtual-real fusion navigation method and system based on BeiDou positioning, which solves the problems of insufficient positioning accuracy, monotonous interactive experience, poor scene atmosphere adaptability and copyright risks, and insufficient adaptability between virtual content and real scene in existing technologies.

[0006] This invention provides a scripted virtual-real fusion guided tour method based on BeiDou positioning, the method comprising: The system collects user location data in real time through a positioning unit that supports the BeiDou positioning system, filters and calibrates the user location data, and collects real-scene images in real time through user equipment. A scripted interaction framework is constructed. When the spatial deviation between the calibrated user location data and the script node position in the scripted interaction framework is not greater than a preset deviation threshold, the script node content corresponding to the script node position is triggered. Based on the script node content, the corresponding virtual character is retrieved from the scripted interaction framework, and the virtual character explains the real-world scene. The scripted interaction task is then issued in conjunction with the user's historical behavior data. An AI composition model is constructed to perform scene emotion recognition on the script node content, the real scene, the historical behavior data and the scripted interaction task, generate comprehensive emotion features, input the comprehensive emotion features into the AI ​​composition model, and output scene-adapted music. Based on the virtual-real fusion rendering algorithm, the interactive interface of the virtual character and the scripted interactive task is merged with the real scene to output a virtual-real fusion guide screen, and the scene-adapted music is played synchronously.

[0007] Preferably, the step of collecting user location data in real time through a positioning unit supporting the BeiDou positioning system and performing filtering and calibration processing on the user location data specifically includes: The positioning unit continuously acquires the user's latitude, longitude, and altitude data during the tour at a preset fixed acquisition frequency, generating raw user location data. The Kalman filter algorithm is used to perform signal filtering and noise suppression on the original user location data, and to reduce invalid location data caused by scenic area obstruction and multipath effect, thereby extracting valid location data. By combining pre-distributed fixed positioning reference points and using the standard position coordinates of the fixed positioning reference points as a reference, real-time differential calibration processing is performed on the effective position data, and the calibrated user position data is output.

[0008] Preferably, based on the scenic area's GIS geographic information and the content of the cultural tourism script creation, a lightweight interactive framework engine is used to construct the scripted interactive framework. The scripted interactive framework includes at least one preset script, which includes multiple script nodes. Each script node is configured with a corresponding script node location, script node content, virtual character identifier, task logic parameters, and incentive strategy parameters. The script node location corresponds one-to-one with the real-world location.

[0009] Preferably, the step of retrieving the corresponding virtual character from the scripted interaction framework based on the script node content, using the virtual character to explain the real-world scene, and issuing the scripted interaction task in conjunction with the user's historical behavior data specifically includes: The corresponding virtual character is retrieved from the scripted interaction framework according to the script node content, wherein the virtual character includes at least an explanation-type virtual character and a task-type virtual character; Based on a preset knowledge base, the virtual character is controlled to explain the knowledge of the real scene through voice synthesis and augmented reality tagging. The system acquires the user's historical behavior data and the scene features of the real-world image, and calculates the target difficulty parameter of the current task to be published using a task difficulty adaptive algorithm. The scripted interactive task is selected from the preset task template library according to the target difficulty parameter, and then the scripted interactive task is published by the task-type virtual character.

[0010] Preferably, the step of acquiring the user's historical behavior data and the scene features of the real-world image, and then calculating the target difficulty parameter of the current task to be published using a task difficulty adaptive algorithm. The corresponding calculation formula is as follows: In the formula, This represents a function for evaluating user capabilities. A quantification function representing the complexity of a scenario; This represents the spatiotemporal context adjustment factor; n represents the total number of historical tasks extracted from historical behavioral data. This represents the completion rate of the i-th historical task extracted from historical behavior data. ; This represents the response latency of the i-th historical task extracted from historical behavior data; Represents the time decay function; , This represents the preset weight coefficients, and satisfies... ; This represents the Sigmoid activation function; This represents the weight matrix of the pre-trained scenario complexity evaluation model; This represents the image depth features extracted from scene features S based on a convolutional neural network; This represents the bias term of the pre-trained scenario complexity evaluation model; This represents a real-time estimate of pedestrian density based on user location data L and the current time context T. This indicates the preset upper limit threshold for pedestrian density.

[0011] Preferably, the construction of the AI ​​composition model involves performing scene emotion recognition on the script node content, the real-world visuals, the historical behavioral data, and the scripted interactive tasks to generate comprehensive emotion features. These comprehensive emotion features are then input into the AI ​​composition model to output scene-adapted music. Specifically, this includes: Obtain the scene type tags in the script node content and the plot emotion tags in the scripted interactive task, and extract the scene features of the real scene and the emotional feedback features in the historical behavior data. The scene type label, the plot emotion label, the scene features, and the emotion feedback features are input into a pre-trained scene emotion recognition model, and the comprehensive emotion features of the current scene are calculated through multimodal feature fusion. The comprehensive emotional features are input into the AI ​​composition model based on the Transformer architecture. Combined with the preset scenic area cultural element library and music style library, the scene-adapted music is generated. The copyright of the scene-adapted music is registered and a copyright traceability code is generated through blockchain.

[0012] Preferably, the virtual-real fusion rendering algorithm, which merges the interactive interface of the virtual character and the scripted interactive task with the real-world scene to output a virtual-real fusion guide screen, specifically includes: The system acquires the real-world image at the current moment and the user's perspective information collected by the terminal sensor, the user's perspective information including the camera device orientation and field of view parameters; The virtual content to be integrated is determined from the scripted interaction framework. The virtual content includes at least the interaction interface of the virtual character and the scripted interaction task, and the anchor position of the virtual content in the real-world coordinate system is obtained. Based on the user location data, the user perspective information, and the anchor position of the virtual content, the target display area and target display ratio of the virtual content in the real scene are calculated through virtual-real fusion projection transformation; Based on the target display area and the target display ratio, a lightweight virtual rendering engine is invoked to render the virtual content in real time, and the rendered virtual content is then pixel-level blended with the real scene to output the virtual-real fusion guide screen.

[0013] A scripted virtual-real fusion tour guide system based on BeiDou positioning, the system comprising: The positioning acquisition module is used to acquire user location data in real time through a positioning unit that supports the BeiDou positioning system, filter and calibrate the user location data, and acquire real-time images through the user equipment. The script triggering module is used to construct a scripted interaction framework. When the spatial deviation between the calibrated user location data and the script node position in the scripted interaction framework is not greater than a preset deviation threshold, the script node content corresponding to the script node position is triggered. The character interaction module is used to retrieve the corresponding virtual character from the scripted interaction framework according to the script node content, explain the real scene through the virtual character, and issue scripted interaction tasks in combination with the user's historical behavior data. The music generation module is used to build an AI composition model, perform scene emotion recognition on the script node content, the real scene, the historical behavior data and the scripted interactive task, generate comprehensive emotion features, input the comprehensive emotion features into the AI ​​composition model, and output scene-adapted music. The fusion rendering module is used to fuse the virtual character and the interactive interface of the scripted interactive task with the real scene based on the virtual-real fusion rendering algorithm, output the virtual-real fusion guide screen, and play the scene-adapted music synchronously.

[0014] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes the steps of a scripted virtual-real fusion navigation method based on BeiDou positioning as described above.

[0015] A readable storage medium storing a computer program, characterized in that, when executed by a processor, the computer program is used to implement the steps of a scripted virtual-real fusion navigation method based on BeiDou positioning as described in any of the above claims.

[0016] Compared with related technologies, the script-based virtual-real fusion tour guide method and system based on BeiDou positioning provided by this invention has the following beneficial effects: This invention collects user location data in real time through a positioning unit supporting the BeiDou Navigation Satellite System, filters and calibrates the user location data, and collects real-world images through user devices. It constructs a scripted interactive framework; when the spatial deviation between the calibrated user location data and the script node position in the scripted interactive framework is no greater than a preset deviation threshold, the corresponding script node content is triggered. Based on the script node content, the corresponding virtual character is retrieved from the scripted interactive framework to explain the real-world scene, and scripted interactive tasks are published in conjunction with the user's historical behavior data. Scene emotion recognition is performed on the script node content, real-world scene, historical behavior data, and scripted interactive tasks to generate comprehensive emotion features. These comprehensive emotion features are input into an AI music composition model to output scene-adapted music. Based on a virtual-real fusion rendering algorithm, the interactive interface of the virtual character and scripted interactive tasks is merged with the real-world scene to output a virtual-real fusion guided tour, simultaneously playing scene-adapted music. This achieves a scripted virtual-real fusion immersive guided tour based on BeiDou high-precision positioning, significantly improving the visitor experience and scenic area operation efficiency.

[0017] This invention relies on the positioning unit of the BeiDou Navigation Satellite System to collect user location data and combines it with Kalman filtering algorithm and real-time differential calibration processing of fixed positioning reference points. This effectively eliminates positioning noise caused by obstructions and multipath effects within scenic areas, ensuring that the calibrated user location data can accurately trigger script node content. It solves the problem of virtual content being disconnected from the real scene due to positioning deviations in existing technologies, significantly improving the accuracy and reliability of guided tour interaction. This invention constructs a scripted interactive framework, combining narration-type and task-type virtual characters, and uses user historical behavior data to release scripted interactive tasks with adaptive difficulty. It also supports natural language interaction to trigger task updates, transforming traditional one-way information transmission into immersive two-way interaction, effectively enhancing visitor participation and the depth of their perception of the scenic area's culture. This invention generates comprehensive emotional features through multi-dimensional scene emotion recognition, driving an AI composition model to dynamically generate scene-adaptive music. It also uses blockchain technology to complete copyright registration, achieving accurate adaptation to different scene atmospheres and building a complete copyright traceability system, effectively avoiding copyright disputes and ensuring the compliant operation of guided tour services. This invention uses a virtual-real fusion rendering algorithm to perform pixel-level fusion of the interactive interface of virtual characters and scripted interactive tasks with real-world scenes. It dynamically adjusts the target display area and target display ratio of virtual content based on user perspective information to ensure a natural and smooth virtual-real fusion. Attached Figure Description

[0018] Figure 1 A flowchart of a scripted virtual-real fusion tour guide method based on BeiDou positioning provided in an embodiment of the present invention; Figure 2 A system block diagram of a scripted virtual-real fusion tour guide system based on BeiDou positioning provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] like Figure 1 The diagram shown is a flowchart of a scripted virtual-real fusion guided tour method based on BeiDou positioning provided by an embodiment of the present invention. Figure 1The execution entity of the method shown can be a software and / or hardware device. The execution entity of this application can include, but is not limited to, at least one of the following: user equipment, network equipment, etc. User equipment can include, but is not limited to, computers, smartphones, personal digital assistants (PDAs), and the aforementioned electronic devices. Network equipment can include, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of computers or network servers. Cloud computing is a type of distributed computing, consisting of a super virtual computer composed of a group of loosely coupled computers. This embodiment does not limit this. Steps S1 to S5 are detailed as follows: S1, real-time acquisition of user location data through a positioning unit supporting the BeiDou positioning system, filtering and calibration of the user location data, and real-time acquisition of real-scene images through user equipment; The process of collecting user location data in real time through a positioning unit supporting the BeiDou positioning system, and filtering and calibrating the user location data specifically includes: The positioning unit continuously acquires the user's latitude, longitude, and altitude data during the tour at a preset fixed acquisition frequency, generating raw user location data. The Kalman filter algorithm is used to perform signal filtering and noise suppression on the original user location data, and to reduce invalid location data caused by scenic area obstruction and multipath effect, thereby extracting valid location data. By combining pre-distributed fixed positioning reference points and using the standard position coordinates of the fixed positioning reference points as a reference, real-time differential calibration processing is performed on the effective position data, and the calibrated user position data is output.

[0021] The BeiDou Navigation Satellite System is a high-precision satellite navigation system used to provide stable satellite signal support for user location acquisition. The positioning unit is a hardware component integrating satellite signal reception, signal filtering, and location analysis functions. It is compatible with mobile devices such as smartphones and scenic area-specific guide terminals, and can operate stably in outdoor cultural and tourism scenarios such as mountaintops, forests, and ancient streets. The real-scene image is real-time video of the scenic area captured by the user. The fixed acquisition frequency refers to a pre-set constant data acquisition interval used to ensure the temporal continuity of location data and meet the real-time requirements of plot point triggering. Latitude and longitude data and altitude data together constitute the user's three-dimensional spatial location information within the scenic area. The raw user location data is the initial, unprocessed location data directly collected by the positioning unit, containing error information caused by environmental interference.

[0022] Furthermore, the Kalman filter algorithm is used to suppress random noise in the raw user location data, improving data stability. Signal filtering is the process of filtering out interference signals in the raw user location data. Noise suppression reduces positioning data fluctuations caused by electromagnetic interference, signal attenuation, and other factors. Multipath effect is the positioning deviation caused by the superposition of satellite signals reflected from the ground, buildings, etc., and direct signals. Invalid location data is location data that exceeds the reasonable positioning error range due to obstructions from scenic spots and multipath effects. Valid location data refers to location data that meets calibration requirements after filtering.

[0023] It is understandable that fixed positioning reference points are fixed points with known precise coordinates evenly distributed within the scenic area, used to provide calibration references. Standard position coordinates are precise three-dimensional coordinate data of the fixed positioning reference points determined by professional measurement. Real-time differential calibration processing is a method of comparing and correcting deviations between the effective position data and the standard position coordinates in real time to further improve positioning accuracy. The calibrated user position data is high-precision position data with a positioning deviation of no more than 3 meters after filtering and differential calibration processing.

[0024] In practical applications, taking a guided tour of an ancient town scenic area as an example, the scenic area sets up a fixed positioning reference point every 500 meters, and its standard position coordinates are determined by professional measurement. After tourists start the tour through a WeChat mini-program, the positioning unit built into the phone continuously acquires the latitude, longitude, and altitude data of the tourists during their tour at a fixed frequency of 1 second, generating raw user location data. When tourists pass near the ancient bridge in the ancient town, the multipath effect caused by tree obstruction and building reflections is quickly suppressed by the Kalman filter algorithm, eliminating invalid data and extracting valid location data. Then, it combines the standard position coordinates of the nearby fixed positioning reference points for real-time differential calibration processing, outputting user location data with a positioning deviation of no more than 3 meters.

[0025] S2, construct a scripted interaction framework. When the spatial deviation between the calibrated user location data and the script node position in the scripted interaction framework is not greater than a preset deviation threshold, trigger the script node content corresponding to the script node position. Based on the scenic area's GIS geographic information and the content of cultural tourism script creation, a lightweight interactive framework engine is used to construct the scripted interactive framework. The scripted interactive framework includes at least one preset script, which includes multiple script nodes. Each script node is configured with a corresponding script node location, script node content, virtual character identifier, task logic parameters, and incentive strategy parameters. The script node location corresponds one-to-one with the real-world location.

[0026] The script-based interaction framework is a logical architecture built upon the scenic area's GIS (Geographic Information System) geographic information and cultural tourism script creation content. It defines the interactive nodes, plot content, and user guidance paths during the tour. A preset script is a complete narrative unit within the script-based interaction framework, containing multiple script nodes organized according to the tour route. Each script node is the smallest triggering unit in the script-based interaction framework, and each script node is configured with its location, content, virtual character identifier, task logic parameters, and incentive strategy parameters.

[0027] It is understandable that the script node location refers to the spatial coordinates corresponding to the script node, which correspond one-to-one with the actual location within the scenic area. The script node content includes the explanatory information triggered by the script node, scripted interactive tasks, and multimedia resources. The virtual character identifier is used to uniquely identify the explanatory or task-oriented virtual character invoked by the script node. Task logic parameters define the execution conditions, process, and objectives of the scripted interactive task. Incentive strategy parameters specify the points, rewards, or story unlocking rules that the user receives upon completion of the task. Spatial deviation refers to the Euclidean distance between the calibrated user location data and the script node location. The preset deviation threshold is a spatial range limit set to trigger script node content. When the spatial deviation between the user location data and the script node location does not exceed this threshold (e.g., a preset deviation threshold of 2 meters), the system determines that the user has entered the node area and automatically invokes the corresponding script node content.

[0028] S3, retrieve the corresponding virtual character from the scripted interaction framework according to the script node content, explain the real scene through the virtual character, and issue the scripted interaction task in combination with the user's historical behavior data; The step of retrieving the corresponding virtual character from the scripted interaction framework based on the script node content, using the virtual character to explain the real-world scene, and issuing the scripted interaction task in conjunction with the user's historical behavior data specifically includes: The corresponding virtual character is retrieved from the scripted interaction framework according to the script node content, wherein the virtual character includes at least an explanation-type virtual character and a task-type virtual character; Based on a preset knowledge base, the virtual character is controlled to explain the knowledge of the real scene through voice synthesis and augmented reality tagging. The system acquires the user's historical behavior data and the scene features of the real-world image, and calculates the target difficulty parameter of the current task to be published using a task difficulty adaptive algorithm. The scripted interactive task is selected from the preset task template library according to the target difficulty parameter, and then the scripted interactive task is published by the task-type virtual character.

[0029] The virtual characters are digital figures generated based on a scripted interactive framework, used for immersive interaction with users. They are specifically divided into narration-based virtual characters and task-based virtual characters. Narration-based virtual characters primarily function to disseminate knowledge; their narration content comes from a pre-built, structured database containing information such as the scenic area's history, culture, and natural features. During the narration, speech synthesis technology is used to convert text information into natural speech for broadcast, and augmented reality tags are used to overlay text, icons, or 3D annotations onto the real-world scene, achieving a fusion of audiovisual elements for the narration effect.

[0030] Following this, the task-oriented virtual character primarily functions as an interactive task publisher and manager. The published scripted interactive tasks are designed based on scenic area scripts, and their difficulty is dynamically calculated by an adaptive task difficulty algorithm based on the user's historical behavior data and the scene characteristics of the real-world visuals. Historical behavior data consists of records of user actions accumulated during their visit, including task completion status, response time, and interaction trajectory. Scene characteristics of the real-world visuals are visual information extracted from the current real-world scene through image recognition, including scene type, object distribution, and spatial complexity.

[0031] Understandably, the task difficulty adaptive algorithm is a computational logic that integrates user ability, scene complexity, and spatiotemporal context. The preset task template library is a pre-configured collection of task templates containing different difficulty levels and types. The task-oriented virtual character selects a suitable template from this library based on the target difficulty parameter and instantiates it into a specific task to be released to the user.

[0032] Using the above method, explanation-type virtual characters and task-type virtual characters are retrieved based on the script node content. The former uses a preset explanation knowledge base to achieve immersive knowledge explanation through voice synthesis and augmented reality tags, while the latter combines the user's historical behavior data and the scene characteristics of the real scene. The target difficulty parameter is calculated by the task difficulty adaptive algorithm, and the scripted interactive task is selected and published from the preset task template library to achieve personalized dynamic adaptation of scripted interactive tasks.

[0033] The process involves acquiring the user's historical behavior data and the scene features of the real-world image, and then using a task difficulty adaptive algorithm to calculate the target difficulty parameter of the task to be published. The corresponding calculation formula is as follows: In the formula, This represents a function for evaluating user capabilities. A quantification function representing the complexity of a scenario; This represents the spatiotemporal context adjustment factor; n represents the total number of historical tasks extracted from historical behavioral data. This represents the completion rate of the i-th historical task extracted from historical behavior data. ; This represents the response latency of the i-th historical task extracted from historical behavior data; Represents the time decay function; , This represents the preset weight coefficients, and satisfies... ; This represents the Sigmoid activation function; This represents the weight matrix of the pre-trained scenario complexity evaluation model; This represents the image depth features extracted from scene features S based on a convolutional neural network; This represents the bias term of the pre-trained scenario complexity evaluation model; This represents a real-time estimate of pedestrian density based on user location data L and the current time context T. This indicates the preset upper limit threshold for pedestrian density.

[0034] The user capability assessment function quantifies a user's task execution ability based on a weighted average of historical task completion rates and response latency. The scene complexity quantification function extracts image depth features from real-world scenes using a convolutional neural network, maps these features through a scene complexity assessment model, and outputs the result to characterize the visual complexity of the scene. The spatiotemporal context adjustment factor reflects the degree of interference between the current spatiotemporal environment and user interaction by comparing the real-time estimated pedestrian density with a preset upper limit threshold. The target difficulty parameter is calculated by geometrically averaging the above three factors, enabling the difficulty of the task to dynamically adapt to individual user capabilities, inherent scene characteristics, and real-time environmental changes.

[0035] Furthermore, the total number of historical tasks is the number of valid tasks extracted from users' historical behavior data. The completion rate of a historical task is a quantified value of the user's degree of completion, ranging from 0 to 1, where 0 represents incomplete and 1 represents complete completion. The response latency of a historical task is the time interval between receiving and completing the task, reflecting task execution efficiency. The time decay function is used to weaken the impact of historical tasks with long response latencies on user ability assessment; the longer the latency, the lower the weight on the assessment result. Preset weighting coefficients are pre-defined proportional parameters based on the needs of the navigation scenario, one emphasizing the impact of historical task completion rate on user ability, and the other emphasizing the impact of response latency. The Sigmoid activation function is used to ensure the stability and comparability of the scenario complexity values.

[0036] Understandably, the scene complexity assessment model is a convolutional neural network model pre-trained on expert-annotated scenic area scene image data. It automatically extracts visual features from real-world images and outputs quantifiable values ​​for scene complexity. The weight matrix is ​​a parameter matrix trained using sample data, used to weight the extracted scene features, highlighting the influence of key scene factors. Image depth features are feature data reflecting attributes such as the spatial structure and detail levels of the real-world image. The bias term is a parameter used in conjunction with the weight matrix to fine-tune the quantifiable results of scene complexity, improving the accuracy of the scene complexity assessment model.

[0037] The real-time crowd density estimate based on user location data and current time context statistics combines the number of tourists per unit area at the user's location within the scenic area and the current time period (e.g., holidays, weekdays) to reflect the level of crowding. The preset upper limit threshold for crowd density is a pre-defined maximum reasonable crowd density for the scenic area, used to normalize the real-time crowd density and ensure that the spatiotemporal context adjustment factor is within a reasonable range.

[0038] In practical applications, taking the stream plot node of the "Forest Treasure Hunt" scenario in a mountain scenic area as an example, four historical tasks were extracted from user historical behavior data, with completion rates of 0.9, 1, 0.85, and 0.95, and response delays of 3 seconds, 2 seconds, 4 seconds, and 2.5 seconds, respectively. The preset weight coefficients were 0.6 and 0.4, respectively, and the result was 0.58 calculated using the user ability evaluation function. The current real-world scene features dense vegetation and crisscrossing paths. After extracting image depth features based on a convolutional neural network, the scene complexity quantification function outputs 0.75. Combining user location data and weekday time context, the real-time pedestrian density was calculated to be 60 people / 100 square meters, with a preset upper limit threshold of 100 and a spatiotemporal context adjustment factor of 0.6. After calculating the target difficulty parameters using a task difficulty adaptive algorithm, virtual props were selected from a preset task template library to find tasks that combined with real-world plant recognition. These tasks were then issued by a task-type virtual character, with the difficulty adapting to the user's ability and the scene conditions.

[0039] Also includes: When the semantic relevance between the user's question and the scene content of the real-world image is identified by natural language processing technology and exceeds a preset similarity threshold, the task-oriented virtual character is triggered and a pre-trained large language model is invoked. Based on a large language model, user location data, scripted interaction tasks, and task logic parameters in the scripted interaction framework are used as input prompts to output task response information. Logically associate scripted interactive tasks with implicit task cues in task response information, collect user interaction behavior data in real time, calculate the matching degree between interaction behavior data and implicit task cues, and generate interaction matching degree. If the interaction matching degree is greater than the preset matching degree threshold, it is determined that the user has completed the target interaction action based on the task response information, triggering the start or status update of the scripted interaction task.

[0040] Natural Language Processing (NLP) technology is used to identify the semantic relationship between user queries and scene content. Queries are questions or inquiries posed by users during browsing, either via voice or text. Semantic relevance is an indicator that quantifies the degree of semantic fit between user queries and scene content. A preset similarity threshold is a pre-defined critical value used to determine whether to trigger an interactive response. A pre-trained large language model is a model trained on massive amounts of data, possessing natural language understanding and generation capabilities, used to output response information that fits the current scene and task. Input prompts are instructions that integrate user location data, scripted interactive tasks, and task logic parameters to guide the large language model in generating responses. Task response information is the content output by the large language model that includes task clues, answers, or guidance.

[0041] Next, implicit task cues are key information hidden within the task response information that guides the user to complete the target interactive action. Logical association is the process of establishing a correspondence between scripted interactive tasks and implicit task cues, ensuring that the cues accurately point to the task completion path. Interactive behavior data is the action data generated by the user after receiving the task response information (such as clicks, taking photos, voice replies, etc.). Matching degree calculation is the process of quantifying the degree of fit between interactive behavior data and implicit task cues.

[0042] Furthermore, the preset matching threshold is a pre-defined critical value for determining whether a user has completed the target interactive action. The target interactive action is a specific operation required to complete a scripted interactive task. When the interaction matching degree exceeds the preset matching threshold, it is determined that the user has completed the target interactive action, thereby triggering the initiation (when not initiated) or status update (when initiated) of the scripted interactive task, thus advancing the tour's plot.

[0043] Through the above methods, natural language interaction between users and the system is achieved. By leveraging a large language model to generate task response information that fits the scene and task, implicit task cues are used to accurately guide users to complete target interactive actions, ensuring the orderly progress of scripted interactive tasks. This not only enhances user participation and the immersive experience of the guided tour, but also adapts to personalized tour needs and ensures the continuity and flexibility of the guided tour process.

[0044] S4, construct an AI composition model, perform scene emotion recognition on the script node content, the real scene, the historical behavior data and the scripted interactive task, generate comprehensive emotion features, input the comprehensive emotion features into the AI ​​composition model, and output scene-adapted music; The construction of the AI ​​music composition model involves performing scene emotion recognition on the script node content, the real-world visuals, the historical behavioral data, and the scripted interactive tasks to generate comprehensive emotion features. These comprehensive emotion features are then input into the AI ​​music composition model to output scene-appropriate music. Specifically, this includes: Obtain the scene type tags in the script node content and the plot emotion tags in the scripted interactive task, and extract the scene features of the real scene and the emotional feedback features in the historical behavior data. The scene type label, the plot emotion label, the scene features, and the emotion feedback features are input into a pre-trained scene emotion recognition model, and the comprehensive emotion features of the current scene are calculated through multimodal feature fusion. The comprehensive emotional features are input into the AI ​​composition model based on the Transformer architecture. Combined with the preset scenic area cultural element library and music style library, the scene-adapted music is generated. The copyright of the scene-adapted music is registered and a copyright traceability code is generated through blockchain.

[0045] The scene type tags are pre-defined tags within the script nodes that identify the category of the current real-world scene, such as mountaintop, ancient bridge, forest, and exhibition hall, used to quickly define the scene's basic attributes. The plot emotion tags are emotional attribute tags set within the scripted interactive tasks, corresponding to the plot's atmosphere, such as suspense, joy, solemnity, and tranquility, clearly indicating the plot's emotional direction. The scene characteristics of the real-world footage are environmental attribute information extracted from real-time captured footage, including vegetation density, architectural style, light intensity, and spatial openness, reflecting the scene's visual and environmental features. The emotional feedback characteristics of historical behavioral data are indicators extracted from the user's past interactive data that reflect the user's emotional state, such as task completion speed, interaction frequency, and dwell time, indirectly reflecting the user's emotional response to the scene.

[0046] Specifically, the scene emotion recognition model is a model trained with multimodal scene and emotion data. Its function is to integrate different types of feature data to accurately determine the overall emotion of the current scene. Multimodal feature fusion is the process of integrating feature data from different dimensions, such as scene type labels, plot emotion labels, scene features, and emotion feedback features, to generate comprehensive feature data with a unified dimension. Comprehensive emotion features are used to quantify the data of the current scene and the user's emotional state, including scene atmosphere attributes and user emotional tendencies.

[0047] Following this, the AI ​​composition model is a music generation model trained on an authorized music database. It possesses the ability to create original music by combining scene and emotional characteristics. The Transformer architecture is the foundational technical architecture of the AI ​​composition model, supporting its efficient learning of music composition rules and stylistic features. The scenic area cultural element library is a resource repository storing unique information such as scenic area names, historical anecdotes, natural features, and regional cultural symbols. The music style library is a resource repository storing the composition parameters and characteristics of different music styles (such as grand, soothing, rustic, and lively).

[0048] Understandably, scene-adapted music is original music generated by an AI composition model that combines comprehensive emotional characteristics, scenic area cultural elements, and musical style, resulting in a high degree of consistency with the current scene and emotional tone of the storyline. Blockchain technology is used for copyright registration and traceability, possessing the characteristics of immutable and traceable data. Copyright registration is the process of recording and solidifying information such as the generation time of the AI-generated music, scene information, AI model parameters, and the proportion of human involvement in creation. The copyright traceability code is a unique identifier generated after copyright registration, used by tourists and scenic areas to query music copyright ownership and generation details, ensuring copyright compliance.

[0049] In practical applications, taking the museum bronze exhibition hall plot node of the "Cultural Heritage Tracing" script as an example, the scene type label "exhibition hall" and the plot emotion label "solemn" are first obtained. Scene features such as the layout of cultural relics and the softness of lighting in the real scene are extracted, as well as emotional feedback features such as long dwell time and detailed interactive operations in the user's historical behavior data. The four types of features are input into a pre-trained scene emotion recognition model, and a comprehensive emotion feature is generated through multimodal feature fusion calculation. After this feature is input into an AI music composition model based on the Transformer architecture, the model combines the relevant historical allusions of bronze cultural relics in the scenic area's cultural element library and the solemn and elegant style parameters in the music style library to generate scene-appropriate music within 5 seconds. At the same time, copyright registration is completed through blockchain, generating a unique copyright traceability code to ensure that the music is appropriate for the scene and compliant with copyright.

[0050] S5, based on the virtual-real fusion rendering algorithm, the virtual character and the interactive interface of the scripted interactive task are merged with the real scene, and the virtual-real fusion guide screen is output, and the scene-adapted music is played synchronously.

[0051] The virtual-real fusion rendering algorithm merges the virtual character and the interactive interface of the scripted interactive task with the real-world scene to output the virtual-real fusion guide screen, specifically including: The system acquires the real-world image at the current moment and the user's perspective information collected by the terminal sensor, the user's perspective information including the camera device orientation and field of view parameters; The virtual content to be integrated is determined from the scripted interaction framework. The virtual content includes at least the interaction interface of the virtual character and the scripted interaction task, and the anchor position of the virtual content in the real-world coordinate system is obtained. Based on the user location data, the user perspective information, and the anchor position of the virtual content, the target display area and target display ratio of the virtual content in the real scene are calculated through virtual-real fusion projection transformation; Based on the target display area and the target display ratio, a lightweight virtual rendering engine is invoked to render the virtual content in real time, and the rendered virtual content is then pixel-level blended with the real scene to output a virtual-real fusion tour screen.

[0052] The virtual-real fusion rendering algorithm is used to accurately overlay virtual content with real-world images, eliminating the visual disconnect between virtual and reality and outputting a natural and smooth virtual-real fusion tour view. The current real-world image is a live video feed of the scenic area captured by the user through their mobile device, ensuring the image is synchronized with the user's viewing perspective. The terminal sensors are built-in components such as cameras and gyroscopes in the mobile device, used to collect data related to the user's perspective. User perspective information reflects parameters reflecting the user's viewing angle. The camera orientation is the direction the terminal lens points, and the field of view parameter is the range of vision the lens can cover; both together determine the display angle of the virtual content within the real-world image.

[0053] Next, virtual content consists of digital elements superimposed on real-world visuals. It includes at least virtual characters with independent interactive logic and interactive interfaces that carry out scripted interactive tasks. The real-world coordinate system is a three-dimensional coordinate system used to locate the virtual content in real-world space, ensuring that the anchor position of the virtual content precisely corresponds to the real-world space. The anchor position is a pre-set, fixed position of the virtual content within the real-world coordinate system, guaranteeing a stable association between the virtual content and a specific area of ​​the real-world scene.

[0054] Understandably, virtual-real fusion projection transformation is a process that converts the anchor position of virtual content into corresponding display coordinates in the real-world scene, ensuring that the virtual content adapts to the real-world perspective. The target display area is the specific display range of the virtual content in the real-world scene, and the target display ratio is the size adaptation ratio of the virtual content in the real-world scene. Both work together to ensure that the virtual content does not exceed the screen area and that the proportions are harmonious. The lightweight virtual rendering engine is a rendering tool adapted for mobile terminals such as smartphones and tablets, which can reduce device performance requirements while ensuring smooth rendering of virtual images.

[0055] Furthermore, real-time rendering is the process of instantly generating and updating virtual content, ensuring that the virtual content adjusts synchronously with changes in the user's perspective and position. Pixel-level blending is a processing method that precisely overlays the rendered virtual content with the real-world scene at the pixel level, eliminating the unnaturalness of the overlay edges. The virtual-real fusion tour guide screen is the output tour guide screen that combines the realism of the real-world scene with the interactivity of the virtual content.

[0056] Through the above methods, the virtual content and real-world scenes are accurately adapted and integrated. The target display area and target display ratio of the virtual content are dynamically adjusted based on the user's location data and perspective information. The lightweight virtual rendering engine ensures real-time rendering effects. Pixel-level fusion eliminates the visual disharmony between the virtual and real scenes, outputting a natural and smooth virtual-real fusion tour guide screen, enhancing the immersive experience and visual effects of the tour guide.

[0057] like Figure 2 The diagram shown is a system block diagram of a scripted virtual-real fusion tour guide system based on BeiDou positioning provided by an embodiment of the present invention. The system includes: The positioning acquisition module is used to acquire user location data in real time through a positioning unit that supports the BeiDou positioning system, filter and calibrate the user location data, and acquire real-time images through the user equipment. The script triggering module is used to construct a scripted interaction framework. When the spatial deviation between the calibrated user location data and the script node position in the scripted interaction framework is not greater than a preset deviation threshold, the script node content corresponding to the script node position is triggered. The character interaction module is used to retrieve the corresponding virtual character from the scripted interaction framework according to the script node content, explain the real scene through the virtual character, and issue scripted interaction tasks in combination with the user's historical behavior data. The music generation module is used to build an AI composition model, perform scene emotion recognition on the script node content, the real scene, the historical behavior data and the scripted interactive task, generate comprehensive emotion features, input the comprehensive emotion features into the AI ​​composition model, and output scene-adapted music. The fusion rendering module is used to fuse the virtual character and the interactive interface of the scripted interactive task with the real scene based on the virtual-real fusion rendering algorithm, output the virtual-real fusion guide screen, and play the scene-adapted music synchronously.

[0058] Figure 2 The apparatus of the illustrated embodiment can be used to perform corresponding actions. Figure 1 The steps in the method embodiments shown are implemented in a similar manner and have similar technical effects, and will not be repeated here.

[0059] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes the steps of a scripted virtual-real fusion navigation method based on BeiDou positioning as described above.

[0060] like Figure 3 The diagram shown is a hardware structure schematic of an electronic device according to an embodiment of the present invention. The electronic device 30 includes: a processor 31, a memory 32, and a computer program; wherein... The memory 32 is used to store the computer program, and the memory may also be flash memory. The computer program is, for example, an application program or functional module that implements the above method.

[0061] Processor 31 is configured to execute the computer program stored in the memory to implement the various steps performed by the device in the above method. For details, please refer to the relevant descriptions in the preceding method embodiments.

[0062] Alternatively, the memory 32 can be either standalone or integrated with the processor 31.

[0063] When the memory 32 is a device independent of the processor 31, the device may further include: Bus 33 is used to connect the memory 32 and the processor 31.

[0064] A readable storage medium storing a computer program, which, when executed by a processor, is used to implement the steps of a scripted virtual-real fusion navigation method based on BeiDou positioning as described above.

[0065] The readable storage medium can be a computer storage medium or a communication medium. A communication medium includes any medium that facilitates the transfer of computer programs from one location to another. A computer storage medium can be any available medium accessible to a general-purpose or special-purpose computer. For example, a readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application-Specific Integrated Circuit (ASIC). Alternatively, the ASIC can be located in a user equipment. Of course, the processor and the readable storage medium can also exist as discrete components in a communication device. The readable storage medium can be a read-only memory (ROM), random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0066] The present invention also provides a program product including executable instructions stored in a readable storage medium. At least one processor of the device can read the executable instructions from the readable storage medium, and the at least one processor executes the executable instructions to cause the device to implement the methods provided in the various embodiments described above.

[0067] In the embodiments of the above-described device, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly manifested as execution by a hardware processor, or execution by a combination of hardware and software modules within the processor.

[0068] Through the above embodiments, this invention collects user location data in real time through a positioning unit supporting the BeiDou positioning system, filters and calibrates the user location data, and collects real-world images through user devices in real time; it constructs a scripted interactive framework, and when the spatial deviation between the calibrated user location data and the script node position in the scripted interactive framework is not greater than a preset deviation threshold, it triggers the script node content corresponding to the script node position; it retrieves the corresponding virtual character from the scripted interactive framework based on the script node content, and uses the virtual character to explain the real-world image, and publishes scripted interactive tasks in combination with the user's historical behavior data; it performs scene emotion recognition on the script node content, real-world image, historical behavior data, and scripted interactive tasks, generates comprehensive emotion features, inputs the comprehensive emotion features into an AI music composition model, and outputs scene-adapted music; based on a virtual-real fusion rendering algorithm, it merges the interactive interface of the virtual character and scripted interactive tasks with the real-world image, outputs a virtual-real fusion tour guide screen, and plays scene-adapted music simultaneously, thereby realizing a scripted virtual-real fusion immersive tour guide based on BeiDou high-precision positioning, which greatly improves the visitor experience and the operational efficiency of the scenic area.

[0069] This invention relies on the positioning unit of the BeiDou Navigation Satellite System to collect user location data and combines it with Kalman filtering algorithm and real-time differential calibration processing of fixed positioning reference points. This effectively eliminates positioning noise caused by obstructions and multipath effects within scenic areas, ensuring that the calibrated user location data can accurately trigger script node content. It solves the problem of virtual content being disconnected from the real scene due to positioning deviations in existing technologies, significantly improving the accuracy and reliability of guided tour interaction. This invention constructs a scripted interactive framework, combining narration-type and task-type virtual characters, and uses user historical behavior data to release scripted interactive tasks with adaptive difficulty. It also supports natural language interaction to trigger task updates, transforming traditional one-way information transmission into immersive two-way interaction, effectively enhancing visitor participation and the depth of their perception of the scenic area's culture. This invention generates comprehensive emotional features through multi-dimensional scene emotion recognition, driving an AI composition model to dynamically generate scene-adaptive music. It also uses blockchain technology to complete copyright registration, achieving accurate adaptation to different scene atmospheres and building a complete copyright traceability system, effectively avoiding copyright disputes and ensuring the compliant operation of guided tour services. This invention uses a virtual-real fusion rendering algorithm to perform pixel-level fusion of the interactive interface of virtual characters and scripted interactive tasks with real-world scenes. It dynamically adjusts the target display area and target display ratio of virtual content based on user perspective information to ensure a natural and smooth virtual-real fusion.

[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A script-based virtual-real fusion guided tour method based on BeiDou positioning, characterized in that, The method includes: The system collects user location data in real time through a positioning unit that supports the BeiDou positioning system, filters and calibrates the user location data, and collects real-scene images in real time through user equipment. A scripted interaction framework is constructed. When the spatial deviation between the calibrated user location data and the script node position in the scripted interaction framework is not greater than a preset deviation threshold, the script node content corresponding to the script node position is triggered. Based on the script node content, the corresponding virtual character is retrieved from the scripted interaction framework, and the virtual character explains the real-world scene. The scripted interaction task is then issued in conjunction with the user's historical behavior data. An AI composition model is constructed to perform scene emotion recognition on the script node content, the real scene, the historical behavior data and the scripted interaction task, generate comprehensive emotion features, input the comprehensive emotion features into the AI ​​composition model, and output scene-adapted music. Based on the virtual-real fusion rendering algorithm, the interactive interface of the virtual character and the scripted interactive task is merged with the real scene to output a virtual-real fusion guide screen, and the scene-adapted music is played synchronously.

2. The script-based virtual-real fusion guided tour method based on BeiDou positioning according to claim 1, characterized in that, The process of collecting user location data in real time through a positioning unit supporting the BeiDou positioning system, and filtering and calibrating the user location data specifically includes: The positioning unit continuously acquires the user's latitude, longitude, and altitude data during the tour at a preset fixed acquisition frequency, generating raw user location data. The Kalman filter algorithm is used to perform signal filtering and noise suppression on the original user location data, and to reduce invalid location data caused by scenic area obstruction and multipath effect, thereby extracting valid location data. By combining pre-distributed fixed positioning reference points and using the standard position coordinates of the fixed positioning reference points as a reference, real-time differential calibration processing is performed on the effective position data, and the calibrated user position data is output.

3. The script-based virtual-real fusion guided tour method based on BeiDou positioning according to claim 1, characterized in that, Based on the scenic area's GIS geographic information and the content of cultural tourism script creation, a lightweight interactive framework engine is used to construct the scripted interactive framework. The scripted interactive framework includes at least one preset script, which includes multiple script nodes. Each script node is configured with a corresponding script node location, script node content, virtual character identifier, task logic parameters, and incentive strategy parameters. The script node location corresponds one-to-one with the real-world location.

4. The script-based virtual-real fusion guided tour method based on BeiDou positioning according to claim 1, characterized in that, The step of retrieving the corresponding virtual character from the scripted interaction framework based on the script node content, using the virtual character to explain the real-world scene, and issuing the scripted interaction task in conjunction with the user's historical behavior data specifically includes: The corresponding virtual character is retrieved from the scripted interaction framework according to the script node content, wherein the virtual character includes at least an explanation-type virtual character and a task-type virtual character; Based on a preset knowledge base, the virtual character is controlled to explain the knowledge of the real scene through voice synthesis and augmented reality tagging. The system acquires the user's historical behavior data and the scene features of the real-world image, and calculates the target difficulty parameter of the current task to be published using a task difficulty adaptive algorithm. The scripted interactive task is selected from the preset task template library according to the target difficulty parameter, and then the scripted interactive task is published by the task-type virtual character.

5. The script-based virtual-real fusion guided tour method based on BeiDou positioning according to claim 4, characterized in that, The process involves acquiring the user's historical behavior data and the scene features of the real-world image, and then using a task difficulty adaptive algorithm to calculate the target difficulty parameter of the task to be published. The corresponding calculation formula is as follows: In the formula, This represents a function for evaluating user capabilities. A quantification function representing the complexity of a scenario; This represents the spatiotemporal context adjustment factor; n represents the total number of historical tasks extracted from historical behavioral data. This represents the completion rate of the i-th historical task extracted from historical behavior data. ; This represents the response latency of the i-th historical task extracted from historical behavior data; Represents the time decay function; , This represents the preset weight coefficients, and satisfies... ; This represents the Sigmoid activation function; This represents the weight matrix of the pre-trained scenario complexity evaluation model; This represents the image depth features extracted from scene features S based on a convolutional neural network; This represents the bias term of the pre-trained scenario complexity evaluation model; This represents a real-time estimate of pedestrian density based on user location data L and the current time context T. This indicates the preset upper limit threshold for pedestrian density.

6. The script-based virtual-real fusion guided tour method based on BeiDou positioning according to claim 1, characterized in that, The construction of the AI ​​music composition model involves performing scene emotion recognition on the script node content, the real-world visuals, the historical behavioral data, and the scripted interactive tasks to generate comprehensive emotion features. These comprehensive emotion features are then input into the AI ​​music composition model to output scene-appropriate music. Specifically, this includes: Obtain the scene type tags in the script node content and the plot emotion tags in the scripted interactive task, and extract the scene features of the real scene and the emotional feedback features in the historical behavior data. The scene type label, the plot emotion label, the scene features, and the emotion feedback features are input into a pre-trained scene emotion recognition model, and the comprehensive emotion features of the current scene are calculated through multimodal feature fusion. The comprehensive emotional features are input into the AI ​​composition model based on the Transformer architecture. Combined with the preset scenic area cultural element library and music style library, the scene-adapted music is generated. The copyright of the scene-adapted music is registered and a copyright traceability code is generated through blockchain.

7. The script-based virtual-real fusion guided tour method based on BeiDou positioning according to claim 1, characterized in that, The virtual-real fusion rendering algorithm merges the virtual character and the interactive interface of the scripted interactive task with the real-world scene to output a virtual-real fusion guide screen, specifically including: The system acquires the real-world image at the current moment and the user's perspective information collected by the terminal sensor, the user's perspective information including the camera device orientation and field of view parameters; The virtual content to be integrated is determined from the scripted interaction framework. The virtual content includes at least the interaction interface of the virtual character and the scripted interaction task, and the anchor position of the virtual content in the real-world coordinate system is obtained. Based on the user location data, the user perspective information, and the anchor position of the virtual content, the target display area and target display ratio of the virtual content in the real scene are calculated through virtual-real fusion projection transformation; Based on the target display area and the target display ratio, a lightweight virtual rendering engine is invoked to render the virtual content in real time, and the rendered virtual content is then pixel-level blended with the real scene to output the virtual-real fusion guide screen.

8. A scripted virtual-real fusion tour guide system based on BeiDou positioning, applied to the scripted virtual-real fusion tour guide method based on BeiDou positioning as described in any one of claims 1-7, characterized in that, The system includes: The positioning acquisition module is used to acquire user location data in real time through a positioning unit that supports the BeiDou positioning system, filter and calibrate the user location data, and acquire real-time images through the user equipment. The script triggering module is used to construct a scripted interaction framework. When the spatial deviation between the calibrated user location data and the script node position in the scripted interaction framework is not greater than a preset deviation threshold, the script node content corresponding to the script node position is triggered. The character interaction module is used to retrieve the corresponding virtual character from the scripted interaction framework according to the script node content, explain the real scene through the virtual character, and issue scripted interaction tasks in combination with the user's historical behavior data. The music generation module is used to build an AI composition model, perform scene emotion recognition on the script node content, the real scene, the historical behavior data and the scripted interactive task, generate comprehensive emotion features, input the comprehensive emotion features into the AI ​​composition model, and output scene-adapted music. The fusion rendering module is used to fuse the virtual character and the interactive interface of the scripted interactive task with the real scene based on the virtual-real fusion rendering algorithm, output the virtual-real fusion guide screen, and play the scene-adapted music synchronously.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor runs the computer program stored in the memory, the processor performs the steps of a scripted virtual-real fusion tour guide method based on BeiDou positioning as described in any one of claims 1-7.

10. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it is used to implement the steps of a scripted virtual-real fusion tour guide method based on BeiDou positioning as described in any one of claims 1-7.