Remote guided tour live broadcast system based on panoramic video

By employing panoramic video capture, intelligent interaction, and augmented reality technologies, combined with a self-evolving optimization module, the system addresses the issues of insufficient immersion and limited interactivity in remote attraction navigation systems. This enables highly immersive multimodal interactive experiences and personalized navigation, while also improving the accuracy of content recommendations and interaction strategies.

CN122120482APending Publication Date: 2026-05-29NANJING NICEBRIDGE INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING NICEBRIDGE INFORMATION TECH CO LTD
Filing Date
2026-04-28
Publication Date
2026-05-29

Smart Images

  • Figure CN122120482A_ABST
    Figure CN122120482A_ABST
Patent Text Reader

Abstract

The application discloses a remote guide live broadcast system based on panoramic video, belonging to the technical field of live broadcast and guide, comprising: a panoramic acquisition module generates a live broadcast stream with a guide voice through video splicing, audio processing and adaptive coding; a tourist terminal realizes matching of explanation content and preloading of data by combining positioning and viewpoint prediction; an intelligent interaction module triggers companion chatting when there is no guide audio, and optimizes experience by fusing emotional computing and multi-modal interaction; a voice recognition module completes voice instruction processing and multi-lingual interpretation; an augmented reality module realizes accurate superposition of real scene and virtual content through pose solution, virtual-real fusion rendering, light field and occlusion processing; a self-evolution optimization module collects tourist behavior feedback, and combines deep reinforcement learning and federated learning to jointly optimize a response strategy model and a content recommendation model, so that the system can realize self-evolution and personalized guide, and the immersion and interactivity of remote guide are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of live streaming and tour guide technology, specifically a remote tour guide live streaming system based on panoramic video. Background Technology

[0002] Current remote tourist attraction guides mostly rely on ordinary video live streaming or static text and images, which suffer from insufficient immersion and limited interactive experiences. Ordinary video live streaming has a fixed perspective, failing to achieve a panoramic immersive tour, and the matching degree between the guide content and the tourist's location and needs is low, easily leading to content disconnect. Furthermore, the interaction of existing guide systems is mostly one-way information push, lacking real-time intelligent chat capabilities, easily resulting in service gaps during periods without guided tours. Voice interaction only supports a single language, failing to meet the needs of tourists from different regions, and the integration of multimodal interaction is also lacking. In addition, the virtual content overlay of traditional guide systems suffers from inconsistencies in lighting and space, resulting in poor integration of virtual and real scenes. Network adaptability is weak, easily causing live streaming interruptions due to bandwidth fluctuations, and data loading lacks predictability, leading to significant latency issues. Optimization and upgrades of existing systems largely rely on manual data processing, failing to achieve self-evolution based on tourist behavior feedback. The accuracy of content recommendation and interaction strategies is difficult to continuously improve, and there are also risks of privacy leaks during user behavior data collection. Overall, these systems fail to meet tourists' needs for immersive, personalized, and intelligent remote guides. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention proposes a remote guided live streaming system based on panoramic video. The panoramic acquisition module generates a live stream with tour guide narration through video stitching, audio processing, and adaptive encoding. The visitor's end combines positioning and viewpoint prediction to match narration content and preload data. The intelligent interaction module triggers chat when tour guide audio is unavailable, integrating affective computing and multimodal interaction to optimize the experience. The speech recognition module processes voice commands and translates between multiple languages. The augmented reality module achieves accurate overlay of real and virtual content through pose calculation, virtual-real fusion rendering, and light field and occlusion handling. The self-evolutionary optimization module collects visitor behavior feedback and, combined with deep reinforcement learning and federated learning, jointly optimizes the response strategy model and content recommendation model, enabling the system to achieve self-evolution and personalized tours, enhancing the immersion and interactivity of remote guided tours.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] The remote guided live streaming system based on panoramic video includes: a panoramic acquisition module, a visitor terminal, an intelligent interaction module, a voice recognition module, an augmented reality module, and a self-evolving optimization module;

[0006] The panoramic acquisition module collects video and audio data from each panoramic camera array, generates a panoramic video stream and a tour guide audio stream, and then merges and outputs a live stream with tour guide voiceover.

[0007] The tourist terminal receives the live stream and obtains location information through its built-in positioning module. Based on the location information, it matches and retrieves the explanation content of the current attraction from the cloud content library and plays it in the live stream.

[0008] During the periods when there is no tour guide audio in the live stream, the intelligent interaction module activates the interaction engine, retrieves preset content from the cloud content library to generate chat voice, and outputs it to the tourist terminal in sync with the live stream. During the interaction, the speech recognition module collects the tourist's voice commands, parses them, retrieves the corresponding response content from the pre-stored cloud content library, and generates chat voice after speech synthesis.

[0009] Based on the position change identified by the positioning module, the augmented reality module retrieves the corresponding augmented reality content from the cloud content library and overlays it onto the live stream using virtual reality technology.

[0010] The self-evolutionary optimization module collects real-time data on tourists' interactive behaviors and feedback, and jointly optimizes the response strategy model of the interaction engine and the content recommendation model of the cloud content library.

[0011] Specifically, the panoramic acquisition module includes multiple array-distributed panoramic camera units and a video stitching server connected to the multiple panoramic camera units. The video stitching server extracts feature points from the raw video data output by each panoramic camera unit based on the scale-invariant feature transform algorithm, and eliminates mismatches of the extracted feature points through the random sampling consistency algorithm. Then, it calculates the homography matrix of the images acquired by adjacent panoramic camera units according to the spatial position relationship of the feature points, and uses the homography matrix to project multiple raw video data onto a preset equidistant cylindrical projection model to complete the spatial alignment and real-time fusion of the images and generate a panoramic video stream.

[0012] The panoramic acquisition module also includes an audio processing unit connected to the video stitching server; the audio processing unit uses a beamforming algorithm to perform sound source localization and enhancement processing on the multi-channel audio data collected by the microphone array built into each panoramic camera unit, extracts the tour guide audio stream, and after timestamping and synchronizing the tour guide audio stream with the panoramic video stream, outputs a live stream with tour guide voice.

[0013] Specifically, the visitor terminal includes a head-mounted display device and a mobile computing unit connected to the head-mounted display device. The mobile computing unit has a built-in positioning module, which obtains the latitude and longitude coordinates of the visitor's location based on a global navigation satellite system receiver, and integrates the three-axis acceleration and angular velocity data collected by the inertial measurement unit inside the head-mounted display device. The module then performs real-time correction and smoothing of the visitor's location using an extended Kalman filter algorithm to generate visitor location information. Based on the visitor location information, the mobile computing unit uses a hash index-based data query method to match and retrieve the current attraction's explanation content from the cloud content library, decodes the explanation content, and overlays it onto the corresponding timeline of the live stream for playback.

[0014] Specifically, the intelligent interaction module includes an interaction status monitoring unit and a dialogue management unit;

[0015] The interactive status monitoring unit analyzes the energy value of the audio track of the live stream in real time, and determines whether there is tour guide audio data in the current time period through short-time energy analysis and zero-crossing rate detection. When no tour guide audio data is detected within a continuous preset time threshold, an interactive trigger signal is generated.

[0016] After receiving the interaction trigger signal, the dialogue management unit starts the interaction engine. Based on the tourist's current location information, it selects preset interactive content that matches the tour scenario from the cloud content library. It then uses a time-domain interpolation algorithm to perform multi-track audio mixing processing on the chat voice generated by the preset interactive content and the live stream. After the volume ratio of the chat voice and the background ambient sound meets the preset loudness standard, it outputs the audio to the tourist's end synchronously.

[0017] Specifically, the speech recognition module includes a front-end speech processing unit and a back-end semantic understanding unit;

[0018] The front-end voice processing unit collects tourists' voice commands through the microphone array at the tourist end, applies the least mean square adaptive filtering algorithm to suppress the collected environmental noise, and then extracts the Mel frequency cepstral coefficients as acoustic features.

[0019] The backend semantic understanding unit inputs the extracted Mel frequency cepstral coefficients into the end-to-end speech recognition model trained based on connection time-series classification, decodes and outputs the corresponding text instructions, and uses a named entity recognition algorithm based on bidirectional long short-term memory network and conditional random field to extract key information including the name of the scenic spot and the operation intention from the text instructions. Based on the key information, a query vector is constructed, and matching response content is retrieved from the cloud content library.

[0020] Specifically, the augmented reality module includes a pose calculation unit and a virtual-real fusion rendering unit;

[0021] The pose calculation unit calculates the six-degree-of-freedom pose parameters of the tourist in three-dimensional space in real time using a visual simultaneous localization and mapping algorithm, based on the tourist's position information output by the positioning module and the posture data collected by the inertial measurement unit inside the head-mounted display device.

[0022] The virtual-real fusion rendering unit, based on six degrees of freedom pose parameters, calls the corresponding augmented reality 3D model resources from the cloud content library, uses a depth testing algorithm based on occlusion detection to calculate the visibility of the augmented reality 3D model from the current viewpoint, and registers and superimposes the augmented reality 3D model onto the corresponding pixel area of ​​the live stream through perspective projection transformation, generating an augmented reality output stream with spatial consistency.

[0023] Specifically, the augmented reality module also includes a dynamic occlusion processing unit;

[0024] Before overlaying augmented reality content onto the live stream, the dynamic occlusion processing unit performs monocular depth estimation on the current frame image of the live stream through a depth estimation network to generate a depth map with the same resolution as the image.

[0025] The depth value of each pixel of the augmented reality 3D model is calculated based on the depth map at the current viewpoint and compared with the depth value of the real scene image. When the depth value of the augmented reality 3D model is greater than the depth value of the real scene image, it is determined that the augmented reality 3D model is occluded by a real object. The area occluded by the real object is made transparent or only the outline of the occluded part is rendered.

[0026] Specifically, the self-evolutionary optimization module includes a data acquisition unit and an offline training unit;

[0027] The data acquisition unit records tourists' interactive behavior data in real time during the interaction process, and after cleaning and normalizing the interactive behavior data, it constructs structured sample data containing context state, tourist actions and feedback scores; the interactive behavior data includes voice request text, click browsing behavior, dwell time and skip behavior data for various types of response content.

[0028] The offline training unit uses a deep deterministic policy gradient algorithm to train the response strategy model inside the interaction engine based on the structured sample data accumulated within a preset period, with the optimization objective of maximizing the cumulative satisfaction score, and updates the network parameters of the response strategy model.

[0029] Specifically, the self-evolutionary optimization module also includes an online update unit;

[0030] The online update unit loads the response strategy model parameters updated by the offline training unit. During real-time interaction, it calculates the selection probability of each candidate response strategy based on the current real-time interaction state of the tourist through forward propagation, and uses an exploration strategy based on Boltzmann distribution for action sampling. At the same time, the online update unit collects feedback data of the current interaction process, uses an online learning algorithm based on a priority experience replay mechanism to fine-tune the value network of the response strategy model in real time, and uses the fine-tuned value network to correct the output of the strategy network of the response strategy model in real time.

[0031] Specifically, the self-evolutionary optimization module also jointly optimizes the content recommendation model of the cloud content library; the content recommendation model adopts an architecture based on deep factorization machine, which includes an embedding layer, a factorization machine layer, a deep neural network layer and a logistic regression layer.

[0032] The embedding layer maps tourists' discrete and continuous features into embedding vectors, where discrete features include historical visit sequence and interaction preference type, and continuous features include current location and duration of stay.

[0033] The factorization machine layer performs second-order feature cross modeling on the embedded vectors and outputs cross feature representations.

[0034] The deep neural network layer takes the result of embedding vector concatenation as input, extracts nonlinear features through a multi-layer fully connected network, and finally concatenates the outputs of the factorization machine layer and the deep neural network layer. Then, it calculates the recommendation score of each candidate content through a logistic regression layer, sorts the candidate content according to the recommendation score, and responds to the visitor's query request.

[0035] Compared with the prior art, the beneficial effects of the present invention are:

[0036] 1. This invention proposes a remote guided live streaming system based on panoramic video, and optimizes and improves its architecture, operation steps and processes. The system has the advantages of simple process, low investment and operating costs and low production and working costs.

[0037] 2. This invention proposes a remote guided tour live streaming system based on panoramic video. This system creates a highly immersive remote guided tour experience through the integrated application of technologies such as panoramic acquisition, augmented reality, and precise positioning. The panoramic acquisition module achieves precise fusion and adaptive encoding and distribution of multiple video and audio streams, adapting to different network conditions to ensure smooth live streaming. The augmented reality module combines pose calculation, light field consistency rendering, and dynamic occlusion processing to ensure that the virtual content is highly consistent with the space and lighting of the real-world live stream. Combined with viewpoint prediction and preloading on the visitor's end, data loading latency is significantly reduced. At the same time, the positioning module accurately matches the explanation content of the attractions after algorithm correction, significantly improving the real-world experience and content relevance of the remote guided tour.

[0038] 3. This invention proposes a remote guided tour live streaming system based on panoramic video. The system relies on intelligent interaction, speech recognition, and self-evolutionary optimization to build a highly intelligent and personalized interactive system, and also achieves continuous iteration of system capabilities. The intelligent interaction module can automatically trigger chat and combine emotion computing and multimodal interaction to meet the needs of tourists. The speech recognition module supports multilingual recognition and translation, breaking down language barriers. The self-evolutionary optimization module collects tourist behavior feedback and combines deep reinforcement learning and federated learning to optimize the response strategy model and content recommendation model. This not only achieves personalized guided tour services, but also completes the joint optimization of system models while protecting user data privacy, allowing the system's interactive capabilities and content recommendation accuracy to continuously improve with use. Attached Figure Description

[0039] Figure 1 This is an architecture diagram of the remote guided live streaming system based on panoramic video of the present invention;

[0040] Figure 2 This is a flowchart illustrating the principle of the remote guided live streaming system based on panoramic video according to the present invention. Detailed Implementation

[0041] This embodiment uses a remote guided tour of a 5A-level natural scenic area's mountain landscape as a typical application scenario. This scenic area encompasses various tourist areas such as mountain trails, viewing platforms, and historical sites. The terrain is complex and the tour routes are scattered. Traditional online guided tour methods suffer from problems such as a single perspective, weak interactivity, and poor user experience. This system addresses these pain points by providing remote tourists with a 720° panoramic immersive guided tour live streaming service through the collaborative operation of multiple modules. Tourists can experience the scenery of the scenic area as if they were there through head-mounted display devices, while also receiving personalized explanations, intelligent interactive companionship, and an augmented reality experience that blends the virtual and real worlds. The following will elaborate on the specific implementation process, technical details, and working principles of each module of the system in this typical scenario.

[0042] Please see Figure 1 and Figure 2 The present invention provides an embodiment of a remote guided live streaming system based on panoramic video, comprising: a panoramic acquisition module, a visitor terminal, an intelligent interaction module, a voice recognition module, an augmented reality module, and a self-evolving optimization module;

[0043] The panoramic acquisition module collects video and audio data from each panoramic camera array, generates a panoramic video stream and a tour guide audio stream, and then merges and outputs a live stream with tour guide voiceover.

[0044] In the mountain scenic area scenario of this embodiment, the panoramic acquisition module deploys multiple array-distributed panoramic camera units at key locations such as viewing platforms, trail nodes, and core areas of historical sites within the scenic area. These units are paired with corresponding audio processing units and adaptive bitrate encoding submodules. Through multi-algorithm collaborative processing, a multi-bitrate live stream with tour guide audio is generated, adaptable to different network conditions, providing tourists with a stable and clear panoramic live stream data source.

[0045] The panoramic acquisition module described in this embodiment includes multiple array-distributed panoramic camera units and a video stitching server connected to the multiple panoramic camera units. In the mountainous scenic area of ​​this embodiment, the panoramic camera units at each key point adopt an 8-lens panoramic camera array, with each camera having a field of view of 120° and an overlap area of ​​at least 30° between adjacent cameras, ensuring that the acquired video data can completely cover a 720° panoramic view. Furthermore, each panoramic camera unit is encapsulated in a waterproof and shockproof protective shell, adapting to the complex outdoor environment of the mountainous scenic area. At the same time, all panoramic camera units are wired to the video stitching server through a dedicated industrial Ethernet connection for the scenic area, ensuring the stability of data transmission.

[0046] A1: The video stitching server extracts feature points from the raw video data output by each panoramic camera unit based on the scale-invariant feature transform algorithm. After eliminating mismatches of the extracted feature points through the random sampling consistency algorithm, it calculates the homography matrix of the images acquired by adjacent panoramic camera units based on the spatial position relationship of the feature points. The homography matrix is ​​then used to project multiple raw video data onto a preset equidistant cylindrical projection model to complete the spatial alignment and real-time fusion of the images, generating a panoramic video stream. The scale-invariant feature transform algorithm and the random sampling consistency algorithm are existing technologies in this field and are not inventive solutions of this application, so they will not be described in detail here.

[0047] Furthermore, the feature point extraction and matching process includes: First, a scale-invariant feature transform algorithm is used to extract feature points from the original images acquired by adjacent panoramic camera units. In terms of the scale-invariant feature transform algorithm parameters, the number of Gaussian difference pyramid layers in the scale space is set to 6, with 3 layers per group. The initial standard deviation of the Gaussian kernel is set to 1.6, the contrast threshold for feature point detection is set to 0.04, and the edge threshold is set to 10, to ensure that the extracted feature points possess scale invariance and rotation invariance. For each image acquired by a camera, feature points are extracted according to the above parameters. The number of feature points extracted from each image is controlled between 2000 and 3000, and the feature points are retained. The pixel coordinates and corresponding scale and orientation information are used. After feature point extraction, the Euclidean distance between the feature point descriptors to be matched is calculated using the Euclidean distance calculation formula. The distance threshold is set to 0.6, meaning that when the Euclidean distance between two feature point descriptors is less than this distance threshold, and the ratio of the nearest neighbor distance to the second nearest neighbor distance is less than 0.6, it is determined to be a valid matching pair. The initial matching results are filtered, and a random sampling consensus algorithm is used to remove erroneous matching pairs. The number of iterations of the random sampling consensus algorithm is set to 2000, the inlier threshold is set to 2 pixels, and the confidence level is set to 0.99. The number of valid matching pairs retained is not less than 70% of the total initial matching pairs to ensure the accuracy of feature point matching.

[0048] Furthermore, the homography matrix is ​​a 3×3 homogeneous matrix. Its calculation is based on the perspective transformation model, using the pixel coordinates and spatial three-dimensional coordinates of the feature points to establish a system of linear equations, including: first, determining the spatial three-dimensional coordinates of the feature points, with the geometric center of the panoramic camera array as the origin of the world coordinate system, the X-axis pointing due east of the scenic area, the Y-axis pointing due north, and the Z-axis perpendicular to the ground upwards; obtaining the intrinsic and extrinsic parameter matrices of each camera through camera calibration. The intrinsic parameter matrix includes parameters such as focal length and principal point coordinates, where the focal length is set to 12 mm, and the principal point coordinates are set to the center position of the image resolution. For example, when the image resolution is 7680×4320, the principal point coordinates are (3840, 2160); the extrinsic parameter matrix includes a rotation matrix and a translation vector, with the three rotation angles of the rotation matrix set as follows: =0°、 =0°、 =15°, the translation vector is set to (0.5 m, 0, 0); the pixel coordinates of the feature points are converted into three-dimensional coordinates in the world coordinate system through the camera intrinsic and extrinsic parameter matrices; then, using the three-dimensional coordinates and pixel coordinates of the feature points of two adjacent images as input, a system of linear equations for the homography matrix is ​​constructed; the least squares method is used to solve the optimal solution of the system of equations; during the solution process, the regularization coefficient is set to 0.001 to avoid overfitting; after obtaining the initial homography matrix, it is optimized again using the random sampling consensus algorithm, with the number of iterations set to 1500 times and the error threshold set to 1 pixel; finally, a homography matrix with an accuracy error of less than 0.5 pixels is obtained, which can accurately describe the perspective transformation relationship between adjacent camera images.

[0049] Furthermore, the parameters of the equidistant cylindrical projection model are set as follows: the radius of the projection sphere is consistent with the acquisition radius of the panoramic camera array, which is set to 5 meters; the resolution of the projected image is set to 16384×8192; the horizontal viewing angle coverage range is 0° to 360°; the vertical viewing angle coverage range is -90° to 90°; and the actual angle corresponding to the pixel pitch is 0.02197° / pixel.

[0050] Furthermore, during the projection process, each pixel of each frame of the original video image undergoes coordinate transformation using a homography matrix to calculate its corresponding coordinates in the equidistant cylindrical projection model. Specifically, the homogeneous coordinates of the original image pixels are multiplied by the homography matrix to obtain the projected homogeneous coordinates. These homogeneous coordinates are then normalized to convert them into two-dimensional pixel coordinates in the equidistant cylindrical projection model. For pixel gaps that appear after projection, a bilinear interpolation algorithm is used to fill them. During interpolation, the four nearest valid pixels around the target pixel are selected, and the weighting coefficients are determined based on the inverse square of the pixel spacing to ensure a natural transition of the filled pixel values.

[0051] Furthermore, in the spatial alignment stage, the geometric center of the panoramic camera array is used as a reference to perform positional calibration on all projected images, with the pixel deviation controlled within 1 pixel. By calculating the pixel difference in the overlapping area of ​​adjacent projected images, the translation parameters of the homography matrix are adjusted to make the mean square error of pixels in the overlapping area less than 10, thus completing the spatial alignment of all images and ensuring that there is no obvious misalignment between adjacent images at the stitching point.

[0052] Furthermore, after image spatial alignment, the multi-projected images are fused in real time. The fusion process employs a weighted fusion algorithm. For overlapping regions, a weight coefficient is set to change linearly with pixel position. The width of the overlapping region is set to 100 pixels. At the beginning of the overlapping region, the weight coefficient of the left image is 1, and the weight coefficient of the right image is 0; at the end of the overlapping region, the weight coefficient of the left image is 0, and the weight coefficient of the right image is 1. The weight coefficients at intermediate positions are calculated using linear interpolation to ensure a smooth transition of pixel values ​​in the overlapping region without obvious seams. For non-overlapping regions… The projected pixel values ​​are directly retained, and the frame rate of the fusion processing is kept consistent with the original video capture frame rate, set to 30 frames per second. The fusion processing time for each frame is controlled within 30 milliseconds to meet real-time requirements. Finally, the fused image is encapsulated according to the video encoding standard to generate a panoramic video stream. The video encoding adopts the H.265 encoding standard, and the bitrate control parameters are set as follows: initial bitrate is 80Mbps, minimum bitrate is 40Mbps, maximum bitrate is 100Mbps, and the GOP length is set to 60 frames to ensure that the image quality and transmission efficiency of the panoramic video stream are balanced.

[0053] In this embodiment, the video stitching server adopts a high-performance industrial-grade server, equipped with a multi-core processor and a high-performance graphics processing unit to meet the real-time processing needs of massive video data. The specific processing procedure is as follows: First, the scale-invariant feature transform algorithm is used to extract feature points from each original video frame. The scale-invariant feature transform algorithm can maintain the invariance of feature points under image scaling, rotation, and illumination changes, adapting to the scene of illumination changing with time and terrain in mountainous scenic areas. The extracted feature points contain key information such as position, scale, and orientation. The number of feature points extracted for each frame is controlled between 2000 and 3000, which ensures the accuracy of feature matching while avoiding processing delays caused by too many feature points. Second, for mismatched feature points caused by mountain occlusion and changes in light and shadow, a random sampling consistency algorithm is used to eliminate them. This algorithm eliminates them through multiple... A sample set is generated through random sampling, and a feature matching model is constructed. Outliers that do not meet the model requirements are removed, and the feature point matching accuracy after eliminating mismatches is no less than 98%. Next, the video stitching server calculates the homography matrix of the images acquired by adjacent panoramic camera units based on the spatial position relationship of the feature points after eliminating mismatches. The homography matrix can describe the projection transformation relationship between two images. In this embodiment, the homography matrix is ​​solved by the least squares method to ensure the accuracy of the matrix solution. Finally, the multiple original video data are projected onto a preset equidistant cylindrical projection model through the homography matrix. This equidistant cylindrical projection model can convert the spherical panoramic image into a cylindrical planar image, which is adapted to the display device of the visitor. During the projection process, the overlapping areas of the images are feathered and blended to eliminate the seams of the image stitching, and finally a seamless 720° panoramic video stream with a resolution of 8K and a frame rate of 30 frames per second is generated.

[0054] A2: The panoramic acquisition module also includes an audio processing unit connected to the video stitching server; the audio processing unit performs sound source localization and enhancement processing on the multi-channel audio data collected by the microphone array built into each panoramic camera unit through a beamforming algorithm, extracts the tour guide audio stream, and after timestamping and synchronizing the tour guide audio stream with the panoramic video stream, outputs a live stream with tour guide voice. The beamforming algorithm is existing technology in this field and is not an inventive solution of this application, so it will not be described in detail here.

[0055] In this embodiment, each panoramic camera unit has a built-in 6-microphone ring array with a sampling rate of 48kHz and a bit depth of 24 bits, enabling accurate acquisition of environmental sounds and tour guide voice information within the scenic area. The specific working process of the audio processing unit is as follows: First, the beamforming algorithm performs sound source localization on the multi-channel audio data. By calculating the time difference and phase difference of the audio signals acquired by different microphones, the sound source location of the tour guide's voice is determined. In mountainous scenic areas, this effectively shields the interference of environmental noise such as wind noise, birdsong, and tourist noise, with a sound source localization error of no more than 0.5 meters. Second, based on the sound source localization results... The system enhances the tour guide's audio signal, amplifying the audio signal of the target sound source and suppressing noise from non-target sound sources, achieving a signal-to-noise ratio of no less than 30dB for the processed tour guide audio. Next, the audio processing unit performs noise reduction and echo cancellation on the enhanced tour guide audio stream to ensure clarity. Finally, the processed tour guide audio stream is timestamped and encapsulated with the panoramic video stream generated by the video stitching server. The synchronization accuracy of the timestamps is controlled within 10 milliseconds to avoid audio-visual asynchrony. The encapsulation format uses TS format, adapted for live broadcast distribution, ultimately outputting a live stream with tour guide audio.

[0056] Furthermore, the panoramic acquisition module also includes an adaptive bitrate encoding submodule. In the remote tour scenario of the mountain scenic area in this embodiment, the network environments of remote tourists vary, including fiber optic, 5G, 4G, and wireless networks. Network quality parameters such as network bandwidth, packet loss rate, and round-trip latency fluctuate significantly. The adaptive bitrate encoding submodule can dynamically adjust the video encoding bitrate and resolution according to the tourist's real-time network status to ensure that the tourist's end can smoothly play the panoramic live stream. The specific process includes:

[0057] (1) The adaptive bitrate encoding submodule monitors the bandwidth, packet loss rate, and round-trip delay of the uplink network in real time, and inputs the parameters into the bitrate decision model based on reinforcement learning. The bitrate decision model uses the Q-learning algorithm to select the optimal combination of video encoding bitrate and resolution based on the current network state. In this embodiment, the network monitoring module of the adaptive bitrate encoding submodule collects the three core network quality parameters of the uplink network—bandwidth, packet loss rate, and round-trip delay—in real time at a period of 500 milliseconds. The collected parameter data is smoothed and filtered to eliminate the influence of accidental fluctuations. This serves as the input to the bitrate decision model. The bitrate decision model is built upon reinforcement learning, using the Q-learning algorithm as its core algorithm. It categorizes network conditions into four levels: excellent, good, average, and poor. A high-quality network has bandwidth above 20Mbps, a packet loss rate below 0.1%, and a round-trip latency below 50ms. A good network has bandwidth between 10-20Mbps, a packet loss rate between 0.1%-0.5%, and a round-trip latency between 50-100ms. A average network has bandwidth between 4-10Mbps, a packet loss rate between 0.5%-2%, and a round-trip latency between 100-20ms. Between 0ms and 1ms, poor network conditions include bandwidth below 4Mbps, packet loss rate above 2%, and round-trip latency above 200ms. Simultaneously, the bitrate decision model pre-sets five combinations of encoding bitrate and resolution: 8K / 30fps (80Mbps), 4K / 30fps (40Mbps), 2K / 30fps (20Mbps), 1080P / 30fps (10Mbps), and 720P / 30fps (5Mbps), corresponding to different network conditions. The Q-learning algorithm establishes network conditions through continuous trial and error and learning. A Q-value table is provided between the state and the encoding combination. The Q-value represents the cumulative reward for selecting any encoding combination in any network state. The reward function comprehensively considers the smoothness of video playback and the image quality. When there is no lag and the image quality is clear on the user's end, a positive reward is given. When there are problems such as lag or screen tearing, a negative reward is given. After the model is trained, it can quickly look up the Q-value table based on the real-time monitored network quality parameters and select the encoding bitrate and resolution combination with the largest Q-value as the optimal decision result. The Q-learning algorithm is existing technology in this field and is not an inventive solution of this application. It will not be described in detail here.

[0058] (2) The video stitching server transcodes the generated panoramic video stream in real time according to the optimal combination of encoding bitrate and resolution, outputs a multi-bitrate live stream adapted to the current network conditions, and distributes it through the Hypertext Transfer Protocol (HTTP) live streaming method. In this embodiment, the video stitching server is equipped with a hardware transcoding chip, which can realize real-time transcoding of the panoramic video stream. The transcoding delay does not exceed 200 milliseconds. According to the optimal encoding combination output by the bitrate decision model, the basic 8K panoramic video stream is transcoded. At the same time, in order to adapt to the network status changes of different tourists, the server will also generate video streams of the other 4 encoding combinations simultaneously, forming 5 multi-bitrate live streams with different bitrates. All live streams are distributed through the Hypertext Transfer Protocol (HTTP) live streaming method. The edge node server deployed in the scenic area serves as the live streaming distribution node, providing tourists with the nearest live stream retrieval service, effectively reducing the transmission delay of the live stream. In the mountain scenic area, the average delay of tourists retrieval of the live stream is controlled within 500 milliseconds, improving the viewing experience of tourists.

[0059] The tourist terminal receives the live stream and obtains location information through its built-in positioning module. Based on the location information, it matches and retrieves the explanation content of the current attraction from the cloud content library and plays it in the live stream.

[0060] The visitor terminal includes a head-mounted display device and a mobile computing unit connected to the head-mounted display device. In this embodiment, the head-mounted display device is a lightweight VR all-in-one machine equipped with two 4K resolution AMOLED displays with a refresh rate of 90Hz and a field of view of 110°, providing visitors with a clear and smooth 720° panoramic visual experience. The device also incorporates sensors such as a gyroscope, accelerometer, and inertial measurement unit to collect real-time head movement postures of visitors. The mobile computing unit uses a high-performance portable computing box equipped with an octa-core processor and an independent graphics processing unit, possessing powerful real-time data processing and graphics rendering capabilities. It achieves a wired connection to the head-mounted display device via a Type-C interface, ensuring bandwidth and stability of data transmission. Furthermore, the mobile computing unit supports multiple network connection methods such as 5G, 4G, and wireless networks, adapting to different network environments for visitors.

[0061] B1: The mobile computing unit has a built-in positioning module. This positioning module obtains the latitude and longitude coordinates of the tourist's location based on the Global Navigation Satellite System receiver, and integrates the three-axis acceleration and angular velocity data collected by the inertial measurement unit inside the head-mounted display device. The extended Kalman filter algorithm is used to perform real-time correction and smoothing of the tourist's location to generate tourist location information. The extended Kalman filter algorithm is existing technology in this field and is not an inventive solution of this application, so it will not be described in detail here.

[0062] In this embodiment, the global navigation satellite system receiver of the positioning module supports joint positioning using GPS, BeiDou, and GLONASS systems, with a positioning accuracy of no less than 10 meters, accurately acquiring the physical latitude and longitude coordinates of the tourist. The inertial measurement unit inside the head-mounted display device collects the three-axis acceleration and angular velocity data of the tourist's head at a frequency of 100Hz, reflecting the tourist's head movement state and relative position changes. Since the tourist's physical location in the remote tour of the mountain scenic area is not the actual location within the scenic area, the latitude and longitude coordinates collected by the positioning module are mainly used to match the scenic area that the tourist currently wants to visit, while the sensor data from the inertial measurement unit is used to assist in correcting the tourist's relative position in the virtual scenic area, avoiding position discrepancies. The offset; the extended Kalman filter algorithm, as the core position correction algorithm, can fuse latitude and longitude data collected by the global navigation satellite system receiver and sensor data collected by the inertial measurement unit. Specifically, it includes: first, predicting the current position state based on the position information of the previous moment and the uniform motion model in the prediction stage; then, comparing the current observation data with the predicted value in the update stage, calculating the Kalman gain, correcting the predicted value, and finally generating smooth and accurate tourist position information. The update frequency of the position information is 50Hz to ensure that it can match the scenic spot's explanation content in real time. The uniform motion model is existing technology in this field and is not an inventive solution of this application, so it will not be described in detail here.

[0063] Furthermore, this embodiment employs a uniform motion model based on an inertial measurement unit (IMU). This model is designed for virtual positioning scenarios in scenic area remote navigation and head-mounted displays, equating tourist head rotation, perspective switching, and virtual location roaming as continuous, smooth spatial motion. This eliminates positioning deviations caused by satellite positioning drift, attitude jumps, and terrain occlusion, ensuring real-time matching between the narration content and the tourist's location. In the prediction phase, based on the filtered and optimized position, velocity, attitude, and sensor zero bias from the previous moment, state prediction is performed according to the uniform motion law. First, based on the previous position and velocity, combined with the three-axis acceleration collected by the IMU, the predicted position is updated. Then, based on the acceleration data and sensor zero bias, the predicted velocity is updated. Simultaneously, based on the angular velocity data, the roll angle, pitch angle, and yaw angle are updated through quaternion integration to obtain the predicted attitude. The zero biases of the accelerometer and gyroscope remain stable in a short time using a random walk method, completing the predicted zero bias update. Finally, based on the system motion characteristics and process noise, the predicted state uncertainty is calculated, yielding the prior state and prior covariance. During the update phase, the predicted state is compared with the actual observation data and corrections are performed. First, the virtual scenic area coordinates output by the satellite positioning module and the head attitude directly output by the inertial measurement unit (IMU) are acquired as observation data. Then, the observation data is compared with the predicted state to obtain the state prediction deviation. Subsequently, the Kalman gain is calculated based on the prediction uncertainty and the observation noise intensity, serving as the optimal weighting coefficient. The Kalman gain is used to weight and correct the predicted position, velocity, attitude, and zero bias to obtain the smooth and accurate optimal state for the current moment. Finally, the state uncertainty is updated to provide an accurate benchmark for the prediction of the next moment. The input to the uniform motion model is the satellite positioning coordinates, the three-axis acceleration, and the three-axis angular velocity output by the IMU. The IMU samples at 100 Hz, and the satellite positioning is updated at 1 Hz. The model outputs smoothed position and attitude information at a uniform frequency of 50 Hz. The output results are directly used for matching scenic spot explanation content, overlaying augmented reality content, and preloading viewpoint predictions, effectively suppressing positioning jumps and attitude jitter, ensuring the synchronization and immersion of remote guided tours.

[0064] B2: The mobile computing unit uses a hash index-based data query method to match and retrieve the explanation content of the current attraction from the cloud content library based on the tourist's location information, and then decodes the explanation content and superimposes it onto the corresponding timeline of the live stream for playback.

[0065] Furthermore, the cloud-based content library is pre-built using a hybrid storage architecture combining a distributed file system and a relational database. The distributed file system stores multimedia files such as panoramic videos, audio guides, and augmented reality 3D models, while the relational database stores metadata corresponding to the multimedia files, including content identifiers, geographic location tags, semantic tags, content types, and access frequencies. The cloud-based content library assigns unique geographic location tags to each attraction within the scenic area, along with corresponding guide content, panoramic video resources, and augmented reality resources. Each geographic location tag corresponds one-to-one with the actual latitude and longitude coordinates of the attraction. The cloud-based content library provides a proximity search interface based on geographic location and a similarity search interface based on semantic vectors. The proximity search interface uses an R-tree-based index structure to achieve fast spatial retrieval of attraction content, while the semantic similarity search interface uses a pre-trained model to convert the query content into semantic vectors to retrieve similar content. The mobile computing unit performs hash operations on the latitude and longitude coordinates of the tourist's location information to generate a unique hash index value. This hash index value is then used to quickly query the geographic location index table in the cloud-based content library to match the attraction guide content corresponding to the tourist's current location, achieving low-latency, high-accuracy content retrieval and synchronized playback.

[0066] In this embodiment, the cloud-based content library configures exclusive explanation content for each scenic spot in the mountain area, including audio guides, text descriptions, and image materials. All explanation content is tagged with corresponding geographic locations, each corresponding to the latitude and longitude coordinates of the scenic spot. The mobile computing unit employs a hash-index-based data query method. First, it performs a hash operation on the latitude and longitude coordinates of the tourist's location information to generate a unique hash index value. Then, it uses this hash index value to quickly query the geographic location index table in the cloud-based content library, matching the explanation content of the scenic spot closest to the tourist's current virtual location within 10 milliseconds. Compared to traditional methods… The traditional linear query method improves query efficiency by over 80%. The retrieved explanation content is stored as an audio stream in MP3 format. The mobile computing unit decodes it in real time and overlays the audio explanation stream onto the corresponding timeline of the live stream. It is then mixed with the tour guide's voice in the live stream. During the mixing process, the volume ratio of the explanation content and the tour guide's voice is adjusted according to the tourist's settings. The default ratio is 1:1, and tourists can also manually adjust it. Ultimately, the explanation content and the live stream are played synchronously. While watching the panoramic live stream, tourists can hear the corresponding explanations of the attractions, enhancing the knowledge and interest of the guided tour.

[0067] Furthermore, the visitor terminal also includes a viewpoint prediction and preloading submodule;

[0068] (1) The viewpoint prediction and preloading submodule is based on the tourist's historical head movement trajectory. By training a viewpoint prediction model based on a long short-term memory network, it takes the head posture sequence within a preset time period as input and outputs the predicted head posture value within the preset time period. In this embodiment, the head-mounted display device collects the tourist's head posture data at a frequency of 100Hz, including yaw angle, pitch angle, and roll angle, which constitute the tourist's head movement trajectory. The viewpoint prediction model is constructed based on a long short-term memory network. The long short-term memory network can effectively capture long-term dependencies in time-series data and is suitable for processing time-series data such as head posture sequences. The training of the viewpoint prediction model The dataset consists of massive amounts of head movement trajectory data from tourists viewing panoramic views of scenic spots, covering tourists of different ages and with different travel preferences, ensuring the model's generalization ability. The model's input is a sequence of head postures over the past 2 seconds, totaling 200 data points, and the output is a predicted head posture value over 1 second, totaling 100 data points. After training, the model achieves a prediction accuracy of no less than 85%, accurately capturing the head movement patterns of tourists. For example, when tourists are viewing the panoramic view from the observation deck, they are likely to first turn their heads to the left to view the mountain scenery, and then turn their heads to the right to view the historical sites. The model can accurately predict the tourist's next viewing angle based on this historical movement trajectory.

[0069] Furthermore, the viewpoint prediction model is a temporal prediction network specifically built for panoramic guided head-mounted display devices. It adopts an encoder-decoder structure, entirely composed of stacked long short-term memory (LSM) network layers. This structure is used to capture the long-term dependencies and viewpoint change patterns of continuous head movements. The encoder extracts the temporal features of historical head postures and consists of two LSM layers connected sequentially. Each layer has a fixed number of hidden units, enabling segment-by-segment feature extraction and temporal information compression of the input historical posture sequence. The decoder generates the posture prediction results and also consists of two LSM layers. It receives the compressed features from the encoder and outputs predicted posture data time-by-time. The input is a sequence of head postures collected by the head-mounted display device over a past period. The input data includes yaw, pitch, and roll angles from multiple consecutive frames, with each frame acquired at a fixed sampling frequency. The model's output is a predicted sequence of head postures over a period of time, also including predicted values ​​for yaw, pitch, and roll angles.

[0070] Furthermore, the viewpoint prediction model training process is completed using massive amounts of tourist head movement trajectory data collected in panoramic tour scenarios. During training, historical posture sequences are used as model inputs, and the corresponding subsequent real postures are used as supervision labels. An adaptive moment estimation optimizer is used to update network parameters. A fixed initial learning rate is set and gradually decays during training. At the same time, reasonable training batches and maximum iterations are set. Training is stopped when the prediction error of the model on the validation set no longer decreases after several consecutive iterations, and the optimal model parameters are saved.

[0071] In actual operation, the head-mounted display device collects tourists' head posture data at a frequency of 100 Hz, including yaw angle, pitch angle, and roll angle, forming a continuous head movement trajectory. The viewpoint prediction model takes 200 frames of head posture data from the past two seconds as input and predicts 100 frames of head posture data from the next second. The output predicted posture is used to determine the field of view area that the tourist will focus on, and then initiates the preloading of augmented reality 3D models and high-resolution texture data of the corresponding area to the cloud content library. The preloaded content is cached locally on the tourist's device, achieving smooth viewpoint switching and low data loading latency, thus improving the immersiveness and smoothness of the panoramic tour.

[0072] (2) Based on the head posture prediction value, determine the field of view area that the tourist may focus on next, request a preloading instruction from the cloud content library for the augmented reality model and high-resolution texture data corresponding to the field of view area, and cache the preloaded data in the tourist's local memory; in this embodiment, based on the yaw angle, pitch angle, and roll angle of the head posture prediction value, calculate the field of view area that the tourist may focus on within 1 second. The field of view area is centered on the tourist's virtual perspective, with a field of view angle of 60°, covering the core area that the tourist is likely to view; subsequently, the mobile computing unit sends a preloading instruction to the cloud content library. The preloading instruction contains the scenic area location and perspective information corresponding to the field of view area. The cloud content library quickly retrieves the corresponding augmented reality model according to the instruction, such as the scenic area location and perspective information. The system utilizes 3D models of ancient buildings, virtual labeled models of mountain landscapes, and high-resolution texture data from panoramic videos of the area. The augmented reality model employs the lightweight glTF format, while the texture data uses a compressed texture format, effectively reducing the data volume. After the retrieved data is transmitted to the visitor's device via the network, it is cached in the local solid-state drive of the mobile computing unit. The local storage has a cache capacity of 128GB, which can store a large amount of pre-loaded data. When the visitor's head rotates into the pre-loaded field of view, the mobile computing unit directly reads the data from the local storage for rendering, without needing to fetch it from the cloud in real time. This effectively reduces data transmission latency and network bandwidth usage, making the visitor's perspective switching process smooth and lag-free, thus enhancing the immersive viewing experience.

[0073] Furthermore, the cloud-based content library adopts a hybrid storage architecture combining a distributed file system and a relational database. Multimedia files are stored in the distributed file system, while corresponding metadata, including content identifiers, geographic location tags, semantic tags, content types, and access frequencies, is stored in the relational database. In this embodiment, the distributed file system uses HDFS, deployed on a server cluster consisting of 10 high-performance servers, with a total storage capacity of 100TB. This capacity is capable of storing large multimedia files such as 8K panoramic video data of mountain scenic areas, high-definition images, augmented reality models, and audio narration files. The distributed file system features high fault tolerance and high scalability. Through a data replication mechanism, each multimedia file is stored in three copies, distributed across different server nodes. Even if any server node fails, normal data retrieval is guaranteed, with data availability not less than 99.99%. The relational database uses MySQL. Deployed on a separate database server, it stores the metadata corresponding to multimedia files. Each piece of metadata contains a unique content identifier, corresponding one-to-one with the multimedia file in the distributed file system. It also includes a geographic location tag, accurate to the latitude and longitude coordinates of each scenic spot. Semantic tags cover keywords such as scenic spot name, type, and features. Content types are divided into panoramic videos, audio explanations, augmented reality models, text materials, image materials, and interactive content. Access frequency records the number of times each content is retrieved, providing data support for the content recommendation model. The metadata storage adopts a structured table structure, supporting efficient CRUD operations. The query time for a single piece of metadata is no more than 1 millisecond.

[0074] Furthermore, the cloud-based content library provides a location-based proximity search interface and a semantic vector-based similarity search interface. The proximity search interface uses an R-tree-based index structure to quickly retrieve content about nearby attractions based on the tourist's location information. The semantic similarity search interface uses a pre-trained model based on a bidirectional encoder representation to convert the tourist's text query into a semantic vector, and then retrieves semantically relevant content through an approximate nearest neighbor search algorithm.In this embodiment, the nearest search interface is the most frequently used query interface in the system. It primarily provides proximity query services for attraction content to the location module and intelligent interaction module on the tourist side. This interface uses an R-tree-based index structure, dividing the geographical space of the scenic area into multiple rectangular regions. Each rectangular region corresponds to a sub-region of the scenic area, storing the metadata of the attractions within that region. The index nodes of the R-tree record the latitude and longitude of the rectangular region's boundaries. When a tourist's location information query request is received, the system first determines the rectangular region where the tourist is located based on their latitude and longitude coordinates, and then performs a precise query within that region to quickly retrieve the nearest surrounding attractions to the tourist's current location. The query response time is no more than 50 milliseconds, which can meet the real-time query requirements of the system. The semantic similarity search interface mainly provides semantically related content query services for the speech recognition module and the intelligent interaction module. When tourists make specific query requests via voice, such as "query information on ancient temples in the scenic area" or "recommend the best location for mountain viewing", this interface will come into play. The specific working process is as follows: First, a pre-trained model based on bidirectional encoder representation is used. This model has been trained on massive amounts of tourism-related text data and can accurately capture the semantic information of the text, converting the tourist's text query into a 768-dimensional semantic vector; then, through an approximate nearest neighbor search algorithm, in the cloud... The system searches for content with the highest similarity to a given semantic vector in the semantic vector library of the content repository. This semantic vector library is a dedicated vector index pre-built during the construction phase of the cloud-based content repository. It consists of a standardized set of vectors generated from all textual content within the cloud-based content repository, including explanatory texts, interactive scripts, attraction introductions, and guide instructions, after unified semantic encoding. Each piece of text corresponds to a unique semantic vector, forming a one-to-one binding relationship with the original content, geographic location tag, attraction number, and content type. The semantic vector library is maintained synchronously with the addition, updating, and deletion of content in the cloud-based content repository to ensure that the vectors and text content remain consistent. The approximate nearest neighbor search algorithm employs locality-sensitive hashing. The algorithm is implemented to quickly find similar vectors in a massive amount of semantic vectors. The similarity calculation adopts the cosine similarity algorithm, and the similarity threshold is set to 0.8. When the cosine similarity is higher than this similarity threshold, it is determined to be semantically related content. Finally, the retrieved semantically related content metadata is returned to the request module, which retrieves the corresponding multimedia file based on the metadata. The response time of semantic similarity search is no more than 100 milliseconds, which can accurately meet the personalized query needs of tourists. Among them, the approximate nearest neighbor search algorithm, the locality sensitive hash algorithm, and the cosine similarity algorithm are existing technologies in this field and are not the inventive solutions of this application, and will not be described in detail here.

[0075] During the periods when there is no tour guide audio in the live stream, the intelligent interaction module activates the interaction engine, retrieves preset content from the cloud content library to generate chat voice, and outputs it to the tourist terminal in sync with the live stream.

[0076] The intelligent interaction module includes an interaction status monitoring unit and a dialogue management unit;

[0077] C1: The interactive state monitoring unit analyzes the energy value of the audio track of the live stream in real time, and determines whether there is tour guide audio data in the current time period through short-time energy analysis and zero-crossing rate detection. When no tour guide audio data is detected within a continuous preset time threshold, an interactive trigger signal is generated.

[0078] Furthermore, the energy value of the audio track of the live stream is analyzed in real time, and the presence of tour guide audio data in the current time period is determined through short-time energy analysis and zero-crossing rate detection, including:

[0079] (1) The interactive status monitoring unit first separates the independent audio track data from the live stream, and converts the separated audio track data into a mono time-domain audio signal. During the conversion process, the sampling rate is set to 16000 Hz and the sampling precision is 16 bits. The audio track data is then standardized and output as a standardized mono time-domain audio signal.

[0080] (2) Perform frame segmentation on the standardized mono time-domain audio signal, set the frame length to 25 milliseconds and the frame shift to 10 milliseconds, so that the single frame time-domain audio signal contains 400 sampling points and a 15-millisecond overlap region is formed between two adjacent frames of time-domain audio signals. After completing the frame segmentation, output the framed multi-frame time-domain audio signal, and add Hanning windows to the framed multi-frame time-domain audio signal frame by frame. The window function length of the Hanning window is consistent with the frame length. The Hanning window suppresses the spectral leakage problem at the frame edge and outputs the windowed multi-frame time-domain audio signal.

[0081] (3) Taking the windowed multi-frame time-domain audio signal as input, perform short-time energy calculation frame by frame, including: taking the square value of the amplitude of each sampling point in a single frame, summing the square values ​​of all sampling points in a single frame, obtaining the original short-time energy value corresponding to each frame, and outputting the original short-time energy value of multiple frames after completing the calculation of all frames. The maximum reference energy value is preset. This value is the maximum short-time energy value obtained by statistical analysis of a large number of audio samples under normal tour guide narration volume. The specific value is 8000. Using this maximum reference energy value as the normalization benchmark, the original short-time energy values ​​of all frames are mapped to the range of zero to one, and the short-time energy normalization process is completed. The normalized short-time energy values ​​of multiple frames are then output.

[0082] (4) Compare the normalized short-time energy values ​​of multiple frames with the preset energy threshold frame by frame, and statistically analyze the comparison results of consecutive frames. If the normalized short-time energy values ​​of 30 consecutive frames are all greater than the energy threshold, it is preliminarily determined that there is tour guide audio data in this period, and the preliminary determination result of the existence of tour guide audio is output. If the normalized short-time energy values ​​of 30 consecutive frames are all less than the energy threshold, it is preliminarily determined that there is no tour guide audio data in this period, and the preliminary determination result of the absence of tour guide audio is output. If the normalized short-time energy values ​​within 30 consecutive frames fluctuate around the energy threshold, it is determined to be a frame energy value fluctuation state, and the determination result of frame energy value fluctuation is output. The energy threshold is a value determined by statistical analysis of a large number of scenic area environmental background noise samples, and is specifically set to 0.05.

[0083] (5) Using the windowed multi-frame time-domain audio signal as input, perform zero-crossing rate calculation frame by frame, including: for each sampling point in a single frame, compare the amplitude sign of the current sampling point with that of the previous sampling point in turn. If the signs are different, it is determined to be a zero crossing. Count the total number of zero crossings in a single frame. Divide the total number of zero crossings in a single frame by the frame length of 25 milliseconds to obtain the zero crossing rate corresponding to each frame, with the unit being times per millisecond. After completing the calculation of all frames, output the zero crossing rate values ​​of the multi-frames.

[0084] (6) Compare the zero-crossing rate values ​​of multiple frames with the preset zero-crossing rate threshold frame by frame, and count the comparison results of consecutive frames. If the zero-crossing rate values ​​of 30 consecutive frames are all less than the zero-crossing rate threshold, then it is determined that there is tour guide audio data in this period, and the auxiliary determination result of having tour guide audio is output; if the zero-crossing rate values ​​of 30 consecutive frames are all greater than the zero-crossing rate threshold, then it is determined that there is no tour guide audio data in this period, and the auxiliary determination result of having no tour guide audio is output; the zero-crossing rate threshold is a value determined by statistical analysis of a large number of tour guide narration voice samples, and is specifically set to 0.3 times per millisecond.

[0085] (7) Using the preliminary judgment result in (4) and the auxiliary judgment result in (6) as dual inputs, perform a comprehensive judgment according to the preset fusion judgment rules, and output the final result of the existence status of the tour guide audio. The fusion judgment rules are as follows: if the preliminary judgment result is that tour guide audio data exists and the auxiliary judgment result is that tour guide audio data exists, then the current time period is determined to have tour guide audio data; if the preliminary judgment result is that there is no tour guide audio data and the auxiliary judgment result is that there is no tour guide audio data, then the current time period is determined to have no tour guide audio data; if the preliminary judgment result is that the frame energy value fluctuates, then the multi-frame normalized short-time energy value and multi-frame zero-crossing rate value are used as inputs, the detection time is extended to 60 frames, and the average normalized short-time energy value and average zero-crossing rate value within 60 frames are calculated. If the average normalized short-time energy value is greater than the energy threshold and the average zero-crossing rate value is less than the zero-crossing rate threshold, then the tour guide audio data exists; otherwise, the tour guide audio data is determined to have no audio data.

[0086] In this embodiment, the interactive state monitoring unit performs frame-by-frame processing on the audio track of the live stream, using a frame length of 25 milliseconds and a frame shift of 10 milliseconds. It extracts the short-time energy value and zero-crossing rate of each audio frame. The short-time energy value reflects the energy level of the audio frame; the short-time energy value of the tour guide's voice is significantly higher than that of the ambient noise. The zero-crossing rate reflects the number of times the signal crosses the zero level in the audio frame; the zero-crossing rate of the voice signal differs significantly from that of the ambient noise. The specific judgment process is as follows: First, the short-time energy value and zero-crossing rate of each audio frame are calculated and compared with a preset threshold. When the short-time energy value is higher than the energy threshold and the zero-crossing rate is lower than the threshold, the signal is considered to have passed the threshold. When the zero-crossing rate threshold is reached, it is determined that the frame contains tour guide audio data. The energy threshold and zero-crossing rate threshold are obtained through training with a large amount of tour guide voice data in scenic areas, which can accurately distinguish between tour guide voice and environmental noise. Then, the continuous audio frames are judged. When it is detected that no tour guide audio data is detected in all audio frames within 2 consecutive seconds, it is determined that there is no tour guide audio in the current time period, an interaction trigger signal is generated, and the interaction trigger signal is sent to the dialogue management unit. The preset duration threshold of 2 seconds can not only avoid false triggers caused by short pauses in the tour guide, but also trigger interactive services in a timely manner during the gaps in the tour guide's explanation, thereby improving the tourist experience.

[0087] C2: After receiving the interaction trigger signal, the dialogue management unit starts the interaction engine. Based on the tourist's current location information, it selects preset interactive content that matches the tour scenario from the cloud content library. The chat voice generated by the preset interactive content is mixed with the live stream through a time domain interpolation algorithm. After the volume ratio of the chat voice and the background ambient sound meets the preset loudness standard, it is synchronously output to the tourist terminal.

[0088] Furthermore, the chat voice generated from the preset interactive content is mixed with the live stream using a temporal interpolation algorithm, including:

[0089] (1) The intelligent interaction module first extracts two core audio data from the system. The first is the chat voice audio stream generated by speech synthesis of the preset interactive content, and the second is the original audio track data of the live stream. The two audio data are subjected to standardized preprocessing: the sampling rate is uniformly set to 48000 Hz, the sampling precision is uniformly set to 24 bits, and both audio data are converted into mono time domain audio signals to eliminate the mixing asynchrony problem caused by the difference in sampling parameters and number of channels. The standardized chat voice time domain signal and the standardized live stream time domain signal are output.

[0090] (2) With two standardized time-domain audio signals as input, and the system timestamp of the live stream as the reference time axis, extract the system timestamp of each audio frame in the audio track of the live stream and complete the calibration. According to the generation time of the interaction trigger signal of the intelligent interaction module, assign a timestamp matching the reference time axis to the chat voice time-domain signal, and mark its starting position on the reference time axis to ensure that the starting positions of the two audio time domains are accurately matched. Output the chat voice time-domain signal and the live stream time-domain signal with the timestamp synchronization and time domain position calibration completed.

[0091] (3) For the two audio signals that have completed the timestamp synchronization, configure the core parameters of the time domain interpolation algorithm: the interpolation sampling interval is set to 1 / 48000 seconds, the interpolation time precision is set to 1 millisecond, the interpolation window length is set to 10 sampling points, and at the same time, determine that a linear interpolation kernel function is used. The linear interpolation kernel function does not need to be trained. The weight coefficients are linearly distributed with the time domain position of the sampling points, and the weight coefficients of adjacent sampling points are summed to 1, forming a complete set of time domain interpolation algorithm configuration parameters.

[0092] (4) Taking the output completed timestamp synchronized chat voice time domain signal as input, combined with the configuration parameters of the time domain interpolation algorithm, perform interpolation stretching processing: first extract the frame length parameter of the live stream time domain signal, the frame length is 25 milliseconds, and each frame has 960 sampling points. Based on this, determine whether the chat voice frame length matches; if the chat voice frame length is shorter than 25 milliseconds, insert new sampling points between adjacent sampling points through the linear interpolation kernel function; if the frame length is longer than 25 milliseconds, compress the frame length through the linear interpolation kernel function, and finally make the number of sampling points of each chat voice frame uniform to 960, and output the chat voice time domain signal with standardized frame length;

[0093] (5) Using the time domain signal of the live stream with timestamp synchronization completed as input, and following the configuration parameters of the time domain interpolation algorithm, perform interpolation completion processing: check the integrity of 960 sampling points in each frame frame by frame. If there are missing sampling points, use the valid sampling points before and after the missing position as the reference, calculate and insert the complete sampling points through the linear interpolation kernel function; perform interpolation smoothing processing on the completed signal, adjust the slope of amplitude change in the amplitude change region within the frame, make the amplitude transition continuous, and output the live stream time domain signal with completed interpolation completion and smoothing.

[0094] (6) Perform amplitude normalization processing on the time domain signal of the chat voice with frame length normalization and the time domain signal of the live stream after interpolation and smoothing: map the amplitude of all sampling points to the interval between 0 and 1, where the reference amplitude of the chat voice is set to 0.7 and the reference amplitude of the live stream environment audio is set to 0.3; assign mixing weight coefficients to the two signals, where the mixing weight coefficient of the chat voice is 0.7 and the mixing weight coefficient of the live stream audio is 0.3, the weight sum is 1, and output the time domain signal of the chat voice and the time domain signal of the live stream after amplitude normalization and weight allocation.

[0095] (7) Taking the two time-domain audio signals that have completed amplitude normalization and weight allocation as input, and based on the previous timestamp synchronization and frame length matching results, perform point-by-point mixing calculation: according to the reference time axis of the live stream, align the sampling points at the same time-domain position one by one, multiply the amplitude of the chat voice sampling point by 0.7 and the amplitude of the live stream sampling point by 0.3, and add the two results to obtain the mixed amplitude at that position; traverse all time-domain positions to complete the calculation and output the preliminary mixed time-domain audio signal;

[0096] (8) Using the initial mixed temporal audio signal as input, call the temporal interpolation algorithm to configure parameters and perform interpolation optimization processing: detect amplitude continuity frame by frame, insert transition sampling points in discontinuous areas to adjust amplitude change trend; reconstruct the mixed signal according to the original frame structure of the live stream, add a timestamp consistent with the live stream to each frame, and output the mixed audio temporal signal that has completed interpolation optimization and frame structure reconstruction.

[0097] In this embodiment, after receiving the interaction trigger signal, the dialogue management unit immediately activates the built-in interaction engine. This engine is a hybrid dialogue engine based on rules and deep learning, enabling natural and fluent conversational interaction. First, the interaction engine sends a query request to the cloud content library based on the tourist's current location information sent by the tourist's terminal. It then filters out preset interactive content matching the scenic area's corresponding scene. For example, if the tourist is in the ancient temple area, the engine filters interactive content about the temple's history, architectural features, and cultural anecdotes; if the tourist is in the mountain trail area, the engine filters interactive content about mountain scenery, hiking techniques, and the scenic area's ecology. All preset interactive content is pre-recorded and generated voice content, covering various aspects of scenic area knowledge, travel suggestions, and fun Q&A, meeting the diverse interactive needs of tourists. Then, the interaction engine processes the filtered preset interactive content... The audio is converted into a chat voice stream, generated using TTS (Text-to-Speech) technology. The synthesized voice is a natural and friendly female voice, with a speech rate of 200 words per minute, conforming to the general public's listening habits. Next, a multi-track audio mixing process is performed on the audio tracks of the chat voice stream and the live stream using a time-domain interpolation algorithm. This algorithm can accurately align the audio timeline, avoiding audio stuttering or dropouts after mixing. Finally, during the mixing process, the volume ratio of the chat voice to the background ambient sound is strictly controlled. The preset loudness standard is that the volume of the chat voice is 10dB higher than the background ambient sound, ensuring that the chat voice is clear and audible without masking the scenic environment sounds in the live stream, maintaining an immersive experience. The mixed audio stream and the panoramic video stream are simultaneously output to the head-mounted display device on the tourist's end, achieving synchronized playback of the chat voice and the live stream.

[0098] Furthermore, the intelligent interaction module also includes an emotion computing submodule. In this embodiment, tourists' emotional states during remote guided tours vary, such as curiosity, pleasure, doubt, and boredom. The emotion computing submodule can accurately identify tourists' emotional states and adjust the interaction strategy accordingly, making the intelligent companionship more personalized and humanized. The specific implementation process is as follows:

[0099] (1) During the conversational interaction, the acoustic features of the tourist's voice commands are obtained through the speech recognition module, and the fundamental frequency trajectory, energy jitter, and speech rate features of the voice commands are extracted. These features are then input into the emotion classification model based on support vector machines to output the tourist's current emotion state category. In this embodiment, while collecting the tourist's voice commands, the speech recognition module synchronously transmits the original acoustic data of the voice to the emotion computing submodule. The emotion computing submodule extracts three core emotion features from the acoustic data, namely the fundamental frequency trajectory, energy jitter, and speech rate features. The fundamental frequency trajectory reflects the pitch change of the voice. Pleasant emotions usually correspond to a higher and more stable fundamental frequency, while confused emotions usually correspond to a sudden change in the fundamental frequency. Energy jitter reflects the volume fluctuation of the voice. Active emotions typically correspond to larger energy fluctuations, while speech rate features reflect the speed of speech; boredom usually corresponds to a slower speech rate, and curiosity usually corresponds to a moderate speech rate. After normalization, the extracted features are input into a support vector machine-based emotion classification model. This model, trained on massive amounts of human emotional speech data, categorizes tourists' emotional states into five categories: pleasure, curiosity, doubt, boredom, and indifference. The model's classification accuracy is no less than 90%. Based on the input features, the model outputs the probability distribution of the tourist's current emotional state category and selects the category with the highest probability as the final emotional state judgment result. For example, when the probability of the pleasure category is 0.85, and the probabilities of the other categories are all below 0.2, the tourist's current emotional state is determined to be pleasure.

[0100] (2) The dialogue management unit adjusts the speech rate, tone synthesis parameters, and affinity level of the reply content of the accompanying voice according to the emotional state category. In this embodiment, the dialogue management unit presets corresponding accompanying voice synthesis parameters and affinity level of the reply content for different emotional state categories. The specific adjustment strategy is as follows: when the tourist's emotional state is pleasant, the speech rate of the accompanying voice is adjusted to 220 words per minute, the tone is increased by 5Hz, the affinity level of the reply content is adjusted to the highest, and more lively and interesting interactive content is selected; when the tourist's emotional state is curious, the normal speech rate and tone of the accompanying voice are maintained, the affinity level of the reply content is adjusted to high, and more detailed and professional scenic spot knowledge content is selected to satisfy the tourist's curiosity; when the tourist's emotional state is confused, the accompanying voice is adjusted to a higher level. The speech rate is adjusted to 180 words per minute, the pitch is lowered by 3Hz, and the friendliness level of the replies is adjusted to medium-high. More easily understood and patiently explained content is selected to accurately address tourists' questions. When tourists are bored, the speech rate is adjusted to 210 words per minute, the pitch is raised by 3Hz, and the friendliness level of the replies is adjusted to medium. More interesting and novel interactive content is selected, such as interesting anecdotes about the scenic area and interactive Q&A, to stimulate tourists' interest. When tourists are indifferent, the normal parameters of the speech are maintained, the friendliness level of the replies is adjusted to a basic level, and conventional scenic area tour content is selected to maintain normal interactive communication. Through these personalized adjustments, the intelligent chatbot better meets the emotional needs of tourists and enhances their interactive experience.

[0101] Furthermore, the intelligent interaction module also includes a multimodal interaction fusion submodule. In the remote tour scenario of the mountain scenic area in this embodiment, when tourists wear head-mounted display devices, simple voice interaction may have problems such as inconvenience in operation and inaccurate expression of instructions. The multimodal interaction fusion submodule supports multimodal interaction of voice and gestures, and can fuse tourists' text instructions and gestures to accurately identify tourists' interaction intentions, thereby improving the accuracy and convenience of interaction. The specific implementation process is as follows: During the interaction process, the multimodal interaction fusion submodule simultaneously receives text instructions output by the voice recognition module and gesture image data collected by the tourist's camera. The gesture image data is processed using a gesture classification network based on a convolutional neural network to identify the gesture category and pointing direction. A multimodal fusion network based on an attention mechanism is used to align and fuse the semantic features of the text instructions with the gesture category features to generate a comprehensive interaction intention vector containing location pointing and semantic description. The corresponding response content is retrieved from the cloud content library based on the comprehensive interaction intention vector.

[0102] It should be explained that the gesture classification network is a convolutional neural network model specifically built for panoramic tour guide scenarios. It uses a residual network as the backbone structure and consists of an input layer, multiple convolutional layers, pooling layers, fully connected layers, and an output layer. The convolutional layers are used to extract visual features such as contours, key points, and joint positions of gesture images layer by layer. The pooling layers are used to reduce the dimensionality of visual features and retain key information. The fully connected layers are used to map high-dimensional visual features into classification features. The output layer is used to output the gesture category and pointing direction. The training of this gesture classification network is completed using a large number of gesture image samples collected in remote tour guide scenarios. The training process includes image acquisition, data annotation, data augmentation, dataset partitioning, forward propagation calculation, loss value calculation, backpropagation to update network parameters, and iterative convergence. After training, it can stably recognize eight common gestures: pointing, clicking, waving, clenching fist, pointing up, pointing down, pointing left, and pointing right. It also calculates the pointing direction based on the key point positions of the hand bones. The recognition results meet the positioning accuracy requirements of panoramic tour guide interaction.

[0103] It should also be explained that the attention-based multimodal fusion network is a fusion model specifically adapted for voice and gesture collaborative interaction. It consists of a text semantic encoding branch, a gesture feature encoding branch, an attention-weighted alignment layer, a feature fusion layer, and an intent output layer connected sequentially. The text semantic encoding branch converts the text commands output by speech recognition into fixed-dimensional semantic feature vectors, extracting core semantic information such as attraction names, operation intentions, and query content. The gesture feature encoding branch converts the gesture category and pointing direction output by the gesture classification network into fixed-dimensional spatial pointing feature vectors. The attention-weighted alignment layer automatically calculates the correlation weights between text semantics and gesture features, strengthening key information related to the interaction intent and weakening irrelevant information to achieve accurate alignment of the two modal features. The feature fusion layer concatenates and integrates the weighted aligned semantic features and gesture features to generate comprehensive interactive features that cannot be fully expressed by a single text command or gesture command. The intent output layer converts the fused features into a fixed-dimensional comprehensive interactive intent vector, which simultaneously contains the tourist's semantic query needs and the specific location information of the gesture pointing. This multimodal fusion network is trained using real interactive samples from panoramic tours. During training, text commands, gesture images, and real interactive intentions are used as supervisory information. The network weights are continuously optimized through backpropagation, enabling the model to accurately understand composite commands combining voice and gestures. In actual operation, the multimodal interaction fusion submodule sends the generated comprehensive interactive intention vector to the cloud content library. The cloud content library accurately retrieves matching explanations, tour information, or interactive responses based on the semantic and location information in the vector, achieving more accurate and more tailored tour interactive responses to the tourists' intentions.

[0104] In this embodiment, the head-mounted display device at the visitor end has a built-in binocular depth camera, capable of acquiring hand gesture image data of the visitor at a frequency of 30 frames per second. The camera's recognition range is 0.5-2 meters in front of the visitor, accurately capturing various simple gestures such as pointing, clicking, waving, and clenching fists. The multimodal interaction fusion submodule simultaneously receives text commands from the visitor output by the speech recognition module and gesture image data acquired by the camera. First, the gesture image data is processed and recognized through a gesture classification network based on a convolutional neural network. This network uses ResNet50 as the backbone network and, after training with massive amounts of gesture image data, can recognize eight common gesture categories, including pointing, clicking, waving, clenching fists, pointing up, pointing down, pointing left, and pointing right. Simultaneously, it can accurately determine the pointing direction of the gesture based on the position of the key points of the hand bones, with a pointing direction recognition error of no more than 5°. The overall accuracy of gesture classification and pointing direction recognition is no less than 95%. Then, semantic features are extracted from the text commands using a semantic feature extraction method based on a bidirectional long short-term memory network. The feature extraction model converts text commands into 512-dimensional semantic feature vectors. Next, a multimodal fusion network based on an attention mechanism aligns and fuses the semantic feature vectors of the text commands with the category and direction feature vectors of gesture recognition. The attention mechanism automatically focuses on key information related to the interaction intent in the text commands and gestures. For example, when a tourist says "the history of this attraction" and simultaneously points to the ancient temple in the live stream, the attention mechanism focuses on the "history of the attraction" in the text and the "pointing to the ancient temple" information in the gesture, achieving accurate feature fusion. Finally, a 1024-dimensional comprehensive interaction intent vector containing location and semantic description is generated. This vector includes both the semantic content the tourist wants to query and the specific location of the scenic spot the tourist is pointing to. Finally, the multimodal interaction fusion submodule sends a query request to the cloud content library based on this comprehensive interaction intent vector, retrieves the corresponding response content, such as the historical explanation of the ancient temple, and converts it into a chat-like voice output to the tourist, achieving accurate multimodal interactive response.

[0105] It should be explained that the semantic feature extraction model is a text semantic encoding network specifically built for remote scenic area navigation scenarios. It consists of an input embedding layer, a bidirectional long short-term memory layer, a temporal pooling layer, and a feature output layer connected sequentially. The input embedding layer converts segmented text instructions into fixed-dimensional word vectors, mapping text symbols to computable numerical features. The bidirectional long short-term memory layer includes two branches: forward temporal computation and backward temporal computation. It can simultaneously extract contextual semantic information from both the preceding and following text, effectively capturing key navigation-related semantics such as scenic spot names, operational intentions, query objects, and location descriptions. The temporal pooling layer compresses and aggregates the temporal features output by the bidirectional long short-term memory layer, retaining global semantic information and removing redundant temporal data. The feature output layer maps the aggregated features into a fixed-length 512-dimensional semantic feature vector for multimodal fusion. This semantic feature extraction model is trained using massive amounts of voice command text from the scenic area tour guide domain. The training steps are as follows: text acquisition, corpus annotation, data cleaning, dataset partitioning, word vector initialization, forward propagation calculation, loss function optimization, backpropagation updating network parameters, and iterative convergence. The training process uses text semantic classification and intent recognition as supervised objectives. After optimization, it can accurately extract the core semantics of text in tour guide scenarios and output a stable and unified 512-dimensional semantic feature vector. Next, a multimodal fusion network based on an attention mechanism is used to align and fuse the semantic feature vector of text commands with the category feature vector and pointing direction feature vector of gesture recognition. The attention mechanism can automatically focus on key information related to the interaction intent in text commands and gesture actions. For example, when a tourist says "the history of this attraction" and points to the ancient temple in the live stream, the attention mechanism will focus on the "history of the attraction" in the text and the "pointing to the ancient temple" information in the gesture, achieving accurate feature fusion.

[0106] During the interaction, the voice recognition module collects tourists' voice commands, parses them, retrieves the corresponding response content from the pre-stored cloud content library, and generates a chat voice after speech synthesis.

[0107] The speech recognition module includes a front-end speech processing unit and a back-end semantic understanding unit;

[0108] D1: The front-end voice processing unit collects tourists' voice commands through the microphone array at the tourist end, applies the least mean square adaptive filtering algorithm to suppress the collected environmental noise, and extracts the Mel frequency cepstral coefficients as acoustic features. The least mean square adaptive filtering algorithm and the calculation formula of the Mel frequency cepstral coefficients are both existing technologies in this field and are not inventive solutions of this application, and will not be described in detail here.

[0109] In this embodiment, the head-mounted display device at the visitor end has a built-in 4-microphone linear microphone array with a sampling rate of 48kHz and a bit depth of 16 bits, enabling accurate acquisition of the visitor's voice commands. The microphone array also features beamforming capabilities, effectively improving the acquisition of target sound sources. Since environmental noise, such as home ambient noise or outdoor noise, may exist during the remote guided tour, affecting the quality of voice command acquisition, the front-end voice processing unit applies a least mean square adaptive filtering algorithm to suppress noise in the acquired audio data. This algorithm dynamically adjusts the filtering coefficient based on real-time changes in environmental noise, effectively suppressing background noise while preserving the visitor's voice signal. The signal-to-noise ratio of the voice commands is no less than 25dB. After noise suppression, the front-end voice processing unit preprocesses the voice signal, including pre-emphasis, framing, and windowing. Pre-emphasis uses a first-order high-pass filter to enhance the high-frequency components of the voice signal. Framing is done with a frame length of 25 milliseconds and a frame shift of 10 milliseconds. Hanning windowing is used to reduce signal leakage between frames. Finally, Mel-frequency cepstral coefficients are extracted from the preprocessed voice signal as acoustic features. Mel-frequency cepstral coefficients can simulate the auditory characteristics of the human ear and accurately capture the acoustic features of the voice. In this embodiment, 13-dimensional Mel-frequency cepstral coefficients and their first and second-order differences, totaling 39-dimensional features, are extracted as input to the back-end semantic understanding unit.

[0110] D2: The backend semantic understanding unit inputs the extracted Mel frequency cepstral coefficients into the end-to-end speech recognition model trained based on connection time-series classification, decodes and outputs the corresponding text instruction, and uses a named entity recognition algorithm based on bidirectional long short-term memory network and conditional random field to extract key information including the name of the scenic spot and the operation intention from the text instruction. Based on the key information, a query vector is constructed, and the matching response content is recalled from the cloud content library.

[0111] Furthermore, the specific steps of D2 include:

[0112] (1) The back-end semantic understanding unit receives the Mel frequency cepstral coefficients extracted by the front-end speech processing unit. The coefficients are 39-dimensional and include static coefficients and first-order and second-order difference coefficients. The Mel frequency cepstral coefficients are format adapted and the feature sequence is divided according to the time step. Each time step corresponds to a 10-millisecond speech segment. The length of the feature sequence is dynamically adjusted according to the duration of the speech command so that the time dimension of the input feature matches the requirements of the input layer of the end-to-end speech recognition model. The output is the format adapted Mel frequency cepstral coefficient feature sequence.

[0113] (2) An end-to-end speech recognition model based on connection-time classification training was constructed. The end-to-end speech recognition model adopts an encoder-decoder architecture, and the model training parameters are configured and trained simultaneously. The encoder module consists of a 6-layer unidirectional long short-term memory network with 512 hidden units in each layer. It can receive 39-dimensional Mel-frequency cepstral coefficient feature sequences and extract temporal acoustic features, outputting a 1024-dimensional encoded feature sequence. The decoder module is a fully connected layer structure based on the attention mechanism. The attention window size is 15 time steps, the number of neurons in the fully connected layer is 512, and the activation function is the ReLU function. The connection-time classification loss is used. The first layer is the model loss calculation module, with the blank label weight coefficient set to 0.1, used to calculate the loss value between the predicted sequence and the real text sequence. The model training uses a speech command corpus in the scenic area tour scenario, containing 100,000 valid speech samples. The training batch size is set to 64, the initial learning rate is set to 0.001, and a cosine annealing strategy is used to adjust the learning rate. The learning rate decreases by 0.1 every 10 training rounds, and the total number of training rounds is set to 100 rounds. The validation set accounts for 20% of the total dataset. Training stops when the character error rate of the validation set no longer decreases for 10 consecutive rounds. The optimal model parameters are saved, and the trained end-to-end speech recognition model is output.

[0114] (3) Based on the format-adapted Mel frequency cepstral coefficient feature sequence, load the trained end-to-end speech recognition model to perform decoding operation, including: the encoder module performs temporal feature encoding on the input feature sequence, the decoder module focuses on the key time step of the encoded feature sequence through the attention mechanism, and combines the training parameters of the connected temporal classification loss layer to map the acoustic feature sequence into a text character sequence. The decoding process adopts the beam search algorithm, the beam width is set to 10, and finally outputs the text instruction corresponding to the speech instruction. The text instruction is a natural language string that contains the tourist's interactive intention and related information.

[0115] (4) Construct a named entity recognition model based on bidirectional long short-term memory network and conditional random field adapted to scenic spot tour scenarios, and complete the model training parameter configuration and training. The input layer of the named entity recognition model can receive word vector sequences after text instruction word segmentation, and the word vector dimension is set to 256 dimensions; the bidirectional long short-term memory network layer contains two layers of bidirectional long short-term memory network, each with 128 forward and backward hidden units, which can extract the contextual semantic features of the word vector sequence and output a 512-dimensional semantic feature sequence; the conditional random field layer is the model sequence labeling layer, and the label set includes scenic spot name, operation intention, The three categories of irrelevant information labels were used. The initial values ​​of the transition matrix were initialized with a uniform distribution, ranging from -0.1 to 0.1. The model was trained using a corpus of 50,000 scenic area guide text instructions with labeled entity types. The batch size was set to 32, the learning rate was set to 0.0005, the Adam optimizer was used, the weight decay coefficient was set to 0.0001, the training epochs were set to 80, and the validation set accounted for 15% of the total dataset. The labeling accuracy was used as the model evaluation metric. Training was stopped when the validation set accuracy reached 95%, the model parameters were saved, and the trained named entity recognition model was output.

[0116] (5) Based on the output text instruction, perform word segmentation processing, divide the text instruction into word sequences of the smallest semantic units, and then convert the word sequence into a 256-dimensional word vector sequence compatible with the named entity recognition model, and output the converted word vector sequence.

[0117] (6) Based on the word vector sequence, load the trained named entity recognition model and perform recognition operation. Specifically, input the word vector sequence into the bidirectional long short-term memory network layer of the model to extract contextual semantic features, and then perform sequence labeling through the conditional random field layer to label the semantic units in the text instruction that belong to the scenic spot name and operation intention. Perform post-processing on the labeling results, merge consecutive labeling units of the same type, and finally extract the key information of the scenic spot name and operation intention in the text instruction, and output a set of structured key information containing the scenic spot name string and the operation intention string.

[0118] (7) Based on the output set of structured key information, construct a query vector, including: first, using a semantic coding model in the scenic area domain to convert the scenic spot name and operation intention into 256-dimensional semantic vectors respectively. The semantic coding model has 3 hidden layers and 256 neurons in each layer. Then, the two 256-dimensional semantic vectors are concatenated into a 512-dimensional query vector. The dimension of the query vector is consistent with the dimension of the feature vector of the reply content in the cloud content library, and the 512-dimensional query vector is output.

[0119] (8) Based on the 512-dimensional query vector, the response content recall operation is performed in the cloud content library. All response content in the cloud content library is pre-encoded as a 512-dimensional feature vector and associated with the corresponding scenic spot name and operation intention tag. The cosine similarity between the input query vector and the feature vector of all response content in the cloud content library is calculated. The similarity threshold is set to 0.85. Response content with similarity greater than the similarity threshold is filtered out. The filtered response content is sorted from high to low similarity. The top three response content are taken as the recall result. The set of matched response content is output. The entire processing flow of the backend semantic understanding unit is completed.

[0120] In this embodiment, the core of the backend semantic understanding unit is an end-to-end speech recognition model trained based on connection-time classification. This model adopts an architecture combining convolutional neural networks, bidirectional long short-term memory networks, and connection-time classification. It eliminates the need for complex phoneme alignment and can directly convert acoustic features into text commands. The model has been trained on massive amounts of Chinese speech data, covering different accents and speaking speeds, and incorporates a large number of tourism-related professional terms. It can accurately recognize tourists' scenic area guide-related voice commands with an accuracy rate of no less than 92%, even if tourists have slight accents. After processing the input 39-Vimel frequency cepstral coefficients, the model decodes and outputs the corresponding text commands. Subsequently, a named entity recognition algorithm based on bidirectional long short-term memory network and conditional random field is adopted to extract key information from text commands. This algorithm can accurately identify named entities and operation intentions in text. In the scenic area guide scenario, the main key information extracted includes the name of the attraction and the operation intention, and the accuracy of named entity recognition is no less than 95%. After extracting the key information, the backend semantic understanding unit constructs a query vector based on this key information. The query vector contains the features of the name of the attraction and the operation intention. Then, the query vector is sent to the semantic similarity search interface of the cloud content library. The interface retrieves the matching response content from the cloud content library and returns the response content to the intelligent interaction module to achieve accurate response to voice commands.

[0121] Furthermore, the speech recognition module also includes a multilingual recognition and translation submodule. After collecting tourists' voice commands, the multilingual recognition and translation submodule first identifies the language of the voice command using a phoneme-level language recognition algorithm. When the identified language is inconsistent with the default language of the explanatory content in the cloud content library, it calls a neural machine translation model based on the Transformer architecture to translate the voice command into a text query in the default language, and then uses speech synthesis technology to reverse synthesize the reply content retrieved from the cloud content library into a chat voice output in the tourist's language. In this embodiment, the default language of the content in the cloud-based content library is Chinese. The multilingual recognition and translation submodule supports the recognition and translation of six mainstream languages: Chinese, English, Japanese, Korean, French, and Spanish. First, a phoneme-level language recognition algorithm is used to identify the language of the collected tourist voice commands. This algorithm extracts the phoneme features of the voice commands. Since phoneme features differ significantly between languages, the algorithm determines the language of the voice command by matching the extracted phoneme features with preset phoneme feature templates for each language. The accuracy rate of language recognition is no less than 98%. When the identified language is Chinese, the voice command is directly transmitted to the backend semantic understanding unit for processing without translation. When the identified language is another foreign language, such as English or Japanese, a neural machine translation model based on the Transformer architecture is immediately invoked for translation. This model has been trained on a massive amount of bilingual parallel corpus in the tourism field and can accurately achieve mutual translation between foreign languages ​​and Chinese, with an accuracy rate of no less than 90%. It can accurately capture the semantic information of tourist voice commands, for example, the English command "Tell me the history of this ancient..." The accurate translation of "temple" is "to explain the history of this ancient temple." The translated Chinese text query is transmitted to the backend semantic understanding unit for key information extraction and content retrieval, retrieving the corresponding Chinese response from the cloud content library. Finally, the retrieved Chinese response is synthesized into a chat voice in the tourist's language using TTS speech synthesis technology. The synthesized speech pronunciation conforms to the pronunciation norms of that language, with a natural and fluent tone. Ultimately, the synthesized foreign language chat voice is output to the tourist's end, enabling barrier-free intelligent interaction for overseas tourists and allowing them to smoothly experience remote guided tours of mountain scenic areas.

[0122] Based on the changes in the tourist's location identified by the positioning module, the augmented reality module retrieves the corresponding augmented reality content from the cloud content library and overlays it onto the live stream using virtual reality technology.

[0123] The augmented reality module includes a pose calculation unit and a virtual-real fusion rendering unit;

[0124] E1: The pose calculation unit calculates the six-degree-of-freedom pose parameters of the tourist in three-dimensional space in real time through a visual simultaneous positioning and mapping algorithm based on the tourist position information output by the positioning module and the posture data collected by the inertial measurement unit inside the head-mounted display device.

[0125] Furthermore, the six-DOF pose parameters of the tourist in three-dimensional space are calculated in real time using a visual simultaneous localization and mapping (SMR) algorithm, including:

[0126] (1) The pose calculation unit receives two data streams. The first stream is the tourist location information output by the positioning module, which includes the tourist's three-dimensional spatial coordinates and positioning accuracy data based on the global three-dimensional coordinate system of the scenic area. The second stream is the attitude data collected by the inertial measurement unit inside the head-mounted display device, which includes three-axis acceleration, three-axis angular velocity, and three-axis magnetometer data. The sampling frequency of the inertial measurement unit is set to 100 Hz. Then, the two data streams are subjected to standardized preprocessing: the coordinate data of the tourist location information is converted into floating-point numerical format, and the unit is unified to meters; the inertial measurement unit attitude data is subjected to dimension conversion, with the acceleration unit unified to meters per second squared, the angular velocity unit unified to radians per second, and the magnetometer data normalized to the numerical range of zero to one; at the same time, the two data streams are time-stamped and synchronized, with the synchronization accuracy set to 1 millisecond, and the tourist location information and inertial measurement unit attitude data that have completed standardized preprocessing and time-stamp synchronization are output.

[0127] (2) For the remote tour scenario of scenic spots, the parameters of the visual simultaneous localization and map building algorithm are configured, including: the feature point extraction threshold is set to 0.04, the feature point matching distance threshold is set to 0.6, the algorithm frame processing frequency is set to 30 frames per second, the motion estimation window size is set to 10 frames, the map point update frequency is set to once every 5 frames, the key frame selection threshold is set to the feature point matching number is less than 80, the global optimization execution interval is set to 20 frames, and the inertial measurement unit data fusion weight coefficients are set to acceleration 0.3, angular velocity 0.7, and magnetometer 0.2. Each weight coefficient is an independent confidence coefficient of each sensor data, not a normalized weighted sum, used to characterize the credibility and influence of acceleration, angular velocity, and magnetometer data in attitude calculation and position fusion. It can be independently adjusted according to the actual noise characteristics of the sensor data in the scenic tour scenario. Finally, the configured visual simultaneous localization and map building algorithm parameter set is output.

[0128] (3) Based on the real-time video frames of the live stream, combined with the parameter set of the visual simultaneous localization and map construction algorithm, visual feature extraction and matching are performed, including: using the scale-invariant feature transformation algorithm to extract feature points of each frame image, the number of extractions is controlled between 200 and 300, retaining the pixel coordinates, scale, and orientation information of the feature points, and then calculating the similarity of feature point descriptors of adjacent frames through Euclidean distance. The smaller the Euclidean distance, the more similar the feature point descriptors are. The distance threshold is set to 0.6, and feature points with an Euclidean distance less than or equal to 0.6 are judged as valid matching pairs. The random sampling consistency algorithm is used to remove incorrect matching pairs, with 200 iterations and an inlier threshold of 2 pixels. The effective feature point set of each frame image and the effective matching pairs of feature points of adjacent frames are output.

[0129] (4) Based on the effective matching pairs of feature points in adjacent frames, combined with the global three-dimensional coordinate system parameters of the scenic area, the initial calculation of visual motion estimation is performed. Specifically, based on the pixel coordinates of the feature points and the camera intrinsic parameters, the two-dimensional pixel feature points are converted into three-dimensional spatial rays; the coordinates of the three-dimensional map points corresponding to the feature points are calculated through the triangulation measurement algorithm, the triangulation reprojection error threshold is set to 1 pixel, and the point cloud data with excessive reprojection error is removed; the camera preliminary motion transformation matrix containing rotation and translation information is solved by utilizing the changes in the three-dimensional map point coordinates of adjacent frames, and the camera preliminary motion transformation matrix based on visual features is output. The triangulation measurement algorithm is the prior art in this field and is not an inventive solution of this application, so it will not be described in detail here.

[0130] (5) Based on the standardized inertial measurement unit attitude data and combined with the fusion weight coefficient, perform pre-integration processing: integrate the three-axis angular velocity data with an integration time step of 0.01 seconds to obtain the tourist head rotation attitude increment, and update the zero bias compensation every 10 frames during the integration process; first remove the influence of gravitational acceleration from the three-axis acceleration data, and then integrate it a second time to obtain the position increment, and combine it with the magnetometer data to correct the rotation attitude, where the correction angle error threshold is 1 degree; fuse the rotation attitude increment, position increment, and magnetometer correction data according to the fusion weight coefficient, and output the attitude increment and position increment data after pre-integration of the inertial measurement unit;

[0131] (6) Based on the camera's initial motion transformation matrix and the attitude and position increment data after pre-integration by the inertial measurement unit, perform multi-source data fusion optimization: convert the attitude and position increments of the inertial measurement unit into motion transformation matrix format and align them with the camera's initial motion transformation matrix; use the extended Kalman filter algorithm to fuse the data, set the state vector dimension to 15 dimensions, the diagonal elements of the process noise covariance matrix to 0.01, and the diagonal elements of the observation noise covariance matrix to 0.05, and calculate and correct the motion transformation matrix error through prediction step and update step iterations, and output the fused and optimized camera motion transformation matrix;

[0132] (7) Based on the fused and optimized camera motion transformation matrix and effective feature point set, perform key frame selection and local map construction: according to the key frame selection threshold, if the number of feature point matches is less than 80, determine the new key frame; associate the key frame feature points with the corresponding 3D map points to construct a local map with a coverage area of ​​10 meters; deduplicate and optimize the 3D map points in the local map, remove duplicates and error exceeding the standard points, and output the selected key frame and the constructed local 3D map.

[0133] (8) Taking keyframes and local 3D maps as input, perform local map optimization and initial pose solution, including: optimizing the local 3D map using bundle adjustment method, setting the number of iterations to 50 and the convergence threshold to 0.001, the optimization objective is to minimize the feature point reprojection error, and the optimization variables include keyframe pose parameters and 3D map point coordinates; based on the optimized local 3D map, solve the camera pose parameters corresponding to the current frame, i.e., the initial pose solution of the tourist's six degrees of freedom, including the X, Y, and Z position parameters in the global 3D coordinate system of the scenic area, the tourist's head yaw angle, pitch angle, and roll angle attitude parameters, and output the initial pose solution parameters of the tourist's six degrees of freedom. Among them, the bundle adjustment method is the prior art in this field and is not an inventive solution of this application, so it will not be described in detail here.

[0134] (9) Based on the initial six-degree-of-freedom pose parameters of the tourist and the tourist position information, perform global optimization and precise pose calculation: take the tourist position information output by the positioning module as the reference benchmark, and correct the initial pose parameters in combination with the constraints of the global three-dimensional coordinate system of the scenic area; when the position parameters of the initial pose solution deviate from the position information output by the positioning module by more than 0.5 meters, start the global optimization process, and use the graph optimization algorithm to optimize all key frame poses and the global map. In the graph optimization algorithm, the nodes are key frame poses and the edges are motion constraints between key frames. Set the number of optimization iterations to 30 times; after correcting the global error of the pose parameters, the final six-degree-of-freedom pose parameters of the tourist in three-dimensional space are obtained, where the calculation accuracy is not less than 0.1 meters and the pose calculation accuracy is not less than 1°. Output the precise pose parameters to complete the entire pose calculation process.

[0135] In this embodiment, the six-degree-of-freedom pose parameters include three position parameters (X, Y, Z) and three attitude parameters, including yaw angle, pitch angle, and roll angle, which can accurately describe the tourist's position and head viewing posture in the three-dimensional space of the virtual scenic area. The input to the pose calculation unit is the tourist's position information output by the tourist-end positioning module and the attitude data collected by the inertial measurement unit inside the head-mounted display device. The inertial measurement unit collects the three-axis acceleration, three-axis angular velocity, and three-axis magnetometer data of the tourist's head at a frequency of 100Hz, reflecting the real-time attitude changes of the tourist's head. The core calculation algorithm is a visual simultaneous localization and mapping (VLM) algorithm, which can combine visual data and inertial measurement data to achieve real-time pose calculation and mapping. In this embodiment, the visual data for the simultaneous visual localization and map building algorithm comes from panoramic video frames of the live stream. By extracting and matching feature points in the video frames and combining them with inertial measurement data, multi-sensor data fusion is achieved. First, feature points are extracted from the panoramic video frames using a scale-invariant feature transform algorithm. The extracted feature points are consistent with those extracted by the panoramic acquisition module, ensuring the accuracy of feature matching. Then, the extracted visual feature points are fused with the inertial measurement data, and the extended Kalman filter algorithm is used for state estimation to calculate the six-degree-of-freedom pose parameters of the tourist in real time. The pose calculation update frequency is 50Hz, which can accurately track the changes in the tourist's position and head posture.

[0136] E2: The virtual-real fusion rendering unit, based on six-degree-of-freedom pose parameters, calls the corresponding augmented reality 3D model resources from the cloud content library, uses a depth testing algorithm based on occlusion detection to calculate the visibility of the augmented reality 3D model from the current viewpoint, and accurately registers and superimposes the augmented reality 3D model onto the corresponding pixel area of ​​the live stream through perspective projection transformation, generating an augmented reality output stream with spatial consistency. The specific process of perspective projection transformation is existing technology in this field and is not an inventive solution of this application, so it will not be described in detail here.

[0137] Furthermore, a depth testing algorithm based on occlusion detection is used to calculate the visibility of the augmented reality 3D model from the current viewpoint, including:

[0138] (1) First, load the three-dimensional model data of the scenic area to be detected. The three-dimensional model of the scenic area uses triangular facets as the basic geometric units. Each triangular facet contains the three-dimensional coordinates and normal vector information of the three vertices. The coordinate system of the model is consistent with the global three-dimensional coordinate system of the scenic area for tourist pose calculation. At the same time, load the current view parameters, including the six-degree-of-freedom pose parameters output by the pose calculation unit and the camera intrinsic parameters of the head-mounted display device. Convert the view parameters into the observation matrix of the three-dimensional model space. The translation component of the observation matrix corresponds to the position coordinates of the tourist, and the rotation component corresponds to the yaw angle, pitch angle and roll angle of the tourist's head. Output the loaded three-dimensional model triangular facet data and the current view observation matrix.

[0139] (2) Configure the core parameters of the depth test algorithm based on occlusion detection: The resolution of the depth buffer is set to 1920×1080, which is consistent with the display resolution of the head-mounted display device; the storage precision of the depth value is set to 32-bit floating point type, and the depth range is set to 0.1 meters to 100 meters, covering the effective observation distance of the scenic area guide scene; the depth threshold of occlusion detection is set to 0.05 meters, that is, when the difference in depth values ​​between two triangular facets is less than 0.05 meters, it is judged as near-distance occlusion; the near clipping face distance of the view frustum clipping is set to 0.1 meters, and the far clipping face distance is set to 100 meters, removing the three-dimensional model triangular facets outside the view frustum, and outputting the configured depth test algorithm parameter set;

[0140] (3) Based on the triangular facet data of the 3D model and the current view observation matrix, combined with the configured algorithm parameters, the triangular facet projection and depth value calculation are performed, including: converting the vertex coordinates of the triangular facets of the 3D model into coordinates in the camera coordinate system through the observation matrix, and then converting them into pixel coordinates in the 2D screen coordinate system through the perspective projection matrix. The vertical field of view of the perspective projection matrix is ​​set to 60 degrees and the aspect ratio is set to 16:9. For each pixel area covered by the triangular facet, the depth value in the depth buffer is calculated. The depth value is the Z-axis coordinate of the vertex in the camera coordinate system. The depth value of the pixels inside the triangular facet is calculated by linear interpolation. The distribution of the screen pixel area and the depth value in the depth buffer corresponding to each triangular facet is output.

[0141] (4) Based on the depth value distribution of the triangular facets, perform depth comparison and marking for occlusion detection. Specifically, the triangular facets are layered according to the spatial position of the 3D model and sorted from near to far from the camera viewpoint, with a sorting step size of 0.1 meters. Each triangular facet is traversed sequentially, and its depth value is compared with the depth value stored at the corresponding pixel position in the depth buffer. If the depth value of the current triangular facet is greater than the stored depth value and the depth difference is greater than the set depth threshold of 0.05 meters, the triangular facet is determined to be occluded by the facet in front and marked as invisible. If the depth value of the current triangular facet is less than the stored depth value, the depth value of the corresponding pixel in the depth buffer is updated and the triangular facet is marked as visible. If the depth difference is less than 0.05 meters, the angle between the normal vector of the triangular facet and the viewpoint direction is further detected. If the angle is greater than 90 degrees, it is determined to be back occlusion and marked as invisible. The set of triangular facets that have completed the visibility status marking is output.

[0142] (5) Based on the set of triangular facets marked with visibility status, perform view frustum clipping and visible region filtering, including: according to the configured near clipping and far clipping parameters, remove triangular facets that exceed the view frustum range; for the remaining triangular facets, check whether they are completely inside the view frustum, partially inside the view frustum, or completely outside the view frustum, clip the facets that are partially inside the view frustum, retain the part inside the view frustum and recalculate its depth value; filter out all triangular facets marked as visible and located inside the view frustum, and output the visible triangular facet subset of the 3D model;

[0143] (6) Based on the subset of visible triangular facets, perform visibility verification and region merging. Specifically, for each visible triangular facet, verify the angle between its normal vector and the view direction. The angle threshold is set to 120 degrees. Triangular facets with angles exceeding the angle threshold are removed. Adjacent visible triangular facets are merged. The merging criteria are that the vertex distance between the triangular facets is less than 0.1 meters and the angle between their normal vectors is less than 10 degrees, forming a continuous visible region. The proportion of the visible region to the entire surface area of ​​the 3D model is calculated. The proportion threshold is set to 1%. Small regions with an area proportion lower than this threshold are determined as invalid visible regions and removed. The final visible region data of the 3D model is output, including the set of visible triangular facets, the spatial range of the visible region, and the area proportion. The visibility of the 3D model under the current view is calculated.

[0144] In this embodiment, the cloud content library is equipped with corresponding augmented reality 3D model resources for each scenic spot in the mountain area, such as 3D restoration models of ancient temples, virtual annotation models of mountain landscapes, and virtual image models of historical figures. All models adopt the lightweight glTF format, and after model simplification and texture compression, they are suitable for real-time rendering. The virtual-real fusion rendering unit first determines the current virtual viewpoint and position of the tourist based on the six-DOF pose parameters output by the pose calculation unit, and sends an augmented reality 3D model retrieval request to the cloud content library to retrieve the model resources corresponding to that viewpoint. If the model resources have been preloaded by the viewpoint prediction and preloading submodule of the tourist's end, they are directly read from the local memory to improve the speed of model retrieval. Then, a depth testing algorithm based on occlusion detection is used to calculate the visibility of the augmented reality 3D model in the current tourist's viewpoint. The depth testing algorithm obtains the depth information of each pixel in the real scene by generating a depth map of the current video frame of the live stream, and then the augmented reality 3D model is used. The depth values ​​of each vertex in the augmented reality model are compared with the corresponding pixel depth values ​​in the real-world depth map to determine whether the 3D model is occluded by objects in the real-world scene, such as whether the model of an ancient temple is occluded by trees in the real-world scene. For occluded parts of the model, they are made invisible to ensure the realism of the virtual-real fusion. Next, through perspective projection transformation, the 3D coordinates of the augmented reality 3D model are converted into 2D pixel coordinates of the live video frame, achieving accurate model registration. The parameters of the perspective projection transformation are adjusted in real time according to the six degrees of freedom pose parameters of the tourists and the intrinsic parameters of the head-mounted display device to ensure that the registered position of the model is consistent with the spatial position of the real-world scene, without offset or misalignment. Finally, the registered augmented reality 3D model is merged and rendered with the live video stream, and the brightness and contrast of the model are adjusted to make it consistent with the visual effect of the real-world scene. Finally, an augmented reality output stream with spatial consistency is generated, which is transmitted to the head-mounted display device at the tourist's end to present the tourist with an immersive guided tour screen that blends virtual and real elements.

[0145] Furthermore, the augmented reality module also includes a light field consistency rendering submodule. In the mountainous scenic area of ​​this embodiment, lighting conditions change constantly with time, weather, and terrain, such as strong sunlight on sunny days, soft light on cloudy days, sidelight in the early morning and evening, and shadows formed by mountain obstruction. If the lighting effect of the augmented reality 3D model is inconsistent with the lighting conditions of the real scene, it will lead to a sense of disharmony in the fusion of virtual and reality, reducing the immersive experience. The light field consistency rendering submodule can achieve matching of the lighting effects of augmented reality content with the real scene, improving the realism of the fusion of virtual and reality. The specific implementation process is as follows:

[0146] (1) Before superimposing the augmented reality 3D model onto the live stream, the light field consistency rendering submodule acquires the current environment's lighting probe data through the panoramic acquisition module, extracts the intensity, color, and main direction features of the ambient light, dynamically adjusts the material shader parameters of the 3D model based on these lighting features, and uses image-based lighting technology to generate specular reflection and environmental occlusion effects consistent with the lighting of the real environment for the augmented reality 3D model, so that the light and shadow performance of the augmented reality content matches the lighting conditions in the live stream; In this embodiment, each panoramic camera unit of the panoramic acquisition module has a built-in lighting sensor, which can collect the current environment's lighting probe data in the scenic area in real time. The lighting probe data is stored in the form of a cube map, covering 360° of ambient lighting information, and the acquisition frequency is 10 frames / second, which can reflect the changes in the lighting conditions of the scenic area in real time;

[0147] (2) The light field consistency rendering submodule first extracts three core lighting features from the lighting probe data: ambient light intensity, color, and principal direction. Ambient light intensity is obtained by calculating the mean pixel brightness of the lighting probe data, ambient light color is obtained by calculating the mean pixel color of the lighting probe data, and the principal direction of ambient light is obtained by performing principal component analysis on the lighting probe data, which can accurately determine the main incident direction of the light. Then, based on the extracted lighting features, the material shader parameters of the augmented reality 3D model are dynamically adjusted. The material shader adopts a physically based rendering shader, which supports the adjustment of various lighting effects such as diffuse reflection, specular reflection, and ambient light occlusion. The specific adjustment strategy is as follows: the ambient light intensity parameter of the shader is adjusted according to the ambient light intensity; the stronger the light, the higher the ambient light intensity parameter. The ring light parameter of the shader is adjusted according to the ambient light color. The ambient light color parameters are adjusted to ensure the model's ambient light color matches the real scene. The parallel light direction and intensity parameters of the shader are adjusted according to the main direction of the ambient light to simulate direct light effects in the real scene. Finally, image-based lighting technology is used to map lighting probe data onto the surface of the augmented reality 3D model, generating specular reflections and ambient occlusion effects consistent with the real-world ambient lighting. Specular reflections simulate the reflection of real-world light on the model's surface, such as the reflection of sunlight on the wooden surface of an ancient temple. Ambient occlusion simulates the light and shadow occlusion between the 3D model and real-world objects, such as the shadows cast by trees in the real scene. Through these processes, the lighting and shadow representation of the augmented reality 3D model perfectly matches the lighting conditions in the live stream, significantly enhancing the realism of the virtual-real fusion. Viewers cannot distinguish the differences in lighting and shadow between the virtual model and the real-world scene.

[0148] The augmented reality module also includes a dynamic occlusion processing unit;

[0149] F1: Before overlaying augmented reality content onto the live stream, the dynamic occlusion processing unit performs monocular depth estimation on the current frame image of the live stream through a depth estimation network to generate a depth map with the same resolution as the image. The depth estimation network aims to infer the distance information of objects in the scene from the image. In this embodiment, the depth estimation network is the prior art in this field and is not an inventive solution of this application, so it will not be described in detail here.

[0150] In this embodiment, the depth estimation network adopts a monocular depth estimation model based on convolutional neural networks. The monocular depth estimation model adopts an Encoder-Decoder architecture with a ResNet101 backbone network. It has been trained on massive outdoor scene depth image data and can accurately achieve depth estimation of monocular images, adapting to complex outdoor scenes in mountainous scenic areas. The input of the monocular depth estimation model is the current frame of the live stream's 8K panoramic image, and the output is an 8K depth map with the same resolution as the image. The gray value of each pixel in the depth map corresponds to the depth information of that location in the real scene. The higher the gray value, the greater the depth. The depth estimation error does not exceed 0.5 meters, which can accurately reflect the depth distribution of the real scene. The depth estimation processing speed is 30 frames / second, which is consistent with the frame rate of the live stream, enabling dynamic depth map generation.

[0151] F2: Calculate the depth value of each pixel of the augmented reality 3D model at the current viewpoint based on the depth map, and compare it with the depth value of the real scene image. When the depth value of the augmented reality 3D model is greater than the depth value of the real scene image, it is determined that the augmented reality 3D model is occluded by a real object. The area occluded by the real object is made transparent or only the outline of the occluded part is rendered.

[0152] In this embodiment, the dynamic occlusion processing unit first converts the three-dimensional coordinates of the augmented reality 3D model into two-dimensional pixel coordinates of the current frame image of the live stream based on the six-DOF pose parameters output by the pose calculation unit, and calculates the depth value of each pixel of the 3D model at these pixel coordinates to generate a depth mask for the 3D model. Then, the depth mask of the 3D model is compared pixel by pixel with the depth map of the real scene generated by the depth estimation network. When the depth value of any pixel of the 3D model is greater than the depth value of the corresponding pixel in the real scene depth map, it is determined that the part of the 3D model corresponding to that pixel is occluded by a real object in the real scene. For the occluded model area, two processing methods are provided, which can be selected according to the visitor's settings: one is to make it transparent. The processing involves two methods: adjusting the transparency of the occluded area of ​​the model to 80%, making it semi-transparent so that the model's outline is visible without obscuring the real-world view; and rendering only the outline of the occluded part using white outlines to highlight the model's outline without affecting the viewing of the real-world view. The dynamic occlusion processing is done in real time with a processing delay of no more than 30 milliseconds, accurately tracking dynamic occlusion changes in the real-world view. For example, when a tourist walks past the augmented reality 3D model of the ancient temple, the part of the 3D model occluded by the tourist will immediately be made transparent or have its outline rendered. After the tourist passes, the model will resume normal rendering, ensuring that the blend of virtual and real images is always realistic, smooth, and without visual incongruity.

[0153] The self-evolutionary optimization module collects real-time interactive behavior data and feedback data from tourists, and based on this, jointly optimizes the response strategy model of the interaction engine and the content recommendation model of the cloud content library.

[0154] The self-evolutionary optimization module includes a data acquisition unit and an offline training unit;

[0155] G1: The data acquisition unit records the interactive behavior data of tourists in real time during the interaction process, and after cleaning and normalizing the interactive behavior data, it constructs a structured sample data containing context state, tourist actions and feedback scores; the interactive behavior data includes voice request text, click browsing behavior, dwell time and skip behavior data for various types of reply content.

[0156] In this embodiment, the data acquisition unit collects all interactive behavior data of tourists during the remote guided tour in real time through the interfaces of various modules of the system. The scope of collection covers all modules related to tourist interaction, such as the intelligent interaction module, voice recognition module, tourist terminal, and cloud content library. The collected interactive behavior data is rich in types, specifically including: voice request text, i.e., all query instructions made by tourists through voice; click browsing behavior, i.e., all click, swipe, selection and other operations performed by tourists on the tourist terminal; dwell time on reply content, i.e., the dwell time of tourists listening to various types of chat voice and explanation content output by the system; skip behavior data, i.e., whether tourists skipped the reply content output by the system and the time of skipping; it also includes auxiliary data such as tourist location information, interaction time, and emotional state; the data collection frequency is synchronized with the tourist's interactive behavior to ensure the real-time and completeness of the data; the collected raw interactive behavior data may contain some invalid data, missing data, and abnormal data, such as due to network fluctuations. To address issues such as duplicate data collection, invalid clicks due to tourist errors, and incomplete data with missing fields, the data acquisition unit cleans the raw data. This involves deduplication, completion, and outlier removal to eliminate invalid data and ensure data quality. The cleaned, valid data is then normalized, converting data of different dimensions and ranges into a unified range. For example, dwell time is converted to a normalized value of 0-1, and tourist location information is converted to relative coordinates. Finally, the processed dataset is structured into a data set containing contextual states, tourist actions, and feedback ratings. Contextual states include tourist location, interaction time, emotional state, and historical interaction records. Tourist actions include voice requests, clicks, and skipping actions. Feedback ratings are indirectly calculated based on tourist behavior; for example, the longer a tourist stays on a particular response, the higher the rating, and vice versa. The structured data is stored in a unified format.

[0157] G2: The offline training unit uses the deep deterministic policy gradient algorithm to train the response strategy model inside the interaction engine based on the structured sample data accumulated within a preset period. The optimization objective is to maximize the cumulative score of tourists' long-term satisfaction. The network parameters of the response strategy model are updated. The deep deterministic policy gradient algorithm is existing technology in this field and is not an inventive solution of this application. It will not be described in detail here.

[0158] In this embodiment, the preset training cycle of the offline training unit is 24 hours, meaning that offline model training is performed once every morning on the structured sample data accumulated the previous day, ensuring that the model can absorb the latest tourist interaction behavior data in a timely manner. The core training algorithm is the deep deterministic policy gradient algorithm, which is a continuous action space optimization algorithm based on reinforcement learning, suitable for training continuous action space models such as interaction engine response strategy models. The algorithm treats the interaction process between the interaction engine and the tourist as a Markov decision process, where the state is the tourist's context state, the action is the interaction engine's response strategy selection, the reward is the tourist's feedback score, and the optimization objective is to maximize the tourist's long-term cumulative satisfaction score, i.e. This approach enables the interaction engine's response strategy to consistently provide satisfactory feedback to tourists throughout long-term interactions. During training, the deep deterministic policy gradient algorithm stores tourist interaction experiences in an experience replay pool with a capacity of 1 million records, providing access to massive amounts of interaction sample data. Samples are randomly selected from the experience replay pool for model training, avoiding overfitting caused by sample correlation. Through continuous trial and error and learning, the network parameters of the response strategy model are adjusted, allowing the model's output response strategy to achieve higher tourist feedback ratings. After training, the model's response strategy recommendation accuracy can be improved by 5%-10%, enhancing the intelligence and precision of the interaction engine.

[0159] Furthermore, the response policy model constructed by the offline training unit includes a policy network and a value network;

[0160] (1) The strategy network adopts a cascaded structure of multi-layer convolutional neural network and long short-term memory network. It takes the current tourist's interaction history sequence and location information as input, extracts local features through convolutional layers, and then inputs them into the long short-term memory network layer to capture temporal dependencies. Finally, it outputs the selection probability distribution of each candidate response strategy in the preset action space through a fully connected layer. The value network adopts a hybrid structure of fully connected layer and gated recurrent unit. It takes the candidate response strategy output by the strategy network and the tourist's context state as joint input and outputs the long-term value assessment value corresponding to the response strategy.

[0161] In this embodiment, the core function of the strategy network is to select the optimal response strategy from the preset action space based on the real-time interaction status of tourists. The preset action space includes 12 core response strategies in the scenic area tour scenario, namely, explanation of scenic spot knowledge, popularization of historical anecdotes, recommendation of tour routes, interactive Q&A, introduction of ecological environment, interpretation of cultural features, safety reminders, weather information, facility location guidance, answering tourists' questions, emotional interaction response, and personalized content push. Each strategy is further subdivided into 5-8 specific response content templates, forming a two-level action space system of strategy category-content template, ensuring the richness and relevance of response strategies.

[0162] Furthermore, the specific network structure design of the policy network is as follows: The input layer receives a concatenated feature vector, which consists of tourist interaction history sequence features and location information features. The interaction history sequence features are 512-dimensional vectors transformed from the past 5 rounds of interaction records by the embedding layer. The location information features are 128-dimensional vectors of the tourist's current virtual location latitude and longitude and the scenic area functional area identifier after one-hot encoding. After concatenation, a 640-dimensional input feature vector is formed. After the input layer, three one-dimensional convolutional layers are connected. The first convolutional layer has 64 convolutional kernels with a kernel size of 3 and a stride of 1, and uses ReLU activation function to extract local correlation features in the interaction sequence, such as continuous question-type command features and preference-type click features. The second convolutional layer has 128 convolutional kernels with a kernel size of 3 and a stride of 1, and uses ReLU activation function to further enhance the fusion of local correlation features. The third convolutional layer has 256 convolutional kernels with a kernel size of 2 and a stride of 1, and uses ReLU activation function to reduce the local correlation features. Dimensionality and depth extraction: The feature maps output from the convolutional layers are flattened and input into a two-layer bidirectional long short-term memory network. The first bidirectional long short-term memory network layer has 128 hidden units, and the second bidirectional long short-term memory network layer has 64 hidden units. This structure can effectively capture long-term dependencies in the interaction history sequence. For example, after visiting an ancient temple, tourists are likely to ask questions related to history and culture. The bidirectional structure can utilize both forward and backward sequence information to improve the comprehensiveness of temporal feature capture. The output of the bidirectional long short-term memory network layer is processed by a fully connected layer. The first fully connected layer has 128 neurons and uses ReLU as the activation function. The number of neurons in the second fully connected layer is consistent with the number of policy categories in the preset action space, and the activation function is Softmax. Finally, the probability distribution of 12 response policies is output. The policy category with the highest probability value is the core response policy recommended by the response policy model. Then, specific response content is selected from the content templates under this category, combined with the semantic needs of tourists.

[0163] (2) The value network adopts a dual neural network structure similar to the policy network structure, but takes the action output by the policy network and the current state as input, outputs the expected cumulative reward estimate of the action, updates the value network parameters through backpropagation of time difference error, and updates the policy network parameters through backpropagation of policy gradient.

[0164] Furthermore, the specific network structure of the value network is as follows: The input layer receives a joint feature vector, which is composed of the candidate response strategy features output by the policy network and the tourist context state features, totaling 268 dimensions; after the input layer, two fully connected layers are connected. The first fully connected layer has 128 neurons and uses ReLU activation function, while the second fully connected layer has 64 neurons and uses ReLU activation function to achieve linear transformation and nonlinear fusion of features; the output of the fully connected layer is input to a gated recurrent unit layer with 32 hidden units. This structure can effectively handle the dynamic correlation between the response strategy and the tourist context state, and capture the temporal features of the behavioral feedback that the tourist may produce after the strategy is executed; the output of the gated recurrent unit layer passes through the last fully connected layer to output the long-term value evaluation value of the candidate response strategy. This value is a continuous real number. The higher the value, the higher the long-term value of the strategy, and the more it can improve the tourist's long-term satisfaction.

[0165] Furthermore, during model training, the policy network first outputs candidate response strategies based on the contextual states of tourists in the training samples; then, the value network evaluates the value of the strategy to obtain the target value; next, the action loss of the policy network and the mean squared error loss of the value network are calculated. The action loss is calculated based on the target value and the probability distribution of the strategy, and the value loss is the mean squared error between the target value and the value network's predicted value; finally, the parameters of the policy network and the value network are updated respectively using the gradient descent algorithm, with the learning rate set to 0.001, the batch size to 256, the training rounds to 100, and validation performed every 10 rounds. When the cumulative tourist satisfaction score on the validation set no longer increases, training is stopped, and the optimal model parameters are saved.

[0166] The self-evolutionary optimization module also includes an online update unit;

[0167] H1: The online update unit loads the response strategy model parameters updated by the offline training unit. During real-time interaction, it calculates the selection probability of each candidate response strategy based on the current real-time interaction state of the tourist through forward propagation. It adopts an exploration strategy based on Boltzmann distribution for action sampling. The calculation process of forward propagation and Boltzmann distribution is the prior art in this field and is not an inventive solution of this application. It will not be described in detail here.

[0168] Furthermore, the Boltzmann distribution-based exploration strategy parameter configuration includes: the core parameter is the temperature parameter, which is used to balance the exploratory and exploitative aspects of the strategy. The initial value is set to 0.8 and dynamically adjusted with the number of interactions: the temperature parameter decreases by 0.05 for every 100 visitor interactions, with a minimum value set to 0.1, ensuring that as interaction data accumulates, candidate response strategies gradually shift from exploration to exploitation. The precision of the exponential operation during calculation is set to 0.0001 to avoid numerical overflow. At the same time, the minimum probability threshold for sampling candidate response strategies is set to 0.001. If the probability value of any candidate response strategy after Boltzmann distribution transformation is lower than the minimum probability threshold, it is removed from the sampling pool. The configured Boltzmann exploration strategy parameters and the filtered candidate response strategy sampling pool are output.

[0169] Furthermore, the Boltzmann distribution sampling process includes: first, substituting the original probability values ​​of candidate response strategies into the Boltzmann distribution formula to calculate the sampling probability of each candidate response strategy; constructing a probability distribution array based on the sampling probabilities, and using a roulette wheel method for random sampling, with the random number seed generated based on the system timestamp to ensure sampling randomness; prioritizing the selection of strategies with the top 10% sampling probabilities during sampling, if the same strategy is obtained in three consecutive samplings, the temperature parameter is readjusted, temporarily increased by 0.1, and sampling is repeated to avoid strategy homogenization. After sampling is completed, the selected response strategy identifier, such as the strategy number and strategy text content, is output, along with the sampling probability of the strategy, the current value of the temperature parameter, and the sampling timestamp, for strategy effectiveness evaluation.

[0170] H2: Simultaneously, the online update unit collects real-time feedback data of the current interaction process, uses an online learning algorithm based on a priority experience replay mechanism to fine-tune the value network of the response strategy model in real time, and uses the fine-tuned value network to correct the output of the strategy network of the response strategy model in real time.

[0171] The self-evolutionary optimization module also jointly optimizes the content recommendation model of the cloud content library; the content recommendation model adopts an architecture based on deep factorization machine, which includes an embedding layer, a factorization machine layer, a deep neural network layer and a logistic regression layer.

[0172] K1: The embedding layer maps the discrete and continuous features of tourists into embedding vectors, where discrete features include the sequence of historical visited attractions and interaction preference types, and continuous features include the current location and duration of stay.

[0173] K2: The factorization machine layer performs second-order feature cross modeling on the embedded vector and outputs cross feature representation. The second-order feature cross is existing technology in this field and is not an inventive solution of this application. It will not be described in detail here.

[0174] Furthermore, the second-order interaction feature computation involves mapping the embedded vectors to an 8-dimensional latent vector space and calculating the interaction term between any two features using the second-order interaction formula of the factorization machine.

[0175] K3: The deep neural network layer takes the result of embedding vector concatenation as input, extracts high-order nonlinear features through a multi-layer fully connected network, and finally concatenates the outputs of the factorization machine layer and the deep neural network layer. Then, it calculates the recommendation score of each candidate content through the logistic regression layer, sorts the candidate content according to the recommendation score, and responds to the visitor's query request.

[0176] Furthermore, when extracting high-order nonlinear features, the embedded vector concatenation result is used as input, and a four-layer fully connected network is used to extract high-order nonlinear features. The input layer has 673,152 neurons, the first hidden layer has 2,048 neurons, the second hidden layer has 1,024 neurons, the third hidden layer has 512 neurons, and the output layer has 128 neurons. All hidden layers use the ReLU activation function with a negative interval slope of 0.01. After each fully connected layer operation, batch normalization and random dropout processing are performed sequentially. The moving average decay rate of batch normalization is set to 0.99, the variance decay rate is set to 0.999, and the dropout rate of random dropout is set to 0.2. The embedded vector concatenation result is fed into each hidden layer through the input layer, and activation, batch normalization, and dropout processing are performed sequentially. Finally, a 128-dimensional high-order nonlinear feature vector in the range of 0 to 1 is output.

[0177] Furthermore, the calculation process for the recommendation score of the logistic regression layer is as follows: the 1-dimensional low-order feature interaction score output by the factorization machine layer is concatenated with the 128-dimensional high-order feature vector output by the deep neural network layer to obtain a 129-dimensional fused feature vector; the fused feature vector is input into the logistic regression layer, with a weight dimension of 129×1, initialized using Xavier, and a bias term of 0.05. After linear transformation, the original score is obtained, and then converted into a recommendation score in the range of 0 to 1 using the Sigmoid function; recommendation scores below 0.01 are set to 0.01, and those above 0.99 are set to 0.99, and the calibrated recommendation score for each candidate content is output.

[0178] Furthermore, the self-evolutionary optimization module also includes a federated learning submodule. The federated learning submodule divides the model update process based on the deep deterministic policy gradient algorithm in the offline training unit into two stages: local training and global aggregation. On each tourist terminal, the model parameter gradient is calculated based on the locally stored interactive behavior data. Only the encrypted gradient parameters are uploaded to the cloud server. The cloud server uses a federated averaging algorithm to aggregate the gradient parameters uploaded by multiple tourist terminals to update the global model. The updated global model parameters are then distributed to each tourist terminal, thereby achieving joint optimization of the response strategy model without collecting the original data.

[0179] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments under the guidance of the present invention without departing from the spirit and scope of the present invention. All of these variations are within the protection scope of the present invention.

Claims

1. A remote guided live streaming system based on panoramic video, characterized in that, include: Panoramic data acquisition module, visitor terminal, intelligent interaction module, voice recognition module, augmented reality module, self-evolutionary optimization module; The panoramic acquisition module collects video and audio data from each panoramic camera array, generates a panoramic video stream and a tour guide audio stream, and then merges and outputs a live stream with tour guide voiceover. The tourist terminal receives the live stream and obtains location information through its built-in positioning module. Based on the location information, it matches and retrieves the explanation content of the current attraction from the cloud content library and plays it in the live stream. During the periods when there is no tour guide audio in the live stream, the intelligent interaction module activates the interaction engine, retrieves preset content from the cloud content library to generate chat voice, and outputs it to the tourist terminal in sync with the live stream. During the interaction, the speech recognition module collects the tourist's voice commands, parses them, retrieves the corresponding response content from the pre-stored cloud content library, and generates chat voice after speech synthesis. Based on the position change identified by the positioning module, the augmented reality module retrieves the corresponding augmented reality content from the cloud content library and overlays it onto the live stream using virtual reality technology. The self-evolutionary optimization module collects real-time data on tourists' interactive behaviors and feedback, and jointly optimizes the response strategy model of the interaction engine and the content recommendation model of the cloud content library.

2. The remote guided live streaming system based on panoramic video as described in claim 1, characterized in that, The panoramic acquisition module includes multiple array-distributed panoramic camera units and a video stitching server connected to the multiple panoramic camera units. The video stitching server extracts feature points from the raw video data output by each panoramic camera unit based on the scale-invariant feature transform algorithm, and eliminates mismatches of the extracted feature points through the random sampling consistency algorithm. Then, it calculates the homography matrix of the images acquired by adjacent panoramic camera units according to the spatial position relationship of the feature points, and uses the homography matrix to project multiple raw video data onto a preset equidistant cylindrical projection model to complete the spatial alignment and real-time fusion of the images and generate a panoramic video stream. The panoramic acquisition module also includes an audio processing unit connected to the video stitching server; The audio processing unit uses a beamforming algorithm to perform sound source localization and enhancement processing on the multi-channel audio data collected by the microphone array built into each panoramic camera unit, extracts the tour guide audio stream, and then timestamps and encapsulates the tour guide audio stream and panoramic video stream to output a live stream with tour guide voice.

3. The remote guided live streaming system based on panoramic video as described in claim 2, characterized in that, The visitor terminal includes a head-mounted display device and a mobile computing unit connected to the head-mounted display device. The mobile computing unit has a built-in positioning module, which obtains the latitude and longitude coordinates of the visitor's location based on a global navigation satellite system receiver, and integrates the three-axis acceleration and angular velocity data collected by the inertial measurement unit inside the head-mounted display device. The module then performs real-time correction and smoothing of the visitor's location using an extended Kalman filter algorithm to generate visitor location information. Based on the visitor location information, the mobile computing unit uses a hash index-based data query method to match and retrieve the current attraction's explanation content from the cloud content library, decodes the explanation content, and overlays it onto the corresponding timeline of the live stream for playback.

4. The remote guided live streaming system based on panoramic video as described in claim 3, characterized in that, The intelligent interaction module includes an interaction status monitoring unit and a dialogue management unit; The interactive status monitoring unit analyzes the energy value of the audio track of the live stream in real time, and determines whether there is tour guide audio data in the current time period through short-time energy analysis and zero-crossing rate detection. When no tour guide audio data is detected within a continuous preset time threshold, an interactive trigger signal is generated. After receiving the interaction trigger signal, the dialogue management unit starts the interaction engine. Based on the tourist's current location information, it selects preset interactive content that matches the tour scenario from the cloud content library. It then uses a time-domain interpolation algorithm to perform multi-track audio mixing processing on the chat voice generated by the preset interactive content and the live stream. After the volume ratio of the chat voice and the background ambient sound meets the preset loudness standard, it outputs the audio to the tourist's end synchronously.

5. The remote guided live streaming system based on panoramic video as described in claim 4, characterized in that, The speech recognition module includes a front-end speech processing unit and a back-end semantic understanding unit; The front-end voice processing unit collects tourists' voice commands through the microphone array at the tourist end, applies the least mean square adaptive filtering algorithm to suppress the collected environmental noise, and then extracts the Mel frequency cepstral coefficients as acoustic features. The backend semantic understanding unit inputs the extracted Mel frequency cepstral coefficients into the end-to-end speech recognition model trained based on connection time-series classification, decodes and outputs the corresponding text instructions, and uses a named entity recognition algorithm based on bidirectional long short-term memory network and conditional random field to extract key information including the name of the scenic spot and the operation intention from the text instructions. Based on the key information, a query vector is constructed, and matching response content is retrieved from the cloud content library.

6. The remote guided live streaming system based on panoramic video as described in claim 5, characterized in that, The augmented reality module includes a pose calculation unit and a virtual-real fusion rendering unit; The pose calculation unit calculates the six-degree-of-freedom pose parameters of the tourist in three-dimensional space in real time using a visual simultaneous localization and mapping algorithm, based on the tourist's position information output by the positioning module and the posture data collected by the inertial measurement unit inside the head-mounted display device. The virtual-real fusion rendering unit, based on six degrees of freedom pose parameters, calls the corresponding augmented reality 3D model resources from the cloud content library, uses a depth testing algorithm based on occlusion detection to calculate the visibility of the augmented reality 3D model from the current viewpoint, and registers and superimposes the augmented reality 3D model onto the corresponding pixel area of ​​the live stream through perspective projection transformation, generating an augmented reality output stream with spatial consistency.

7. The remote guided live streaming system based on panoramic video as described in claim 6, characterized in that, The augmented reality module also includes a dynamic occlusion processing unit; Before overlaying augmented reality content onto the live stream, the dynamic occlusion processing unit performs monocular depth estimation on the current frame image of the live stream through a depth estimation network to generate a depth map with the same resolution as the image. The depth value of each pixel of the augmented reality 3D model is calculated based on the depth map at the current viewpoint and compared with the depth value of the real scene image. When the depth value of the augmented reality 3D model is greater than the depth value of the real scene image, it is determined that the augmented reality 3D model is occluded by a real object. The area occluded by the real object is made transparent or only the outline of the occluded part is rendered.

8. The remote guided live streaming system based on panoramic video as described in claim 7, characterized in that, The self-evolutionary optimization module includes a data acquisition unit and an offline training unit; The data acquisition unit records tourists' interactive behavior data in real time during the interaction process, and after cleaning and normalizing the interactive behavior data, it constructs structured sample data containing context state, tourist actions and feedback scores; the interactive behavior data includes voice request text, click browsing behavior, dwell time and skip behavior data for various types of response content. The offline training unit uses a deep deterministic policy gradient algorithm to train the response strategy model inside the interaction engine based on the structured sample data accumulated within a preset period, with the optimization objective of maximizing the cumulative satisfaction score, and updates the network parameters of the response strategy model.

9. The remote guided live streaming system based on panoramic video as described in claim 8, characterized in that, The self-evolutionary optimization module also includes an online update unit; The online update unit loads the response strategy model parameters updated by the offline training unit. During real-time interaction, it calculates the selection probability of each candidate response strategy based on the current real-time interaction state of the tourist through forward propagation, and uses an exploration strategy based on Boltzmann distribution for action sampling. Meanwhile, the online update unit collects feedback data from the current interaction process, uses an online learning algorithm based on a priority experience replay mechanism to fine-tune the value network of the response strategy model in real time, and uses the fine-tuned value network to correct the output of the strategy network of the response strategy model in real time.

10. The remote guided live streaming system based on panoramic video as described in claim 9, characterized in that, The self-evolutionary optimization module also jointly optimizes the content recommendation model of the cloud content library; the content recommendation model adopts an architecture based on deep factorization machine, which includes an embedding layer, a factorization machine layer, a deep neural network layer and a logistic regression layer. The embedding layer maps tourists' discrete and continuous features into embedding vectors, where discrete features include historical visit sequence and interaction preference type, and continuous features include current location and duration of stay. The factorization machine layer performs second-order feature cross modeling on the embedded vectors and outputs cross feature representations. The deep neural network layer takes the result of embedding vector concatenation as input, extracts nonlinear features through a multi-layer fully connected network, and finally concatenates the outputs of the factorization machine layer and the deep neural network layer. Then, it calculates the recommendation score of each candidate content through a logistic regression layer, sorts the candidate content according to the recommendation score, and responds to the visitor's query request.