A scene multimedia immersion interaction system based on AR technology

CN122530503APending Publication Date: 2026-08-07SUZHOU GOLD MANTIS EXHIBITION DESIGN ENG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU GOLD MANTIS EXHIBITION DESIGN ENG
Filing Date
2026-06-03
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

然而,现有AR系统在复杂物理环境(如强光、多障碍物、动态场景)中普遍存在虚实融合精度不足、环境适应性差、交互方式单一且不自然、内容专业化程度低等问题

Benefits of technology

[0030]本发明的交互系统通过空间定位感知模块中视觉惯性里程计与深度传感器的融合定位,结合语义分割的环境认知,实现了厘米级高精度位姿估计,且能够根据地面、水面、建筑等场景要素约束虚拟模型的放置与碰撞,显著提升了虚实融合的精准度;显示渲染模块利用环境光传感器、光照估计模型、多层感知器和PBR渲染管线,自适应调节虚拟模型在复杂光照下的亮度与色彩,配合模型剪枝与GPU加速,保证了在强光、动态环境下的显示稳定性与流畅度(帧率≥60fps);多模态自然交互模块融合手势识别与语音交互,并通过时间戳同步与状态机仲裁,避免了指令冲突,实现了自然、低延迟(手势<20ms,语音<500ms)的人机交互;虚拟内容构建优化模块采用渐进式网格与姿态预测,兼顾高精度模型表现与终端性能适配,减少定位延迟造成的拖影或偏差;智能功能模块基于终端侧AI框架实现智能问答、实时翻译与个性化推荐,无需依赖网络,保护隐私且响应迅速,极大地丰富了AR场景的信息服务能力;各模块之间信号传递路径清晰、协同关系明确,形成了从环境感知、交互输入到内容生成与渲染显示的完整闭环,系统整体可靠性高、可扩展性强。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122530503A_ABST
    Figure CN122530503A_ABST
Patent Text Reader

Abstract

The application discloses a scene multimedia immersion interaction system based on AR technology. The system comprises: a spatial positioning and sensing module, which calculates the device pose in real time and performs semantic segmentation on scene elements; a display rendering module, which adaptively adjusts the virtual model light and transparency; a multi-modal natural interaction module, which fuses gestures and voice instructions; a virtual content construction and optimization module, which loads and dynamically optimizes three-dimensional models; an intelligent function module, which runs voice recognition, translation, and question and answer models, and provides information superposition and personalized recommendation in combination with a knowledge base. Spatial positioning information is distributed to the display rendering, interaction, and model optimization modules; interaction control signals trigger the intelligent function module; intelligent output and optimized models are sent to the display rendering module, and finally an augmented reality picture is generated. The application realizes centimeter-level positioning, stable display under complex lighting, low-delay multi-modal interaction, and terminal-side intelligent services, and significantly improves the virtual-real fusion quality and scene adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interactive system technology, specifically to a scene multimedia immersive interactive system based on AR technology. Background Technology

[0002] In recent years, AR (Augmented Reality) technology has developed rapidly and has been widely applied in many fields such as gaming, education, healthcare, and cultural heritage display. Augmented reality technology itself is an interdisciplinary research field that integrates various technologies from different research areas, such as virtual reality, computer vision, artificial intelligence, wearable mobile computing, human-computer interaction, and bioengineering. Augmented reality technology can overlay computer-generated virtual information onto the real physical environment observed by the user in real time, thereby achieving an "enhanced" perception of the real world.

[0003] Abroad, research and application of AR technology started earlier, and certain achievements have been made in areas such as the accuracy of virtual-reality fusion and the naturalness of interaction. For example, in the field of cultural heritage display, some countries use AR technology to recreate historical scenes, allowing tourists to immerse themselves in historical culture. However, there are relatively few specialized studies and mature application cases in multimedia immersive interactive systems for specific scenarios such as canal shipping.

[0004] The development of AR technology in China has been rapid, with active exploration in multiple fields. However, existing AR systems generally suffer from problems such as insufficient accuracy in virtual-real fusion, poor environmental adaptability, simplistic and unnatural interaction methods, and low level of content specialization in complex physical environments (such as strong light, multiple obstacles, and dynamic scenes). Specifically, visual positioning errors are often 10-15cm, which cannot meet the requirements for accurate overlay of fine models; virtual models appear dim, shaky, or flickering under strong light; interaction is mainly based on touch screens or simple gestures, lacking natural interaction logic adapted to specific scenarios (such as water transportation and historical scenes); and content presentation is mostly superficial, lacking in-depth knowledge integration and intelligent interaction functions. Summary of the Invention

[0005] This invention provides a scene-based multimedia immersive interactive system based on AR technology. It aims to achieve high-precision spatial positioning, adaptive rendering, multimodal natural interaction, and intelligent content services in complex environments through the collaborative work of multiple modules, thereby significantly improving the virtual-real fusion quality, environmental adaptability, and user experience of the AR system.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a scene multimedia immersive interactive system based on AR technology, comprising:

[0007] The spatial positioning and perception module is configured to calculate the pose of the device in the physical space in real time, and to perform semantic segmentation on scene elements in the physical environment to generate environmental cognitive information.

[0008] The display rendering module is configured to collect real-time ambient lighting parameters and dynamically adjust the lighting and transparency of the virtual model to perform brightness and color adaptation processing between the virtual content and the real scene.

[0009] A multimodal natural interaction module is configured to recognize user interaction commands and merge the interaction commands from different modalities into a unified interaction control signal;

[0010] The virtual content construction optimization module is configured to load 3D models.

[0011] The intelligent function module is configured to run a functional model and, in conjunction with a knowledge base, provide users with information overlay display and personalized recommendations.

[0012] Specifically, the spatial positioning and perception module sends the acquired pose information and environmental awareness information to the display rendering module, the multimodal natural interaction module, and the virtual content construction and optimization module. The multimodal natural interaction module sends the interaction control signal to the intelligent function module. The intelligent function module sends the generated overlaid text information and personalized recommendation content to the display rendering module. The virtual content construction and optimization module sends the optimized high-precision 3D model to the display rendering module. The display rendering module receives the pose information, environmental awareness information, overlaid text information, personalized recommendation content, and high-precision 3D model, and renders and generates the final augmented reality image according to the interaction control signal.

[0013] As a further description of the above technical solution:

[0014] The environmental perception information sent by the spatial positioning and perception module to the virtual content construction and optimization module is used to guide model simplification or detail enhancement strategies; the pose information sent by the spatial positioning and perception module to the multimodal natural interaction module is used to assist in determining the spatial context of user interaction commands.

[0015] As a further description of the above technical solution:

[0016] The spatial positioning and perception module calculates pose by fusing visual inertial odometry and depth sensor, with the fusion method being either tightly coupled or loosely coupled; the depth sensor is a LiDAR or structured light sensor.

[0017] As a further description of the above technical solution:

[0018] The semantic segmentation uses a deep learning model to identify the spatial location of each physical object in the scene, thereby constraining the placement of the virtual model and its physical collision response.

[0019] As a further description of the above technical solution:

[0020] The display rendering module includes: a lighting estimation model configured to obtain high dynamic range lighting parameters of the real environment based on illuminance data collected in real time by an ambient light sensor; a multilayer perceptron configured to optimize the lighting response of the virtual model through an MLPs algorithm; and a PBR rendering pipeline configured to perform lighting adaptation processing between the virtual model and the real environment based on a brightness adaptation model.

[0021] As a further description of the above technical solution:

[0022] The multimodal natural interaction module includes: a gesture recognition unit configured to separate the hand region from the background through semantic segmentation and obtain the gesture command in the interaction command; and a voice interaction unit configured to support multilingual voice input and synthesis to obtain the voice command in the interaction command.

[0023] As a further description of the above technical solution:

[0024] The multimodal natural interaction module further includes: a timestamp synchronization unit configured to unify the time base of data from each modality; and an interaction state machine configured to analyze the logical consistency of the interaction instructions from different modalities.

[0025] As a further description of the above technical solution:

[0026] The virtual content construction and optimization module constructs a 3D model using 3D scanning technology and dynamically adjusts the model accuracy based on the performance of the terminal device using progressive mesh units, simplifying the mesh for non-critical parts. The virtual content construction and optimization module also includes a pose prediction module, which is configured to predict the model pose of the next frame based on historical pose data.

[0027] As a further description of the above technical solution:

[0028] The functional model includes a speech recognition model, a natural language processing model, and a machine translation model deployed based on a terminal-side artificial intelligence framework; the knowledge base supports automatic collection via web crawlers and updates after manual review; the intelligent functional modules include a real-time intelligent question answering unit and a real-time translation unit. The real-time intelligent question answering unit is configured to retrieve answers from the knowledge base and display them overlaid in augmented reality text format. The real-time translation unit is configured to support real-time translation of voice or text input, and the translation results are displayed in the augmented reality scene in text format with adjustable parameters.

[0029] In summary, due to the adoption of the above technical solution, the present invention has the following beneficial effects compared with the prior art:

[0030] The interactive system of this invention achieves centimeter-level high-precision pose estimation through the fusion positioning of visual inertial odometry and depth sensors in the spatial positioning perception module, combined with semantic segmentation for environmental cognition. It can also constrain the placement and collision of virtual models based on scene elements such as ground, water, and buildings, significantly improving the accuracy of virtual-real fusion. The display rendering module utilizes an ambient light sensor, a lighting estimation model, a multilayer perceptron, and a PBR rendering pipeline to adaptively adjust the brightness and color of the virtual model under complex lighting conditions. Combined with model pruning and GPU acceleration, it ensures display stability and smoothness (frame rate ≥ 60fps) in strong light and dynamic environments. The multimodal natural interaction module integrates gesture recognition and voice interaction, and uses time... The system employs a combination of synchronization and state machine arbitration to avoid command conflicts, enabling natural, low-latency (gesture <20ms, voice <500ms) human-computer interaction. The virtual content construction and optimization module utilizes progressive mesh and pose prediction, balancing high-precision model performance with terminal performance adaptation, reducing ghosting or deviations caused by positioning latency. The intelligent function module, based on a terminal-side AI framework, enables intelligent question answering, real-time translation, and personalized recommendations without relying on the network, protecting privacy and responding rapidly, greatly enriching the information service capabilities of AR scenarios. The signal transmission paths between modules are clear, and the collaborative relationships are well-defined, forming a complete closed loop from environmental perception and interactive input to content generation and rendering display, resulting in high overall system reliability and strong scalability. Attached Figure Description

[0031] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is a system module block diagram of a scene multimedia immersive interactive system based on AR technology.

[0033] Figure 2 This is a system block diagram of the spatial positioning and perception module in a scene multimedia immersive interactive system based on AR technology.

[0034] Figure 3 This is a system module block diagram of the display and rendering module in a scene multimedia immersive interactive system based on AR technology.

[0035] Figure 4This is a system module block diagram of a multimodal natural interaction module in a scene multimedia immersive interactive system based on AR technology.

[0036] Figure 5 This is a system module block diagram of a virtual content construction and optimization module in an AR-based scene multimedia immersive interactive system.

[0037] Figure 6 This is a system module block diagram of an intelligent functional module in a scene multimedia immersive interactive system based on AR technology. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0039] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0040] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0041] In the description of the embodiments of the present invention, it should be noted that the terms "upper" and "inner" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of the invention is usually placed when in use. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting the present invention.

[0042] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection, an indirect connection through an intermediate medium, or a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0043] Please see Figure 1-6 This invention provides a technical solution: a scene multimedia immersive interactive system based on AR technology, which can run on multiple terminal devices such as smartphones, AR glasses, and tablets. The following describes in detail the working principles and collaborative processes of each module after the system starts up, following the workflow after system startup.

[0044] In this embodiment, the system is applied to a shipping scenario, aiming to solve problems such as inaccurate integration of virtual and reality, unnatural interaction, and poor display effects in complex environments that exist in current AR technology applications in shipping scenarios. Through five modules—display technology, interaction technology, sensing technology, computer graphics technology, and intelligent functions—a complete AR shipping scenario multimedia immersive interactive system is constructed. This system presents the history, culture, and development of canal shipping to users in an immersive way, allowing them to experience the charm of canal shipping firsthand, enhancing their understanding and knowledge of canal culture, and promoting its inheritance and dissemination. The system includes:

[0045] Firstly, the spatial positioning and perception module is configured to calculate the pose of the device in the physical space in real time, and to perform semantic segmentation of scene elements in the physical environment to generate environmental cognitive information.

[0046] Specifically, the spatial positioning and perception module calculates pose by fusing visual inertial odometry and a depth sensor, with the fusion method being either tightly coupled or loosely coupled; the depth sensor is a LiDAR or structured light sensor. The semantic segmentation utilizes a deep learning model to identify the spatial positions of various objects in the scene, thereby constraining the placement of the virtual model and its physical collision response. The environmental awareness information sent by the spatial positioning and perception module to the virtual content construction and optimization module guides model simplification or detail enhancement strategies; the pose information sent by the spatial positioning and perception module to the multimodal natural interaction module assists in determining the spatial context of user interaction commands.

[0047] The spatial positioning and sensing module is the underlying foundation of the system. Once the device is powered on, this module acquires real-time images from the camera, data from the inertial measurement unit (IMU), and data from the depth sensor (LiDAR or structured light sensor). Its specific working principle is as follows:

[0048] 1. Pose calculation process:

[0049] Visual Inertial Odometry (VIO): Extracts ORB or SuperPoint feature points from consecutive image frames, combines them with IMU pre-integration results, and performs tightly coupled or loosely coupled pose estimation through nonlinear optimization (such as iSAM2) to output the device's six-degree-of-freedom pose (position and attitude angle) in the local coordinate system.

[0050] Depth sensor fusion: For scenes where pure vision is prone to failure, such as sparsely textured water surfaces and repetitive structures (e.g., dock supports), LiDAR directly provides high-precision depth point clouds. The system jointly optimizes the depth point cloud with visual features, corrects the cumulative drift of VIO, and finally outputs centimeter-level pose information (positioning error <3cm).

[0051] 2. Environmental semantic segmentation process:

[0052] Meanwhile, the module embeds a lightweight deep learning semantic segmentation network (such as DeepLabV3+ or BiSeNet) to perform pixel-level classification on each frame of camera image, identifying key scene elements such as ground, water, sky, buildings, bridges, and obstacles. The network outputs a semantic mask map with category labels, and combines it with depth information to calculate the location range of each element in three-dimensional space, forming "environmental cognitive information".

[0053] 3. Signal transmission:

[0054] This module simultaneously sends the calculated pose information and environmental awareness information to three downstream modules at a frequency of ≥60fps:

[0055] It is sent to the display rendering module for subsequent coordinate transformation and occlusion relationship determination of the virtual model;

[0056] Send it to the multimodal natural interaction module to help determine the spatial context of the target of the user's gesture or voice command (for example, if the user says "zoom in on this ship", the specific virtual object referred to by "this" can be determined by combining the pose).

[0057] Send to the virtual content building and optimization module to guide model simplification strategies (e.g., a low-precision model can be used for distant water areas, while the details of the dock model are preserved for nearby areas).

[0058] Secondly, the multimodal natural interaction module is configured to recognize the user's interaction commands and integrate the interaction commands from different modalities into a unified interaction control signal.

[0059] Specifically, the multimodal natural interaction module includes: a gesture recognition unit configured to separate the hand region from the background through semantic segmentation and obtain the gesture commands in the interaction instructions; a voice interaction unit configured to support multilingual voice input and synthesis, specifically, to continuously receive microphone audio streams to obtain the voice commands in the interaction instructions; a timestamp synchronization unit configured to unify the time base of data from different modalities; and an interaction state machine configured to analyze the logical consistency of the interaction commands from different modalities. Its workflow is as follows:

[0060] 1. Gesture recognition process:

[0061] First, a semantic segmentation network is used to separate the hand region from the camera image and remove background interference;

[0062] Subsequently, skin color detection (YCrCb color space) and contour extraction algorithms were used to obtain the binary mask and key points of the hand (fingertips, center of the palm, etc.).

[0063] A trained lightweight convolutional neural network (such as MobileNetV2+SSD) is used to recognize static gestures (such as open fingers and clenched fist) and dynamic gestures (such as swiping and waving), with a recognition response time of less than 20ms.

[0064] To improve robustness against complex backgrounds, data augmentation (random rotation, scaling, local occlusion, etc.) was employed during the training phase.

[0065] 2. Voice interaction process:

[0066] End-to-end speech recognition models based on recurrent neural networks (RNN) or Transformers (such as Conformer) are used to convert speech into text in real time;

[0067] The text-to-speech (TTS) unit can convert system-feedback text into speech output, supporting multiple languages ​​including Chinese and English;

[0068] The voice response latency (from the end of the user's voice to the recognition of the command) is less than 500ms.

[0069] 3. Multimodal fusion process:

[0070] Gesture commands and voice commands may occur simultaneously; for example, a user might say "move there" while pointing in a direction. In this case, the timestamp synchronization unit within the module adds a unified timestamp (milliseconds) to each gesture recognition result and voice recognition result. The interaction state machine analyzes the commands in both modalities:

[0071] If the two logics are consistent (e.g., gesturing to the boat + saying "rotate"), they are merged into a composite control signal;

[0072] If the two conflict (e.g., gesture "zoom in" + voice "zoom out"), then a unique signal will be output according to the preset priority (e.g., voice priority or user-defined).

[0073] Finally, a unified interactive control signal is generated and sent to the intelligent function module.

[0074] Thirdly, the intelligent function module is configured to run a functional model and, in conjunction with a knowledge base, provide users with information overlay display and personalized recommendations.

[0075] Specifically, the functional model includes a speech recognition model, a natural language processing model, and a machine translation model deployed based on a terminal-side artificial intelligence framework; the knowledge base supports automatic collection via web crawlers and updates after manual review; the intelligent functional module includes a real-time intelligent question answering unit and a real-time translation unit. The real-time intelligent question answering unit is configured to retrieve answers from the knowledge base and display them overlaid in augmented reality text form. The real-time translation unit is configured to support real-time translation of voice or text input, and the translation results are displayed in the augmented reality scene in text form with adjustable parameters.

[0076] Working principle and process of intelligent functional modules:

[0077] 1. Knowledge base and question-and-answer process:

[0078] The intelligent function module pre-builds a knowledge graph and structured knowledge base for specific scenarios (such as the cultural history of a certain body of water). The knowledge base is deployed offline and can also be updated incrementally after being manually reviewed by a configured web crawler periodically scraping relevant content from authoritative websites. When a user asks a question via voice or text (e.g., "What is the history of this ship?"), the multimodal natural interaction module sends the parsed text command to this module as an interactive control signal. The intelligent question answering unit uses BERT or a lightweight QA model (based on the CoreML framework) deployed on the terminal to perform intent recognition and entity extraction on the question, retrieves the most relevant answer from the knowledge base, and then returns it to the display rendering module in augmented reality text format (with a background frame, adjustable font size and color).

[0079] 2. Real-time translation process:

[0080] For scenarios requiring multilingual narration, the real-time translation unit receives voice or text input and invokes a machine translation model on the terminal side (such as a miniaturized Transformer-based model) to translate the source language (e.g., Chinese) into the target language (e.g., English, Japanese). The translation result is also output as an overlay of text, and users can adjust the display position, size, and color of the text in real time to avoid obscuring core scene content. All AI models are accelerated using model quantization (FP16 or INT8) and GPU / NPU to ensure low power consumption and low latency.

[0081] 3. Personalized recommendations:

[0082] Based on the user's historical interaction behavior (such as which virtual content they viewed and which questions they asked), the intelligent function module generates personalized recommended content (such as "you may also be interested in the construction process of this dock") through simple collaborative filtering or tag-based recommendation algorithms, and sends it to the display rendering module.

[0083] Fourth, the virtual content construction and optimization module is configured to load a 3D model. Specifically, this module constructs a 3D model using 3D scanning technology and dynamically adjusts the model's accuracy based on the terminal device's performance using progressive mesh units, simplifying the mesh for non-critical parts. The module also includes a pose prediction module, configured to predict the model's pose for the next frame based on historical pose data. Its working principle is as follows:

[0084] 1. Model loading and progressive meshing:

[0085] This module pre-collects high-precision point cloud data of real objects (such as ancient ships and architectural components) using a 3D scanner, and generates a high-fidelity 3D model through mesh reconstruction and texture mapping. During runtime, the module first loads the base model from storage. Then, based on received environmental awareness information (especially the distance of objects to the user and whether they are in the visual focus area) and the real-time performance of the terminal device (CPU / GPU utilization, available memory), it dynamically adjusts the model's mesh precision.

[0086] For close-up or important objects, use a high-precision mesh (details preserved);

[0087] For distant or non-focused objects, automatically simplify the mesh to reduce the number of vertices and faces;

[0088] The implementation uses a progressive mesh, which allows for quick switching between different LOD levels with smooth transitions.

[0089] 2. Attitude prediction process:

[0090] To reduce "model jitter" or "trailing" caused by transmission and rendering latency between the spatial positioning and perception module and the display rendering module, a pose prediction module is incorporated into the system. This module, based on Kalman filtering or a lightweight LSTM network, records pose data (position and rotational angular velocity) from several past frames and predicts the model's pose for the next frame (approximately 16ms later). The predicted pose is then sent along with the 3D model to the display rendering module, thereby compensating for the positioning latency and allowing the virtual model to more closely match the real-world image.

[0091] Fifth, the display rendering module is configured to collect real-time ambient lighting parameters and dynamically adjust the lighting and transparency of the virtual model to perform brightness and color adaptation processing between the virtual content and the real scene. Specifically, the display rendering module includes: a lighting estimation model, configured to obtain high dynamic range lighting parameters of the real environment based on illuminance data collected in real time by an ambient light sensor; a multilayer perceptron, configured to optimize the lighting response of the virtual model through MLPs algorithms; and a PBR rendering pipeline, configured to perform lighting adaptation processing between the virtual model and the real environment based on a brightness adaptation model. Its workflow is as follows:

[0092] 1. Ambient lighting adaptation:

[0093] The display rendering module is the final compositing unit of the system. It reads illuminance data (lux values) in real time from the ambient light sensor and runs a lighting estimation model: using a convolutional neural network to analyze camera images and estimate high dynamic range (HDR) lighting parameters of the real environment, including the direction, color, intensity, and ambient spherical harmonics of the main light source. These parameters are fed into multilayer perceptrons (MLPs) for learning and optimization—the MLPs output the necessary adjustments to diffuse and specular reflection intensities based on the material properties (metallicity, roughness, etc.) of the virtual model. Subsequently, the PBR rendering pipeline uses these lighting parameters to perform real-time lighting calculations on the 3D model received from the virtual content construction and optimization module, and dynamically adjusts the emitted light intensity and transparency of the model according to the brightness adaptation model. This ensures that the virtual model is not overexposed in strong light or underexposed in low light, achieving seamless color and brightness fusion with the real scene.

[0094] 2. Occlusion and Blending Rendering:

[0095] The display rendering module also utilizes pose and environmental awareness information obtained from the spatial positioning and perception module to determine the occlusion relationship between the virtual model and real objects. For example, when a real person or pillar should occlude the virtual model, it performs correct rendering through depth testing or semantic masking. Simultaneously, the module receives overlaid text information and personalized recommendations from the intelligent function module, converting them into a billboard in screen space or world space, and placing them according to the principle of not occluding the core scene.

[0096] 3. Final screen generation:

[0097] Finally, the display rendering module uses the interactive control signals (such as "rotate model", "display translated text", "answer questions") from the multimodal natural interaction module to synthesize all the above elements—high-precision 3D model (with predicted pose), overlaid text, UI controls, etc.—in real time onto the original camera image and outputs a video stream of no less than 60 frames per second to the screen or the optical display of the AR glasses, so that the user can see an immersive multimedia interactive screen that blends the virtual and real worlds.

[0098] The working principle of a scene multimedia immersive interactive system based on AR technology in this embodiment includes:

[0099] 1. Initialization: Each module starts up, and the spatial positioning and perception module begins to track the device pose and environmental semantics;

[0100] 2. Perception and cognition: Pose and semantic information are continuously distributed to the display rendering, interaction, and virtual content modules;

[0101] 3. User interaction: When a user issues a gesture / voice command, the multimodal natural interaction module recognizes and integrates the commands, generates a control signal, and sends it to the intelligent function module.

[0102] 4. Intelligent Response: The intelligent function module queries the knowledge base based on control signals, performs translation or recommendation, and generates text / recommendation content, which is then sent to the display rendering module;

[0103] 5. Model Optimization: The virtual content construction and optimization module loads and dynamically simplifies the model based on environmental perception and terminal performance, adds pose prediction, and then sends it to the display rendering module;

[0104] 6. Rendering and Display: The display and rendering module combines pose, ambient lighting, control signals, text content and 3D model to perform lighting adaptation, virtual and real occlusion and final compositing to output high-quality AR images;

[0105] 7. Closed loop: After seeing the screen, the user can proceed to the next round of interaction. The whole process runs in a loop with low latency, precise integration, and natural interaction.

[0106] Furthermore, this system supports various terminals such as mobile phones, tablets, and AR glasses. For devices without depth sensors, the spatial positioning and perception module can degenerate into a pure visual SLAM solution, with slightly reduced accuracy but still better than traditional solutions; for high-performance devices, LiDAR fusion and higher-precision semantic segmentation are enabled. The intelligent function module can dynamically load knowledge bases from different fields, enabling rapid migration from different scenarios such as "water culture" to "industrial equipment maintenance".

[0107] In summary, due to the adoption of the above technical solutions, the scene multimedia immersive interactive system based on AR technology in this embodiment has the following advantages compared with the prior art:

[0108] The interactive system of this invention achieves centimeter-level high-precision pose estimation through the fusion positioning of visual inertial odometry and depth sensors in the spatial positioning perception module, combined with semantic segmentation for environmental cognition. It can also constrain the placement and collision of virtual models based on scene elements such as ground, water, and buildings, significantly improving the accuracy of virtual-real fusion. The display rendering module utilizes an ambient light sensor, a lighting estimation model, a multilayer perceptron, and a PBR rendering pipeline to adaptively adjust the brightness and color of the virtual model under complex lighting conditions. Combined with model pruning and GPU acceleration, it ensures display stability and smoothness (frame rate ≥ 60fps) in strong light and dynamic environments. The multimodal natural interaction module integrates gesture recognition and voice interaction, and uses time... The system employs a combination of synchronization and state machine arbitration to avoid command conflicts, enabling natural, low-latency (gesture <20ms, voice <500ms) human-computer interaction. The virtual content construction and optimization module utilizes progressive mesh and pose prediction, balancing high-precision model performance with terminal performance adaptation, reducing ghosting or deviations caused by positioning latency. The intelligent function module, based on a terminal-side AI framework, enables intelligent question answering, real-time translation, and personalized recommendations without relying on the network, protecting privacy and responding rapidly, greatly enriching the information service capabilities of AR scenarios. The signal transmission paths between modules are clear, and the collaborative relationships are well-defined, forming a complete closed loop from environmental perception and interactive input to content generation and rendering display, resulting in high overall system reliability and strong scalability.

[0109] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A scene-based multimedia immersive interactive system based on AR technology, characterized in that, include: The spatial positioning and perception module is configured to calculate the pose of the device in the physical space in real time, and to perform semantic segmentation on scene elements in the physical environment to generate environmental cognitive information. The display rendering module is configured to collect real-time ambient lighting parameters and dynamically adjust the lighting and transparency of the virtual model to perform brightness and color adaptation processing between the virtual content and the real scene. A multimodal natural interaction module is configured to recognize user interaction commands and merge the interaction commands from different modalities into a unified interaction control signal; The virtual content construction and optimization module is configured to load 3D models. The intelligent function module is configured to run a functional model and, in conjunction with a knowledge base, provide users with information overlay display and personalized recommendations; Specifically, the spatial positioning and perception module sends the acquired pose information and environmental awareness information to the display rendering module, the multimodal natural interaction module, and the virtual content construction and optimization module. The multimodal natural interaction module sends the interaction control signal to the intelligent function module. The intelligent function module sends the generated overlaid text information and personalized recommendation content to the display rendering module. The virtual content construction and optimization module sends the optimized high-precision 3D model to the display rendering module. The display rendering module receives the pose information, environmental awareness information, overlaid text information, personalized recommendation content, and high-precision 3D model, and renders and generates the final augmented reality image according to the interaction control signal.

2. The scene multimedia immersive interactive system based on AR technology according to claim 1, characterized in that, The environmental perception information sent by the spatial positioning perception module to the virtual content construction and optimization module is used to guide model simplification or detail enhancement strategies. The pose information sent by the spatial positioning and perception module to the multimodal natural interaction module is used to assist in determining the spatial context of user interaction commands.

3. The scene multimedia immersive interactive system based on AR technology according to claim 1, characterized in that, The spatial positioning and perception module calculates the pose by fusing visual inertial odometry and depth sensor, and the fusion method can be tight coupling or loose coupling. The depth sensor is a LiDAR or structured light sensor.

4. The scene multimedia immersive interactive system based on AR technology according to claim 1, characterized in that, The semantic segmentation uses a deep learning model to identify the spatial location of each physical object in the scene, thereby constraining the placement of the virtual model and its physical collision response.

5. The scene multimedia immersive interactive system based on AR technology according to claim 1, characterized in that, The display rendering module includes: The illumination estimation model is configured to obtain high dynamic range illumination parameters of the real environment based on illuminance data collected in real time by an ambient light sensor. A multilayer perceptron configured to optimize the lighting response of a virtual model using MLPs algorithms; The PBR rendering pipeline is configured to adapt the lighting of the virtual model to the real environment based on a brightness adaptation model.

6. The scene multimedia immersive interactive system based on AR technology according to claim 1, characterized in that, The multimodal natural interaction module includes: A gesture recognition unit is configured to separate the hand region from the background through semantic segmentation and obtain the gesture command in the interaction command; A voice interaction unit is configured to support multilingual voice input and synthesis in order to obtain voice commands from the interaction instructions.

7. The scene multimedia immersive interactive system based on AR technology according to claim 1, characterized in that, The multimodal natural interaction module also includes: The timestamp synchronization unit is configured to unify the time base of data from all modalities. An interactive state machine configured to analyze the logical consistency of the interactive instructions across different modalities.

8. The scene multimedia immersive interactive system based on AR technology according to claim 1, characterized in that, The virtual content construction and optimization module constructs a three-dimensional model using three-dimensional scanning technology, and dynamically adjusts the model accuracy according to the performance of the terminal device using progressive mesh units, simplifying the mesh for non-critical parts; The virtual content construction and optimization module also includes a pose prediction module, which is configured to predict the model pose of the next frame based on historical pose data.

9. The scene multimedia immersive interactive system based on AR technology according to claim 1, characterized in that, The functional model includes a speech recognition model, a natural language processing model, and a machine translation model deployed based on a terminal-side artificial intelligence framework; The knowledge base supports automatic collection via web crawlers and manual review and updates. The intelligent function module includes a real-time intelligent question-answering unit and a real-time translation unit. The real-time intelligent question-answering unit is configured to retrieve answers from the knowledge base and display them overlaid in augmented reality text form. The real-time translation unit is configured to support real-time translation of voice or text input, and the translation results are displayed in the augmented reality scene in text form with adjustable parameters.