A WebXR-based immersive interactive interface generation system and method
Patent Information
- Application Number
- CN202511109906.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-08-08
AI Technical Summary
[0003]本申请所要解决的技术问题是相关技术无法基于WebXR技术高效、高质量地生成沉浸式交互界面的问题
[0016] Understandably, the beneficial effects of the WebXR-based immersive interactive interface generation method provided in the second aspect, the electronic device in the third aspect, and the computer-readable storage medium in the fourth aspect can be referenced to the beneficial effects of the first aspect and any of its possible design methods, and will not be repeated here.
Smart Images

Figure CN121209868B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interface rendering technology, and more specifically to an immersive interactive interface generation system and method based on WebXR. Background Technology
[0002] With the rapid development of information technology, WebXR technology, as a cross-platform solution integrating Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR), has been widely and deeply applied in many fields due to its convenience of not requiring the installation of dedicated clients and being accessible directly through a browser, as well as its unique advantages in creating an immersive experience for users. WebXR technology breaks through the limitations of traditional human-computer interaction in terms of time and space, bringing users a completely new immersive experience and greatly expanding the boundaries of human-computer interaction. It has not only changed the way people obtain information, communicate, and work, but also powerfully promoted the digital transformation and innovative development of various industries, becoming an important technological direction for the next generation of internet interaction models. Existing WebXR technologies primarily rely on traditional 2D interface design methods and underlying API development models. In terms of interface design, they employ planar interaction logic, such as clicks and swipes. During development, developers must directly call the underlying WebXR interfaces to implement functionality, and adaptation to different devices often relies on hard-coding. There is a lack of a unified component-based development framework and visual design tools, and the rendering architecture is mainly based on the traditional DOM for page drawing and updating. However, traditional 2D interface design methods and underlying development models are ill-suited to the 3D immersive interactive scenarios required by WebXR technology. This results in significant deficiencies in areas such as interaction dimension matching, spatial awareness adaptation, performance compatibility, development efficiency, UI stability, and occlusion optimization, making it difficult to meet the deep application needs of WebXR technology across multiple fields. Consequently, it becomes impossible to efficiently and effectively generate immersive interactive interfaces based on WebXR technology. Summary of the Invention
[0003] The technical problem this application aims to solve is that related technologies cannot efficiently and effectively generate immersive interactive interfaces based on WebXR technology.
[0004] To address the aforementioned technical problems, this application provides a WebXR-based immersive interactive interface generation system and method, specifically employing the following technical solution: In a first aspect, this invention provides an immersive interactive interface generation system based on WebXR. This system includes: a modular component configuration module, a visual design and editing module, an adaptive rendering module, and a cross-platform operation module. The modular component configuration module can receive device hardware parameter data and raw interactive input data, and process them to generate spatial anchoring data, component configuration data, and standardized interactive event data. The visual design and editing module can process the spatial anchoring data, component configuration data, and standardized interactive event data to generate scene description data and motion effect logic data. The adaptive rendering module can optimize rendering based on the spatial anchoring data, scene description data, and motion effect logic data to generate interactive interface rendering data and optimized spatial anchoring data. The cross-platform operation module can process the interactive interface rendering data and optimized spatial anchoring data, combined with real-time device status data, to generate interactive interface display data and feedback control data, and send them to the display device and feedback device respectively.
[0005] This system, firstly, receives device hardware parameter data and raw interactive input data through a modular component configuration module, processing and generating spatial anchoring data, component configuration data, and standardized interactive event data. Then, a visual design and editing module processes the spatial anchoring data, component configuration data, and standardized interactive event data to generate scene description data and motion effect logic data. Next, an adaptive rendering module optimizes rendering based on the spatial anchoring data, scene description data, and motion effect logic data, generating interactive interface rendering data and optimized spatial anchoring data. Finally, a cross-platform operation module processes the interactive interface rendering data and optimized spatial anchoring data, combined with real-time device status data, to generate interactive interface display data and feedback control data, which are then sent to the display device and feedback device, respectively. This system significantly improves development efficiency through its component-based architecture, visual design tools, and graphical logic orchestration system. By employing layered asynchronous rendering, a WebAssembly acceleration pipeline, and optimized rendering methods with dynamic spatial anchoring, it overcomes rendering performance bottlenecks, effectively improving the efficiency and quality of generating immersive interactive interfaces based on WebXR.
[0006] In conjunction with the first aspect, in one alternative implementation, the aforementioned raw interactive input data includes: 3D point cloud data of the physical environment, gesture sensor data, eye-tracking data, and speech waveform data. Spatial anchoring data includes: spatial anchor point coordinate data and volumetric collision parameter data. Standardized interactive event data includes: gesture feature vectors, gaze point coordinate data, speech command text data, tactile vibration parameters, and spatial audio data. Specifically, the modular component configuration module includes: a spatial primitive component module, an input adaptation layer module, and a feedback system component module. The spatial primitive component module can be used to generate spatial anchor point coordinate data and volumetric collision parameter data based on the 3D point cloud data of the physical environment. The input adaptation layer module can be used to process and generate gesture feature vectors, gaze point coordinate data, and speech command text data based on the gesture sensor data, eye-tracking data, and speech waveform data, respectively. The feedback system component module can be used to generate tactile vibration parameters and spatial audio data based on the gesture sensor data, eye-tracking data, and speech waveform data.
[0007] In conjunction with the first aspect, in one alternative implementation, the aforementioned input adaptation layer module includes: a gesture recognition unit, a gaze tracking unit, and a speech analysis unit. The gesture recognition unit can be used to determine gesture feature vectors based on gesture sensor data using the MediaPipeHands model. The gaze tracking unit can be used to determine gaze point coordinates based on eye-tracking data using a pupil localization algorithm. The speech analysis unit can be used to determine speech command text data based on speech waveform data using the Web SpeechAPI module.
[0008] In conjunction with the first aspect, in one alternative implementation, the aforementioned visual design editing module includes: a scene layout module, a hotspot definition module, and a motion effect orchestration module. The scene layout module can be used to determine the scene layout mode of the interactive interface based on spatial anchoring data and component configuration data; the scene layout mode includes at least one of the following: spherical layout mode, planar projection mode, and following view mode. The hotspot definition module can be used to determine collider data and interactive hotspot coordinate data based on spatial anchoring data and standardized interactive event data to obtain scene description data. The motion effect orchestration module can be used to generate GLSL shader code through an animation editor based on component configuration data and standardized interactive event data to obtain motion effect logic data.
[0009] In conjunction with the first aspect, in one alternative implementation, the aforementioned adaptive rendering module includes: a layered rendering control module, a dynamic spatial anchoring module, and a WebAssembly acceleration module. The layered rendering control module can be used to process scene description data into static, dynamic, and special effects layers, determining layered rendering instructions and LOD control parameters. The dynamic spatial anchoring module can be used to perform feature fusion and storage based on environmental feature data and device motion sensor data, and generate spatial anchor point drift compensation data through dynamic adjustment and compensation to optimize the spatial anchoring data and obtain optimized spatial anchoring data. The WebAssembly pipeline acceleration module can be used to optimize rendering based on layered rendering instructions, LOD control parameters, and motion effect logic data, generating interactive interface rendering data.
[0010] In conjunction with the first aspect, in one alternative implementation, the dynamic spatial anchoring module includes: an acquisition unit, a data preprocessing unit, a feature fusion and storage unit, a data retrieval and update unit, and a dynamic adjustment and compensation unit. The acquisition unit can be used to acquire environmental feature data and equipment motion sensor data. The data preprocessing unit can be used to extract features from the environmental feature data and equipment motion sensor data respectively, determining visual feature data and inertial feature data; and then filtering and denoising the visual feature data and inertial feature data respectively to obtain preprocessed visual feature data and preprocessed inertial feature data. The feature fusion and storage unit can be used to fuse the preprocessed visual feature data and preprocessed inertial feature data to obtain fused feature data, and store the fused feature data. The data retrieval and update unit can be used to retrieve the stored fused feature data in response to retrieval commands, and to update the stored fused feature data. The dynamic adjustment and compensation unit can be used to determine spatial anchor point drift compensation data based on the fused feature data through a compensation machine model.
[0011] In conjunction with the first aspect, in one alternative implementation, the aforementioned WebAssembly pipeline acceleration module specifically includes: a main thread unit, a WASM Worker Pool unit, a SIMD computational core unit, a shared memory unit, a GPU driver layer unit, and a WebXR rendering pipeline unit. The main thread unit can be used to create a WebXR session and obtain layered rendering instructions, LOD control parameters, and animation logic data. The WASM Worker Pool unit can be used to create at least two parallel worker threads based on the WebXR session, generate multiple sub-rendering tasks based on the layered rendering instructions, LOD control parameters, and animation logic data, allocate them to the worker threads, and ensure that the worker threads return the processing results to the main thread unit after processing. The SIMD computational core unit can be used to optimize matrix transformations and lighting calculations in the worker threads. The shared memory unit can be used to store layered rendering instructions, LOD control parameters, animation logic data, and processing result data. The GPU driver layer unit can be used to pass layered rendering instructions, LOD control parameters, and animation logic data to the WebXR rendering pipeline unit. The WebXR rendering pipeline unit can be used to optimize rendering based on layered rendering instructions, LOD control parameters, motion logic data, and sub-rendering tasks, generating interactive interface rendering data.
[0012] In conjunction with the first aspect, in one alternative implementation, the aforementioned cross-platform operation module specifically includes: a device status acquisition module, a display processing module, and a transmission module. The device status acquisition module is used to collect real-time device status data of the display device, which characterizes the pose state of the display device. The display processing module is used to perform data mapping and prediction based on interactive interface rendering data, optimized spatial anchoring data, and real-time device status data to generate interactive interface display data and feedback control data. The transmission module is used to send the interactive interface display data to the display device and the feedback control data to the feedback device.
[0013] Secondly, this application provides a method for generating an immersive interactive interface based on WebXR. This method includes: First, receiving device hardware parameter data and raw interactive input data, and processing them to generate spatial anchoring data, component configuration data, and standardized interactive event data. Then, processing the spatial anchoring data, component configuration data, and standardized interactive event data to generate scene description data and motion effect logic data. Next, optimizing rendering based on the spatial anchoring data, scene description data, and motion effect logic data to generate interactive interface rendering data and optimized spatial anchoring data. Finally, based on the interactive interface rendering data and optimized spatial anchoring data, combined with real-time device status data, processing to generate interactive interface display data and feedback control data, and sending them to the display device and feedback device respectively.
[0014] Thirdly, this application provides an electronic device comprising: a device body and an immersive interactive interface generation system based on WebXR, as described in the first aspect above and any alternative thereof, configured in the device body.
[0015] Fourthly, this application provides a computer-readable storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform the WebXR-based immersive interactive interface generation method provided in the second aspect above.
[0016] Understandably, the beneficial effects of the WebXR-based immersive interactive interface generation method provided in the second aspect, the electronic device in the third aspect, and the computer-readable storage medium in the fourth aspect can be referenced to the beneficial effects of the first aspect and any of its possible design methods, and will not be repeated here. Attached Figure Description
[0017] Figure 1 A schematic diagram of the architecture of the WebXR-based immersive interactive interface generation system provided in the embodiments of this application; Figure 2 This is a schematic diagram of the modular component configuration module provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the visual design editing module provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the adaptive rendering module provided in an embodiment of this application; Figure 5 This is a schematic diagram of the cross-platform operating module provided in an embodiment of this application; Figure 6 This is a flowchart illustrating the WebXR-based immersive interactive interface generation method provided in this application embodiment. Detailed Implementation
[0018] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.
[0019] WebXR technology, as a cross-platform solution integrating virtual reality, augmented reality, and mixed reality, has been widely used in many fields. It breaks through the limitations of traditional human-computer interaction, bringing users a brand-new immersive experience, greatly expanding the boundaries of human-computer interaction, and powerfully promoting the digital transformation and innovative development of various industries.
[0020] Existing WebXR technologies primarily rely on traditional 2D interface design methods and underlying API development models. In terms of interface design, they employ planar interaction logic, such as clicks and swipes. During development, developers must directly call the underlying WebXR interfaces to implement functionality, and adaptation to different devices often relies on hard-coding. There is a lack of a unified component-based development framework and visual design tools, and the rendering architecture is mainly based on the traditional DOM for page drawing and updating. However, existing technologies face several key challenges in 3D immersive environments: First, the interaction dimensions are mismatched; traditional planar interaction cannot effectively support 6-DOF spatial interaction, such as gesture recognition and gaze tracking. Second, spatial awareness is lacking, with insufficient dynamic adaptation capabilities to 3D spatial layouts, leading to a fragmented user experience. Third, performance and compatibility bottlenecks exist; traditional DOM rendering architectures struggle to meet the high frame rate requirements of XR scenarios, cross-device adaptation is labor-intensive, and compatibility with diverse hardware configurations ranging from mobile AR to high-end VR devices is difficult. Fourth, development efficiency is low; the lack of componentization and visualization tools results in long development cycles and low reusability. Fifth, UI element stability is insufficient; drifting is prone to occur in AR scenarios, and there are serious visual occlusion issues, affecting interaction accuracy and immersion. These problems significantly limit the in-depth application and development of WebXR technology across various fields.
[0021] It is evident that traditional 2D interface design methods and underlying development models are ill-suited to the 3D immersive interactive scenarios required by WebXR technology. This results in significant deficiencies in areas such as interaction dimension matching, spatial awareness adaptation, performance compatibility, development efficiency, UI stability, and occlusion optimization, making it difficult to meet the deep application needs of WebXR technology across multiple fields. Consequently, it is impossible to efficiently and effectively generate immersive interactive interfaces based on WebXR technology.
[0022] To address the aforementioned issues, this application provides a WebXR-based immersive interactive interface generation system and method. This system and method can be applied to the generation of immersive interactive interfaces in various fields, including interactive games and entertainment, education and training, healthcare, and industrial manufacturing. Specifically, the system and method are based on WebXR technology and include: a modular component configuration module, a visual design and editing module, an adaptive rendering module, and a cross-platform runtime module. Through a component-based architecture, visual design tools, and a graphical logic orchestration system, development efficiency is significantly improved. By employing layered asynchronous rendering, a WebAssembly acceleration pipeline, and optimized rendering methods with dynamic spatial anchoring, rendering performance bottlenecks are overcome, thereby effectively improving the efficiency and quality of WebXR-based immersive interactive interface generation.
[0023] The solutions provided in the embodiments of this application will be described below with reference to the accompanying drawings.
[0024] Specifically, Figure 1 This is a schematic diagram of the architecture of the WebXR-based immersive interactive interface generation system provided in the embodiments of this application, as shown below. Figure 1 As shown, the WebXR-based immersive interactive interface generation system 100 provided in this application embodiment includes: a modular component configuration module 110, a visual design editing module 120, an adaptive rendering module 130, and a cross-platform operation module 140.
[0025] The modular component configuration module 110 can be used to receive device hardware parameter data and raw interactive input data, and process and generate spatial anchoring data, component configuration data and standardized interactive event data.
[0026] In this embodiment, the original interactive input data may specifically include: 3D point cloud data of the physical environment (e.g., SLAM point cloud data), gesture sensor data, eye-tracking data, and speech waveform data. Spatial anchoring data includes: spatial anchor point coordinate data and volume collision parameter data. Standardized interactive event data includes: gesture feature vectors, gaze point coordinate data, voice command text data, tactile vibration parameters, and spatial audio data.
[0027] In some embodiments, Figure 2 This is a schematic diagram of the structure of the modular component configuration module provided in the embodiments of this application, such as... Figure 2 As shown, the modular component configuration module 110 includes: a spatial primitive component module 111, an input adaptation layer module 112, and a feedback system component module 113.
[0028] Among them, the spatial primitive component module 111 can be used to generate spatial anchor point coordinate data and volume collision parameter data based on three-dimensional point cloud data of the physical environment.
[0029] For example, the spatial primitive component module 111 may specifically include an AnchorPoint component and a VolumetricButton component. The AnchorPoint component can generate spatial anchor point coordinate data by binding UI elements to the physical environment through spatial anchors based on the physical environment's 3D point cloud data. The VolumetricButton component can perform volumetric interaction processing for 3D collision detection based on the physical environment's 3D point cloud data to determine volumetric collision parameter data.
[0030] The input adaptation layer module 112 can be used to process gesture sensor data, eye tracking data and voice waveform data to generate gesture feature vectors, gaze coordinate data and voice command text data respectively, so as to obtain standardized interaction event data.
[0031] In some embodiments, see continue to see Figure 2 The input adaptation layer module 112 specifically includes: a gesture recognition unit 1121, a gaze tracking unit 1122, and a speech analysis unit 1123.
[0032] The gesture recognition unit 1121 can be used to determine the gesture feature vector based on the gesture sensor data and the MediaPipeHands model.
[0033] The gaze tracking unit 1122 can be used to determine the gaze point coordinates based on eye-tracking data and a pupil positioning algorithm.
[0034] The speech analysis unit 1123 can be used to determine the speech command text data based on the speech waveform data through the Web Speech API module.
[0035] The feedback system component module 113 can be used to generate tactile vibration parameters and spatial audio data based on gesture sensor data, eye tracking data and voice waveform data.
[0036] For example, the feedback system component module 113 can determine spatial audio data through the WebAudio API based on eye-tracking data and speech waveform data, and can generate tactile vibration parameters through the WebHaptics API based on gesture sensor data and eye-tracking data, and can drive the device vibration module based on the tactile vibration parameters.
[0037] The aforementioned visualization design and editing module 120 can be used to generate scene description data and motion effect logic data based on spatial anchoring data, component configuration data, and standardized interactive event data processing.
[0038] In some embodiments, Figure 3This is a schematic diagram of the structure of the visual design and editing module provided in the embodiments of this application, such as... Figure 3 As shown, the visual design editing module 120 may specifically include: a scene layout module 121, a hot zone definition module 122, and a motion effect arrangement module 123.
[0039] Specifically, the scene layout module 121 can be used to determine the scene layout mode of the interactive interface based on spatial anchoring data and component configuration data.
[0040] The scene layout modes include at least one of the following: spherical layout mode, planar projection mode, and view-following mode. Specifically, the spherical layout mode distributes UI elements along a sphere centered on the user's head, where the radius of the sphere can be preset according to application requirements. The planar projection mode dynamically projects the 2D interface onto a physical surface (e.g., based on a plane detection API). The view-following mode maintains a relative position between the UI elements and the user's viewpoint. For example, an anti-dizziness algorithm can be used to control movement speed to achieve the view-following mode.
[0041] The hot zone definition module 122 can be used to determine collision body data and interactive hot zone coordinate data based on spatial anchoring data and standardized interactive event data, so as to obtain scene description data.
[0042] For example, the hotspot definition module 122 can use the Bounding Box visualization tool to set hotspot areas based on spatial anchoring data and standardized interactive event data to determine the coordinate data of the interactive hotspots. It also determines the collision volume by dynamically adjusting the collision volume (e.g., Sphere, Box, and Capsule shapes), using the collision volume data and the interactive hotspot coordinate data as scene description data.
[0043] The motion orchestration module 123 can be used to generate GLSL shader code through the animation editor based on component configuration data and standardized interactive event data to obtain motion logic data.
[0044] Specifically, the motion orchestration module 123 can generate GLSL shader code based on component configuration data and standardized interactive event data, using the timeline-based animation editor and physical simulation parameters (elasticity coefficient, damping ratio), to obtain motion logic data and optimize animation performance.
[0045] The aforementioned adaptive rendering module 130 can be used to optimize rendering based on spatial anchoring data, scene description data, and motion effect logic data, generating interactive interface rendering data and optimized spatial anchoring data.
[0046] In some embodiments, Figure 4 This is a schematic diagram of the structure of the adaptive rendering module provided in the embodiments of this application, as shown below. Figure 4 As shown, the adaptive rendering module 130 may specifically include: a layered rendering control module 131, a dynamic spatial anchoring module 132, and a WebAssembly acceleration module 133.
[0047] Specifically, the layered rendering control module 131 can be used to perform layered processing based on scene description data, according to static layer, dynamic layer, and special effects layer, and determine layered rendering instructions and LOD control parameters.
[0048] The static layer is used for scene rendering based on static light sources, such as background elements using pre-baked lightmaps, which can be rendered once per frame. The dynamic layer is used for interactive components that need to be dynamically updated in real-time during scene rendering; for example, interactive components can be updated at a frequency of 60Hz. The effects layer adds post-processing effects after scene rendering to achieve special effects, such as Bloom, depth of field, and asynchronous rendering.
[0049] The dynamic spatial anchoring module 132 can be used to perform feature fusion and storage based on environmental feature data and equipment motion sensor data, and generate spatial anchor point drift compensation data through dynamic adjustment and compensation, so as to optimize the spatial anchoring data and determine the optimized spatial anchoring data.
[0050] In some embodiments, such as Figure 4 As shown, the dynamic spatial anchoring module 132 specifically includes: an acquisition unit 1321, a data preprocessing unit 1322, a feature fusion and storage unit 1323, a data retrieval and update unit 1324, and a dynamic adjustment and compensation unit 1325.
[0051] Specifically, the acquisition unit 1321 can be used to acquire environmental feature data and equipment motion sensor data.
[0052] For example, the acquisition unit 1321 can acquire environmental feature data through a camera device and obtain motion sensor data of the device through the device's inertial sensors (such as accelerometers and gyroscopes).
[0053] The data preprocessing unit 1322 can be used to extract features from environmental feature data and equipment motion sensor data respectively, to determine visual feature data and inertial feature data. It then performs filtering and denoising processing on the visual feature data and inertial feature data respectively, to obtain preprocessed visual feature data and preprocessed inertial feature data.
[0054] For example, the data preprocessing unit 1322 can use computer vision technology to extract key feature points, edges, or texture information from environmental feature data to obtain visual feature data. It can also extract inertial feature data such as acceleration and angular velocity from the device motion sensor data.
[0055] The feature fusion and storage unit 1323 can be used to fuse preprocessed visual feature data and preprocessed inertial feature data to obtain fused feature data and store the fused feature data.
[0056] Specifically, the feature fusion and storage unit 1323 can perform feature fusion on preprocessed visual feature data and preprocessed inertial feature data. For example, feature fusion can be achieved through methods such as data association, feature matching, or Kalman filtering to obtain more accurate positioning information. Furthermore, the feature fusion and storage unit 1323 can store the fused feature data in the device's local storage or on a cloud server. Depending on the real-time requirements and data volume, a storage format and data structure that meets the application requirements can be selected, such as binary files or databases.
[0057] In one implementation, the feature fusion and storage unit 1323 can employ an adaptive storage method based on user behavior when storing fused feature data from multiple users. Specifically, the storage strategy and retrieval method of the fused feature data can be dynamically adjusted according to users' usage habits and behavioral patterns. For example, for users who move frequently, recent fused feature data can be stored first, and the retrieval algorithm can be optimized to improve real-time performance. For users who are static or move infrequently, the storage structure can be optimized to reduce storage space usage.
[0058] The data retrieval and update unit 1324 can be used to retrieve stored fusion feature data in response to a retrieval command, and to update stored fusion feature data.
[0059] Specifically, the data retrieval and update unit 1324 can retrieve the stored fusion feature data in real time as needed to support positioning, tracking, or interactive functions. Furthermore, the data retrieval and update unit 1324 can continuously update the stored fusion feature data as the environment changes or the device moves, to maintain the accuracy and timeliness of the fusion feature data.
[0060] The dynamic adjustment and compensation unit 1325 can be used to determine spatial anchor point drift compensation data based on fused feature data and through a compensation machine model.
[0061] Specifically, the dynamic adjustment and compensation unit 1325 can employ an adaptive compensation method. For example, it can dynamically select either the ICP or Kalman filtering algorithm based on device performance to prioritize the use of low-power mode on the mobile device (actual power consumption can be reduced by 2.1W).
[0062] For example, the selection and switching between ICP and Kalman filtering algorithms includes: the dynamic adjustment and compensation unit 1325 can dynamically select either the ICP or Kalman filtering algorithm for UI element positioning and tracking based on device performance and environmental conditions. For instance, the ICP algorithm can be selected in scenarios requiring high precision, while the Kalman filtering algorithm can be selected in scenarios requiring high real-time performance. By monitoring device performance and environmental changes in real time, a smooth switching between the ICP and Kalman filtering algorithms can be achieved to ensure that UI elements remain stably attached in different scenarios.
[0063] Furthermore, the ICP and Kalman filtering algorithms can be optimized for performance, such as reducing computational load and improving convergence speed, to meet the real-time and accuracy requirements of WebXR applications. The execution efficiency of the algorithms can also be further improved by introducing parallel computing or GPU acceleration technologies.
[0064] The dynamic adjustment and compensation unit 1325 employs an adaptive compensation method to adjust the position and orientation of UI elements in real time based on environmental changes and user behavior. For example, when user movement or changes in ambient light are detected, the strategy automatically adjusts the display position and brightness of UI elements to ensure they are always clearly visible. By introducing machine learning or deep learning algorithms, the adaptive compensation strategy can learn user habits and behavioral patterns, thereby making adjustments and compensations more intelligently.
[0065] Meanwhile, the dynamic adjustment and compensation unit 1325 employs an adaptive compensation method to reduce UI element drift or misalignment caused by environmental changes or device errors, thereby improving its stability and accuracy. By fusing data from other sensors (such as gyroscopes and accelerometers), the adaptive compensation method further enhances the accuracy of UI element attachment. Furthermore, it achieves cross-platform coordinate unification: by using WebXR's XRSpace to transform the spatial coordinate systems of different devices, the error compensation algorithm ensures that the positioning deviation between multiple devices is less than 1.5cm.
[0066] The dynamic adjustment and compensation unit 1325, combined with adaptive compensation methods and ICP and Kalman filtering algorithms, can provide users with a more natural and smooth interactive experience. For example, when performing object grabbing or placing operations in a virtual environment, UI elements can follow the user's hand movements in real time and remain stably attached, thereby improving the accuracy and immersion of the operation.
[0067] The WebAssembly pipeline acceleration module 133 can be used to optimize rendering based on layered rendering instructions, LOD control parameters, and motion logic data to generate interactive interface rendering data.
[0068] Specifically, the WebAssembly pipeline acceleration module 133 can adopt a three-level acceleration pipeline combining WASM Worker Pool, SIMD, and GPU to achieve near-native performance in the browser environment.
[0069] In some embodiments, see continue to see Figure 4 The WebAssembly pipeline acceleration module 133 specifically includes: a main thread unit 1331, a WASM WorkerPool unit 1332, a SIMD computing core unit 1333, a shared memory unit 1334, a GPU driver layer unit 1335, and a WebXR rendering pipeline unit 1336.
[0070] Specifically, the main thread unit 1331 can be used to create WebXR sessions and obtain layered rendering instructions, LOD control parameters, and motion logic data.
[0071] For example, the main thread unit 1331 can initiate an immersive session by calling the WebXR API, i.e., creating a WebXR session, and obtain relevant device and input information, i.e., obtaining layered rendering instructions, LOD control parameters, and motion logic data. The main thread unit 1331 can also be used to load and unload scene resources, and manage objects and interaction logic in the scene. It receives input from the user (such as head tracking, controller actions, etc.) and updates the scene state in response to user actions.
[0072] The WASM WorkerPool unit 1332 can be used to create at least two parallel worker threads based on a WebXR session, generate multiple sub-rendering tasks based on layered rendering instructions, LOD control parameters and motion logic data and assign them to the worker threads, and make the worker threads send the processing result data back to the main thread unit 1331 after processing.
[0073] In this embodiment, the WASM WorkerPool unit 1332 can utilize WebAssembly (WASM) technology to create parallel computing threads (i.e., worker threads) to improve application performance. In WebXR applications, the WASMWorkerPool unit 1332 can be used to perform computationally intensive tasks such as complex physics simulations and collision detection.
[0074] The creation of Worker threads can include: WASM WorkerPool unit 1332 creating multiple Worker threads using WASM-compiled code to achieve parallel computing; assigning sub-rendering tasks to Worker threads for execution to reduce the burden on the main thread; and after a Worker thread completes its task, returning the result to the main thread unit 1331 for further processing.
[0075] The SIMD computing core unit 1333 can be used to optimize the processing of matrix transformations and lighting calculations in the Worker thread.
[0076] For example, the SIMD computing core unit 1333 can process multiple data elements simultaneously, thereby significantly improving computational efficiency. In WebXR applications, the SIMD computing core unit 1333 can be used to optimize computational processes such as matrix transformation and lighting calculation in the Worker thread. Matrix transformation refers to accelerating the matrix transformation calculation of objects in a 3D scene using the SIMD instruction set. Lighting calculation refers to optimizing the calculation process of the lighting model using SIMD technology to improve rendering efficiency.
[0077] Shared memory unit 1334 can be used to store layered rendering instructions, LOD control parameters, motion logic data, and processing result data.
[0078] Specifically, the shared memory unit 1334 enables efficient data transfer between multiple threads or processes. In WebXR applications, the shared memory unit 1334 can transfer rendering data, user input information, etc., between the main thread and worker threads via shared memory.
[0079] For example, memory allocation can be performed by allocating a shared memory region within the main thread. Data transfer can write data such as layered rendering instructions, LOD control parameters, animation logic data, and processing result data into the shared memory region for access by worker threads or the GPU.
[0080] The GPU driver layer unit 1335 can be used to pass layered rendering instructions, LOD control parameters, and motion logic data to the WebXR rendering pipeline unit 1336.
[0081] Specifically, the GPU driver layer unit 1335 can pass layered rendering instructions, LOD control parameters, and motion logic data to the GPU and WebXR rendering pipeline unit 1336 for accelerated rendering. In WebXR applications, the GPU driver layer needs to cooperate with the WebXR rendering pipeline unit 1336 to achieve high-efficiency rendering performance.
[0082] For example, the GPU driver layer unit 1335 can implement the following functions: Rendering command submission, i.e., sending layered rendering instructions, LOD control parameters, and animation logic data to the GPU driver layer for processing. Resource management: managing resources such as textures and shaders in the GPU to ensure smooth rendering. Performance optimization: optimizing rendering performance by adjusting rendering parameters and using hardware acceleration techniques. Synchronization mechanism: using appropriate synchronization mechanisms (such as semaphores, mutexes, etc.) to ensure data consistency and security.
[0083] The WebXR rendering pipeline unit 1336 can be used to optimize rendering based on layered rendering instructions, LOD control parameters, motion logic data, and sub-rendering tasks to generate interactive interface rendering data.
[0084] Specifically, the WebXR rendering pipeline unit 1336 can be used to render 3D scenes onto the user's display device, that is, to optimize rendering based on layered rendering instructions, LOD control parameters, motion effect logic data and sub-rendering tasks to generate interactive interface rendering data.
[0085] For example, the WebXR rendering pipeline unit 1336 can perform the following functions: Scene building: Building 3D scenes using 3D modeling software or engines and exporting them in the format required by WebXR applications. Shader compilation: Writing and compiling vertex shaders and fragment shaders to achieve the scene rendering effect. Rendering loop: In each rendering loop, updating the scene state, executing rendering commands, and outputting the rendering results to the display device. Interaction processing: Updating the scene state based on user input (such as head tracking, controller actions, etc.) and re-rendering the scene in response to user actions.
[0086] In some embodiments, the data resources processed in the adaptive rendering module 130 can employ a resource reclamation mechanism. Specifically, traditional JavaScript implementations are slow (typically 400ms / 1024²). WebGL's native compressed texture format has poor compatibility and requires pre-compression. For UI components in invisible areas, GPU memory reclamation methods can be used, along with WASM-accelerated texture compression algorithms, such as the BC7 format. Specifically, the BC7 format features include: support for 8 compression modes (i.e., 3-6bpp variable bitrate), and each 4×4 pixel block can be independently encoded. Furthermore, it enables real-time processing, meeting the needs of VR applications.
[0087] The cross-platform operation module 140 can be used to process and generate interactive interface display data and feedback control data based on interactive interface rendering data and optimized spatial anchoring data, combined with real-time device status data, and send them to the display device and feedback device respectively.
[0088] Among them, the display device can be, for example, VR glasses, VR helmet display, etc., and the feedback device can be, for example, a controller device, force feedback glove device, etc.
[0089] In some embodiments, Figure 5 This is a schematic diagram of the structure of the cross-platform operating module provided in the embodiments of this application, such as... Figure 5 As shown, the cross-platform operation module 140 includes: a device status acquisition module 141, a display processing module 142, and a transmission module 143.
[0090] Specifically, the device status acquisition module 141 can be used to acquire real-time device status data of the display device, and the real-time device status data is used to characterize the pose state of the display device.
[0091] The display processing module 142 can be used to perform data mapping and prediction based on interactive interface rendering data, optimized spatial anchoring data and real-time device status data, and generate interactive interface display data and feedback control data.
[0092] The sending module 143 can be used to send interactive interface display data to the display device and to send feedback control data to the feedback device.
[0093] In some embodiments, the cross-platform runtime module 140 may include a unified input abstraction layer module, a context prediction engine module, and an anomaly degradation module. The unified input abstraction layer module can map XR controller, gesture, and voice input to a standard event stream and can eliminate tracking jitter through Kalman filtering to improve processing accuracy. The context prediction engine module can predict the next interaction target based on the user's gaze trajectory (i.e., gaze point coordinate data) using the Eye Tracking API to preload resources for the predicted target and reduce interaction latency. The anomaly degradation module can automatically disable post-processing effects when the frame rate is below 45fps and can switch to low-poly mode when violent head movements are detected.
[0094] The WebXR-based immersive interactive interface generation system provided in the above embodiments of this application firstly receives device hardware parameter data and raw interactive input data through a modular component configuration module, and processes and generates spatial anchoring data, component configuration data, and standardized interactive event data. Then, a visual design and editing module processes the spatial anchoring data, component configuration data, and standardized interactive event data to generate scene description data and motion effect logic data. Next, an adaptive rendering module optimizes rendering based on the spatial anchoring data, scene description data, and motion effect logic data to generate interactive interface rendering data and optimized spatial anchoring data. Finally, a cross-platform operation module processes the interactive interface rendering data and optimized spatial anchoring data, combined with real-time device status data, to generate interactive interface display data and feedback control data, and sends them to the display device and feedback device respectively.
[0095] The system offers the following benefits: Improved development efficiency: Compared to native WebXR development, component reuse rate is increased by 85%, and interface setup time is reduced to 1 / 4. Improved performance metrics: Stable 90fps rendering is achieved on Quest 3, with a 40% reduction in GPU utilization. Improved interaction accuracy: The false trigger rate for gesture recognition has decreased from 12.3% to 3.8% (test dataset: 500 gesture samples). Improved cross-platform compatibility: A single codebase adapts to 5 types of XR devices, reducing adaptation workload by 90%.
[0096] In summary, the WebXR-based immersive interactive interface generation system provided in the above embodiments of this application significantly improves development efficiency through a component-based architecture, visual design tools, and a graphical logic orchestration system. By employing layered asynchronous rendering, a WebAssembly accelerated pipeline, and optimized rendering methods with dynamic spatial anchoring, it overcomes rendering performance bottlenecks, thereby effectively improving the efficiency and quality of WebXR-based immersive interactive interface generation.
[0097] This application also provides a method for generating an immersive interactive interface based on WebXR. Figure 6 A flowchart illustrating the WebXR-based immersive interactive interface generation method provided in this application embodiment is shown below. Figure 6 As shown, the method includes the following steps S201-S204: S201: Receive device hardware parameter data and raw interactive input data, process and generate spatial anchoring data, component configuration data and standardized interactive event data.
[0098] In some embodiments, the raw interactive input data includes: 3D point cloud data of the physical environment, gesture sensor data, eye-tracking data, and speech waveform data. Spatial anchoring data includes: spatial anchor point coordinate data and volume collision parameter data. Standardized interactive event data includes: gesture feature vectors, gaze point coordinate data, speech command text data, tactile vibration parameters, and spatial audio data. S201 specifically includes the following steps: First, based on the 3D point cloud data of the physical environment, spatial anchor point coordinate data and volume collision parameter data are generated.
[0099] Then, gesture feature vectors, gaze coordinate data, and voice command text data are generated based on gesture sensor data, eye tracking data, and voice waveform data, respectively.
[0100] Finally, based on gesture sensor data, eye-tracking data, and speech waveform data, tactile vibration parameters and spatial audio data are generated.
[0101] S202. Based on spatial anchoring data, component configuration data, and standardized interactive event data processing, generate scene description data and motion effect logic data.
[0102] In some embodiments, S202 specifically includes the following steps: First, the scene layout mode of the interactive interface is determined based on spatial anchoring data and component configuration data; the scene layout mode includes at least one of the following: spherical layout mode, planar projection mode, and following view mode.
[0103] Then, based on spatial anchoring data and standardized interaction event data, collision body data and interaction hot zone coordinate data are determined to obtain scene description data.
[0104] Finally, based on component configuration data and standardized interactive event data, GLSL shader code is generated through the animation editor to obtain motion effect logic data.
[0105] S203. Optimize rendering based on spatial anchoring data, scene description data, and motion effect logic data to generate interactive interface rendering data and optimized spatial anchoring data.
[0106] In some embodiments, S203 specifically includes the following steps: First, based on the scene description data, the data is processed in layers: static layer, dynamic layer, and special effects layer, and the layer rendering instructions and LOD control parameters are determined.
[0107] Then, feature fusion and storage are performed based on environmental feature data and equipment motion sensor data, and spatial anchor point drift compensation data is generated through dynamic adjustment and compensation to optimize the spatial anchoring data and determine the optimized spatial anchoring data.
[0108] Finally, based on layered rendering instructions, LOD control parameters, and motion effect logic data, the rendering is optimized to generate the interactive interface rendering data.
[0109] S204. Based on the interactive interface rendering data and optimized spatial anchoring data, combined with the device's real-time status data, process and generate interactive interface display data and feedback control data, and send them to the display device and feedback device respectively.
[0110] In some embodiments, S204 specifically includes the following steps: First, real-time status data of the display device is collected, which is used to characterize the pose state of the display device.
[0111] Then, based on the interactive interface rendering data, optimized spatial anchoring data, and real-time device status data, data mapping and prediction are performed to generate interactive interface display data and feedback control data.
[0112] Finally, the interactive interface display data is sent to the display device, and the feedback control data is sent to the feedback device.
[0113] The WebXR-based immersive interactive interface generation method provided in the above embodiments of this application first receives device hardware parameter data and raw interactive input data, and processes them to generate spatial anchoring data, component configuration data, and standardized interactive event data. Then, based on the spatial anchoring data, component configuration data, and standardized interactive event data, scene description data and motion effect logic data are generated. Next, based on the spatial anchoring data, scene description data, and motion effect logic data, optimized rendering is performed to generate interactive interface rendering data and optimized spatial anchoring data. Finally, based on the interactive interface rendering data and optimized spatial anchoring data, combined with real-time device status data, interactive interface display data and feedback control data are generated and sent to the display device and feedback device respectively. This significantly improves development efficiency through a component-based architecture, visual design tools, and a graphical logic orchestration system. By using layered asynchronous rendering, a WebAssembly acceleration pipeline, and optimized rendering methods with dynamic spatial anchoring, the rendering performance bottleneck is overcome, thereby effectively improving the efficiency and quality of WebXR-based immersive interactive interface generation.
[0114] This invention also provides an electronic device, which may include: a device body and a WebXR-based immersive interactive interface generation system configured in the device body as described in the above embodiments.
[0115] This invention also provides a computer-readable storage medium for storing computer instructions for running the above-described WebXR-based immersive interactive interface generation method.
[0116] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0117] In the description of this invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0118] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0119] Similar parts between the embodiments provided in this application can be referred to mutually. The specific implementation methods provided above are only a few examples under the overall concept of this application and do not constitute a limitation on the scope of protection of this application. For those skilled in the art, any other implementation methods extended from the solution of this application without creative effort shall fall within the scope of protection of this application.
Claims
1. A WebXR-based immersive interactive interface generation system, characterized in that, include: Modular component configuration module, visual design and editing module, adaptive rendering module, and cross-platform operation module; among them, The modular component configuration module is used to receive device hardware parameter data and raw interactive input data, and process and generate spatial anchoring data, component configuration data, and standardized interactive event data. The raw interactive input data includes: physical environment 3D point cloud data, gesture sensor data, eye tracking data, and speech waveform data. The spatial anchoring data includes: spatial anchor point coordinate data and volume collision parameter data. The standardized interactive event data includes: gesture feature vectors, gaze point coordinate data, voice command text data, tactile vibration parameters, and spatial audio data. The modular component configuration module includes: a spatial primitive component module, an input adaptation layer module, and a feedback system component module; wherein, the spatial primitive component module is used to generate the spatial anchor point coordinate data and the volume collision parameter data based on the three-dimensional point cloud data of the physical environment; the input adaptation layer module is used to process the gesture sensor data, the eye tracking data, and the speech waveform data to generate the gesture feature vector, the gaze point coordinate data, and the speech command text data, respectively; the feedback system component module is used to generate the tactile vibration parameters and the spatial audio data based on the gesture sensor data, the eye tracking data, and the speech waveform data; The visualization design and editing module is used to generate scene description data and motion effect logic data based on the spatial anchoring data, the component configuration data, and the standardized interactive event data processing. The adaptive rendering module is used to optimize rendering based on the spatial anchoring data, the scene description data, and the motion effect logic data, generating interactive interface rendering data and optimized spatial anchoring data. The adaptive rendering module includes a layered rendering control module, a dynamic spatial anchoring module, and a WebAssembly acceleration module. The layered rendering control module is used to perform layered processing based on the scene description data, according to static layers, dynamic layers, and special effects layers, determining layered rendering instructions and LOD control parameters. The dynamic spatial anchoring module is used to perform feature fusion and storage based on environmental feature data and device motion sensor data, and to generate spatial anchor point drift compensation data through dynamic adjustment and compensation, for optimizing the spatial anchoring data and determining the optimized spatial anchoring data. The WebAssembly acceleration module is used to optimize rendering based on the layered rendering instructions, the LOD control parameters, and the motion effect logic data, generating the interactive interface rendering data. The cross-platform operation module is used to process and generate interactive interface display data and feedback control data based on the interactive interface rendering data and the optimized spatial anchoring data, combined with the device's real-time status data, and send them to the display device and feedback device respectively.
2. The system according to claim 1, characterized in that, The input adaptation layer module includes: a gesture recognition unit, a gaze tracking unit, and a speech analysis unit; wherein... The gesture recognition unit is used to determine the gesture feature vector based on the gesture sensor data using the MediaPipe Hands model; The gaze tracking unit is used to determine the gaze point coordinates based on the eye tracking data using a pupil localization algorithm. The speech analysis unit is used to determine the speech command text data based on the speech waveform data through the Web Speech API module.
3. The system according to claim 1, characterized in that, The visual design and editing module includes: a scene layout module, a hotspot definition module, and an animation choreography module; wherein... The scene layout module is used to determine the scene layout mode of the interactive interface based on the spatial anchoring data and the component configuration data; the scene layout mode includes at least one of the following: spherical layout mode, planar projection mode, and following view mode; The hot zone definition module is used to determine the collision body data and the interactive hot zone coordinate data based on the spatial anchoring data and the standardized interactive event data, so as to obtain the scene description data. The motion effects orchestration module is used to generate GLSL shader code through an animation editor based on the component configuration data and the standardized interactive event data, so as to obtain the motion effects logic data.
4. The system according to claim 1, characterized in that, The dynamic spatial anchoring module includes: a data acquisition unit, a data preprocessing unit, a feature fusion and storage unit, a data retrieval and update unit, and a dynamic adjustment and compensation unit; wherein... The acquisition unit is used to acquire the environmental feature data and the motion sensor data of the device; The data preprocessing unit is used to extract features from the environmental feature data and the device motion sensor data to determine visual feature data and inertial feature data; and to filter and denoise the visual feature data and the inertial feature data to obtain preprocessed visual feature data and preprocessed inertial feature data. The feature fusion and storage unit is used to fuse the preprocessed visual feature data and the preprocessed inertial feature data to obtain fused feature data, and to store the fused feature data. The data retrieval and update unit is used to retrieve the stored fusion feature data in response to a retrieval command, and to update the stored fusion feature data. The dynamic adjustment and compensation unit is used to determine the spatial anchor point drift compensation data based on the fused feature data and through a compensation machine model.
5. The system according to claim 1, characterized in that, The WebAssembly acceleration module includes: a main thread unit, a WASM Worker Pool unit, a SIMD computing core unit, a shared memory unit, a GPU driver layer unit, and a WebXR rendering pipeline unit; wherein... The main thread unit is used to create a WebXR session and obtain the layered rendering instructions, the LOD control parameters, and the motion effect logic data. The WASM Worker Pool unit is used to create at least two parallel Worker threads based on the WebXR session, generate multiple sub-rendering tasks based on the layered rendering instructions, the LOD control parameters and the animation logic data and assign them to the Worker threads, and make the Worker threads send the processing result data back to the main thread unit after processing. The SIMD computing core unit is used to optimize the processing of matrix transformation and lighting calculation in the Worker thread; The shared memory unit is used to store the layered rendering instructions, the LOD control parameters, the animation logic data, and the processing result data; The GPU driver layer unit is used to pass the layered rendering instructions, the LOD control parameters and the motion effect logic data to the WebXR rendering pipeline unit; The WebXR rendering pipeline unit is used to optimize rendering based on the layered rendering instructions, the LOD control parameters, the motion effect logic data, and the sub-rendering tasks, and generate the interactive interface rendering data.
6. The system according to claim 1, characterized in that, The cross-platform operation module includes: a device status acquisition module, a display processing module, and a transmission module; wherein... The device status acquisition module is used to acquire real-time device status data of the display device, and the real-time device status data is used to characterize the pose state of the display device. The display processing module is used to perform data mapping and prediction based on the interactive interface rendering data, the optimized spatial anchoring data and the real-time device status data, and generate the interactive interface display data and the feedback control data. The sending module is used to send the interactive interface display data to the display device and the feedback control data to the feedback device.
7. A method for generating an immersive interactive interface based on WebXR, characterized in that, The system applied to any one of claims 1-6 comprises: Receive device hardware parameter data and raw interactive input data, process and generate spatial anchoring data, component configuration data and standardized interactive event data; Based on the spatial anchoring data, the component configuration data, and the standardized interactive event data processing, scene description data and motion effect logic data are generated. Based on the spatial anchoring data, the scene description data, and the motion effect logic data, the rendering is optimized to generate interactive interface rendering data and optimized spatial anchoring data. Based on the interactive interface rendering data and the optimized spatial anchoring data, combined with the device's real-time status data, interactive interface display data and feedback control data are generated and sent to the display device and feedback device, respectively.
8. An electronic device, characterized in that, include: The main body of the device and the WebXR-based immersive interactive interface generation system configured in the main body as described in any one of claims 1-6.
Citation Information
Patent Citations
Interface interaction method
CN119536572A
Precise positioning method and system for intelligent Internet of Things equipment and universe virtual-real fusion
CN120029465A
Web-side large-scale model performance optimization method
CN120070701A
Interaction method and device, storage medium and equipment
CN120428882A