Multi-independent three-dimensional model double-sensor configurable switching driving ar glasses space virtual-real rendering method
Patent Information
- Application Number
- CN202610823513.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-08-28
AI Technical Summary
本发明的目的在于克服现有 AR 眼镜三维模型渲染控制模式固化、三维模型素材多版本冗余、头部绑定模型运动形式单一、眼动传感无法驱动模型形变、多三维模型运动强制同步、双传感器长期高负载运行、作业依赖手持外设的缺陷,提供一种多独立三维模型双传感可组态切换驱动的 AR 眼镜空间虚实渲染方法
[0009]This invention simplifies the resource system and significantly reduces production and maintenance costs by using only independent 3D models as rendering carriers, eliminating auxiliary materials such as videos and images. It adopts a single resource package architecture, where one set of 3D model resources can be compatible with three driving modes: head, eye-tracking, and dual-channel collaborative. This eliminates the need to create multiple sets of 3D model projects for different control modes, reducing the workload of modeling, parameter debugging, and version management, and significantly reducing the costs of content production, storage, and maintenance. At the same time, it forms a standardized 3D resource encapsulation system.
Smart Images

Figure FT_1
Abstract
Description
Technical Field
[0001] This invention relates to the fields of augmented reality wearable terminals, wearable dual-source human posture sensing and acquisition, 3D virtual-real space fusion rendering, and contactless human-computer interaction. Specifically, it relates to an AR glasses virtual-real rendering implementation scheme that integrates head inertial sensing and embedded eye-tracking sensors, supports configuration switching between single / dual sensor drive, and uses multiple independent 3D models as the sole rendering carrier. This invention can be applied to industrial and consumer-grade AR wearable work scenarios such as 3D inspection and annotation of industrial equipment, outdoor surveying and mapping, virtual-real training for vocational skills, and 3D model auxiliary display at engineering sites. Background Technology
[0002] Current AR glasses 3D model virtual-real rendering systems equipped with dual sensors have revealed numerous technical shortcomings in practical applications. Existing technical solutions cannot simultaneously address the following industry pain points: First, redundant material engineering leads to high production and maintenance costs. Most existing devices only support a single head-linkage mode or a single eye-tracking layer switching mode. To adapt to both control methods, technicians need to create two independent sets of 3D model engineering files, repeatedly modeling, debugging parameters, and packaging and publishing, significantly increasing content production, device storage, and post-maintenance costs. The industry lacks a standardized 3D model resource encapsulation architecture compatible with multiple control modes. Second, the model-driven form is singular, and the utilization rate of eye-tracking sensors is low. In traditional solutions, all 3D models are rigidly bound to the center of the viewport in head-driven mode, only achieving simple linear synchronous displacement and rotation, unable to achieve non-linear dynamic effects such as reverse movement and threshold-triggered deformation; eye-tracking functions are only used for basic operations such as menu selection and layer switching, unable to directly drive core rendering parameter adjustments such as scaling, rotation, transparency, and pose offset of 3D models, and the sensor functions have not been deeply developed. Third, forced synchronization of multiple model movements results in poor image performance. When multiple independent 3D models are loaded into the system, existing technologies apply the same posture response algorithm to all models. Different 3D models with different functions and forms cannot achieve layered, differentiated, and independent movement, resulting in a monotonous virtual-real fusion image and limited dynamic display effects. Fourth, dual sensors lack intelligent scheduling, leading to high device load. AR glasses equipped with head sensing and eye tracking dual modules generally employ a long-term, high-frequency, synchronous data collection mode for both sets of sensors, resulting in a consistently high overall computational load. Low-end wearable terminals are prone to rendering delays, model image jitter, and overall lag, while also exacerbating power consumption. Fifth, rigid model spatial constraints result in insufficient realism in virtual-real fusion. Most existing solutions rigidly anchor the 3D model to the glasses frame and the corresponding position of the wearer's face, preventing the model from freely floating and layering within the real-world 3D space. This violates the visual logic of real space and reduces the immersive AR experience. Sixth, limited interaction methods prevent hands-free operation. Existing systems heavily rely on handheld peripherals such as smartphones and tablets to adjust the pose and shape of 3D models. Operators cannot perform practical tasks such as assembly, maintenance, and surveying while manipulating the model, resulting in a lack of a purely wearable, hands-free interactive system. This invention presents a complete technical solution integrating a purely independent 3D model as the sole carrier, dual-channel independent pose mapping calculation, three-configuration intelligent switching, idle sensor hibernation to reduce load, complete decoupling of multiple model deformations, dual data stream fusion driving, and rigid anchor point-free 3D floating rendering. This combination of technical features fills a clear technological gap. Summary of the Invention
[0003] Purpose of the invention The purpose of this invention is to overcome the shortcomings of existing AR glasses, such as fixed 3D model rendering control modes, redundant 3D model materials with multiple versions, limited movement forms of head-bound models, inability of eye-tracking sensors to drive model deformation, forced synchronization of multiple 3D model movements, long-term high-load operation of dual sensors, and reliance on handheld peripherals. This invention provides a spatial virtual-real rendering method for AR glasses that uses multiple independent 3D models with configurable switching between dual sensors. This invention uses multiple independent 3D models as the sole rendering carrier, enabling free switching between three working modes: head-only driving, eye-tracking-only driving, and dual-channel collaborative driving. A single 3D model resource package is compatible with all operating configurations. Sensors are intelligently started, stopped, and put into sleep mode according to the operating mode, balancing device computing power and power consumption. Each 3D model is configured with independent deformation calculation logic, and the movement between models is completely decoupled. Combined with head movements, it achieves large-scale macroscopic control, while eye movements complete high-precision fine-tuning, enabling hands-free interaction throughout the process. This comprehensively improves the rendering stability, 3D image richness, and multi-scene adaptability of AR glasses. Technical solution
[0004] A method for spatial virtual-real rendering of AR glasses driven by dual sensors and multiple independent 3D models includes the following execution steps:
[0005] S01. Prepare independent 3D model files. Batch produce several independent 3D model files as the sole display carrier for AR virtual and real rendering. The 3D model includes at least one of mesh entity model, material texture model, and skeletal animation model. Each 3D model can be loaded, rendered, and receive posture control commands independently. There is no forced subordination or binding relationship between models.
[0006] S02. The system establishes two logically independent posture mapping systems: the first posture channel is a head-positioning posture mapping system, and the second posture channel is an eye-tracking posture mapping system. For each independent 3D model, dynamically adjustable rendering parameters are configured, including 3D world coordinates, overall scaling, Euler rotation angle, material transparency, and texture brightness. A many-to-many binding relationship is established between the 3D model and the two posture channels: a single 3D model can be bound to the first posture channel, the second posture channel, or both posture channels simultaneously. A single posture channel can mount any number and type of independent 3D models. Three standard linkage operation paradigms are defined: linear following mapping in the same direction, reverse displacement mapping, and posture threshold step deformation triggering. Each 3D model is configured with its own dedicated linkage operation logic, supporting combinations of multiple operation paradigms. All independent 3D model files, the dual-channel binding relationship, and the dedicated linkage operation logic for each 3D model are integrated and compressed using an encrypted format compatible with AR wearable devices, generating a unified resource package compatible with three operating modes.
[0007] S03. AR Glasses Hardware Analysis and Operation Configuration Scheduling: The AR glasses acquire the aforementioned resource package via wireless communication. The device has a built-in dedicated analysis and verification program that sequentially completes resource package decryption, decompression, file integrity verification, tamper identification and interception, ensuring secure resource loading. The AR glasses hardware integrates a first attitude sensing unit, a second attitude sensing unit, and an optical waveguide AR... The system comprises a perspective display unit and a main control scheduling chip. The first attitude sensing unit is a three-axis inertial head posture module used to collect head spatial motion data. The second attitude sensing unit is an eye-tracking camera embedded in the eyeglass frame used to collect eye movement feature data. The main control scheduling chip has three preset, freely switchable stable operating modes: Mode 1: Only the first head attitude sensing unit is activated to collect data at high frequency, while the second eye-tracking unit is powered off and enters a low-power sleep state; Mode 2: Only the second eye-tracking unit is activated to collect data at high frequency, while the first head attitude sensing unit is powered off and enters a low-power sleep state; Mode 3: The first and second attitude sensing units collect data synchronously and in parallel, working together to complete the driving operation. The main control chip issues power and scheduling commands according to the selected operating mode, waking up the working hardware modules and cutting off power to idle hardware modules. The device memory partition caches all channel mapping rules, and the real-time rendering thread only loads the computational logic corresponding to the currently active mode. Idle rules are stored in low-speed flash memory, not occupying real-time rendering computing power.
[0008] S04. Attitude Data Acquisition and Layered Virtual-Real Fusion Rendering: After starting the virtual-real rendering task, the awakened sensor units acquire raw attitude data. After filtering, noise reduction, and normalization conversion, a standardized floating-point attitude data stream is output. When running in Mode 1 (head-driven only): a single head attitude data stream is input into the system, and each independent 3D model iteratively updates its rendering parameters according to its own exclusive linkage calculation rules. When running in Mode 2 (eye-tracking only): a single eye-tracking attitude data stream is input into the system, and each independent 3D model iteratively updates its rendering parameters according to its own exclusive eye-tracking linkage calculation rules. When running in Mode 3 (dual-channel collaborative... (Simultaneous drive): Two attitude data streams are acquired independently without interference; only a 3D model with a single attitude channel is bound, and only the corresponding channel's data stream is responded to and the rendering parameters are updated; simultaneously, 3D models with two attitude channels are bound, and vector superposition and fusion operations are performed on the two attitude data streams to solve for the final set of rendering parameters of the model; the optical waveguide perspective display unit reads the final rendering parameters of each 3D model, and assigns a unique exclusive 3D space node to each 3D model in the real-world coordinate system, without applying any frame or rigid anchoring constraints to the face; multiple independent 3D models are layered and superimposed on the real-world image to complete the local output display of the virtual-real fusion image. Beneficial effects
[0009] This invention simplifies the resource system and significantly reduces production and maintenance costs by using only independent 3D models as rendering carriers, eliminating auxiliary materials such as videos and images. It adopts a single resource package architecture, where one set of 3D model resources can be compatible with three driving modes: head, eye-tracking, and dual-channel collaborative. This eliminates the need to create multiple sets of 3D model projects for different control modes, reducing the workload of modeling, parameter debugging, and version management, and significantly reducing the costs of content production, storage, and maintenance. At the same time, it forms a standardized 3D resource encapsulation system.
[0010] Intelligent hardware scheduling reduces device load and power consumption. This invention dynamically starts and stops idle sensors based on the operating configuration, changing the traditional working mode of long-term synchronous acquisition by dual sensors. This effectively reduces the computing load of AR glasses and avoids problems such as screen lag, model jitter, and rendering delay in low-end wearable terminals. At the same time, it reduces unnecessary power consumption and extends the device's battery life.
[0011] By deeply exploring sensor capabilities, the system enriches the dynamic effects of 3D models, breaking away from the traditional use of eye tracking for layer switching and menu clicking. It directly uses eye-tracking data to drive changes in the pose, scaling, transparency, and rotation of 3D models. It also supports various non-linear operation logics such as synchronous following, reverse offset, and threshold step deformation. Combined with macro-control of the head and fine-tuning of the eyeballs, it enables multiple 3D models to achieve differentiated dynamic performance, enhancing the sense of layering and expressiveness of AR images.
[0012] The models are completely decoupled, and the motion forms are flexible and diverse. Each independent 3D model is configured with different linkage operation logic. The deformation and pose changes between models are completely decoupled, completely eliminating the defects of forced synchronous motion of multiple models. The motion rules of each model can be flexibly set according to the use scenario, adapting to complex 3D display, annotation, and training scenarios.
[0013] By eliminating rigid anchoring, the realism of the virtual-real fusion is enhanced. All 3D models are rendered in independent 3D nodes in the coordinate system of the real world. They are no longer rigidly bound to the frame and face. The models can float freely and be arranged in layers in the real space, which fits the visual logic of the real space and greatly enhances the AR immersive experience.
[0014] Purely motion-sensing, hands-free interaction, adapted to industrial operation scenarios, relies entirely on head and eye movements to control all 3D models, eliminating the need for handheld peripherals such as mobile phones and tablets, completely freeing the operator's hands. It is perfectly suited for operation scenarios that require hands-free operation, such as industrial maintenance, equipment assembly, outdoor surveying, and skills training. At the same time, it can seamlessly switch to another operating mode when a single sensor fails, ensuring strong operation continuity and high reliability, and possessing value for large-scale commercialization.
[0015] With its novel technical architecture, this invention circumvents existing patent barriers by limiting the use of a single independent 3D model as the sole rendering medium. It combines a combined architecture of three-configuration switching, dual independent posture mapping, multi-model decoupling operation, dual-channel data fusion, intelligent sensor sleep mode, and rendering without rigid anchor points. This is different from existing traditional technical solutions that include multiple media such as images and videos, single fixed control mode, forced model synchronization, and constant sensor operation. It has outstanding inventiveness and can effectively circumvent the scope of existing patent protection. Detailed Implementation
[0016] The present invention will be further described in detail below with reference to specific embodiments. The following embodiments are only used to explain the technical solutions of the present invention and do not constitute a limitation on the scope of protection of the present invention.
[0017] Example 1: Reverse linkage rendering of two independent 3D models in head-only single-drive mode
[0018] S01. Create a 3D model of the first mesh sphere and a 3D model of the second mesh box. The two models are independent of each other and have no binding relationship.
[0019] S02. Bind both 3D models to the first head posture channel, configure reverse displacement offset calculation logic for the sphere model, and configure another set of reverse displacement offset calculation logic for the box model. The calculation parameters of the two types of models are independent of each other. Integrate the two 3D models, channel binding relationship, and exclusive linkage logic and seal them into a unified resource package to generate a unified resource package.
[0020] S03: The AR glasses load and verify the resource package, switch to head-only drive mode, and the eye-tracking unit enters sleep mode; the main control chip only powers the head posture sensing unit and starts data acquisition.
[0021] S04. Start virtual and real rendering. The head posture sensing unit collects head rotation and displacement data, and outputs a standardized posture data stream after processing. The sphere model and the box model complete independent reverse floating translation in the real scene 3D space according to their own exclusive logic. Both models are rendered on independent nodes in the real scene coordinate system and are not rigidly bound to the glasses frame or the wearer's face. Finally, the virtual and real blended picture is output.
[0022] Example 2: Composite Deformation Rendering of Skeletal Animation 3D Models in Dual-Channel Collaborative Driving Mode
[0023] S01. Create a 3D model of the ribbon skeleton animation as the sole rendering medium.
[0024] S02. Bind the skeletal animation model to the first head pose channel and the second eye pose channel simultaneously; set the head pose data stream to control the overall three-dimensional spatial coordinates of the model, and the eyeball vertical gaze offset data stream to control the rotation angle of the model itself; configure the vector fusion operation logic of the two data streams; and seal the model file, binding relationship, and fusion operation logic to generate a resource package.
[0025] S03. The AR glasses load and verify the resource package, switch to the dual-channel collaborative drive mode, and the head sensing unit and eye movement capture unit start working simultaneously.
[0026] S04. Start virtual and real rendering. Two sensors simultaneously collect posture data and generate standard data streams. When the wearer turns their head and moves their eyes up and down, the system integrates the two data streams to calculate the final rendering parameters of the model, so that the ribbon model simultaneously produces a composite deformation of overall spatial displacement and its own rotation. The model is suspended on independent nodes in the real scene space without rigid anchoring constraints, and the image output is completed.
[0027] Example 3: Multi-model Differential Threshold Deformation Rendering in Eye-tracking Single-Drive Mode
[0028] S01. Create three 3D models of mechanical parts of different specifications, with each model being independent of the others.
[0029] S02. Bind the three models to the second eye-tracking attitude channel, configure synchronous following logic for the first model, reverse offset logic for the second model, and attitude threshold step deformation logic for the third model; after integrating all files and logic, seal the package.
[0030] S03, the AR glasses switch to eye-tracking-only mode, the head posture sensing unit goes into sleep mode, and only the eye-tracking capture unit works.
[0031] S04. When the wearer moves their eyes, the eye movement data stream drives the three models to move independently according to their own logic: the first model moves synchronously with the gaze, the second model moves in the opposite direction, and when the eye movement data reaches a set threshold, the third model triggers a step deformation; the movement of each model does not interfere with each other, and all of them are suspended in the real-world 3D space to achieve differentiated rendering and display. Attached Figure Description
[0032] Figure 1 This is a schematic diagram of independent 3D model linkage in eye-tracking single mode of the present invention.
Claims
1. A spatial virtual-real rendering method for AR glasses driven by dual sensors and multiple independent 3D models, characterized in that, The process includes the following steps: Step 1: Prepare several independent 3D model files as display carriers for AR virtual-real rendering; Step 2: Construct two independent posture channel mapping systems corresponding to head movement and eye movement, establish a many-to-many binding relationship between each 3D model and the two types of posture channels, configure exclusive linkage operation logic between posture data stream and spatial rendering parameters for each 3D model, and integrate and package the 3D model files, channel binding relationships, and linkage operation logic into a unified resource package; Step 3: Equip the AR glasses with head posture sensing components and eye movement tracking sensing components, and configure three switchable operating modes: enable the head posture channel alone, enable the eye movement posture channel alone, and run the dual channels synchronously and collaboratively. Unused sensing components enter a low-power sleep state; Step 4: Start the virtual-real rendering task. The enabled sensing components output standardized posture data streams, and drive the 3D model to update the rendering state according to the binding relationship and linkage operation logic; at the same time, the 3D models bound to the two types of channels fuse the two data streams to solve the final rendering parameters, and the AR glasses display unit outputs a fused image of the real scene and the 3D model.
2. The method according to claim 1, characterized in that, The linkage operation logic includes at least one of the response forms of synchronous following, reverse offset, and threshold step deformation; multiple 3D models under the same attitude channel are configured with different linkage operation logics, and the deformation iteration processes of each model are decoupled from each other without forced synchronization constraints.
3. The method according to claim 1, characterized in that, The head posture sensing component collects spatial motion data of the wearer's head, and the eye movement capturing sensing component collects characteristic data of the wearer's eye movements.
4. The method according to claim 1, characterized in that, The wearer can adjust the shape and pose of the 3D model through at least one of head movements or eye movements, without needing to hold any external touch-screen interactive device.
5. The method according to claim 1, characterized in that, The resource package adopts an encrypted encapsulation structure, and the AR glasses have a built-in parsing and verification program to realize resource package integrity verification, tamper detection and secure loading.
6. The method according to claim 1, characterized in that, AR glasses render each 3D model to its own 3D spatial position in the coordinate system of the real world, and all 3D models are not rigidly anchored to the glasses frame or the wearer's face.
7. The method according to claim 1, characterized in that, The 3D model includes at least one of the following: a mesh solid model, a model with texture mapping, and a skeletal animation model.
8. The method according to claim 1, characterized in that, A single 3D model can be bound to a single attitude channel or two attitude channels simultaneously; a model bound to only a single channel only responds to the corresponding data stream, while a model bound to both attitude channels performs vector superposition and fusion operations on the two attitude data streams.
9. The method according to claim 1, characterized in that, The AR glasses main control module partitions and caches the channel mapping rules according to the current operating mode. Only the calculation logic corresponding to the current active mode is loaded into the real-time rendering thread, and the idle rules are stored in the low-speed storage area, which does not occupy the real-time computing power.