A digital glasses touch interaction method based on real-time hand image virtualization mapping

CN122526424APending Publication Date: 2026-08-07BEIJING YUHUANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING YUHUANG TECHNOLOGY CO LTD
Filing Date
2026-05-30
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

本发明的目的在于克服现有XR智能眼镜手势交互依赖预设手势、学习成本高、交互自由度低、可视化反馈缺失、对位精度差、体验割裂的技术缺陷,提供一种基于实时手部影像虚拟化映射的数字眼镜触控交互方法,彻底摒弃传统XR专属交互规则,实现无预设手势、全原生触屏行为复刻、虚实同步可视化交互,大幅提升数字眼镜人机交互的自然度、精准度与普适性

Benefits of technology

[0010] First, it completely eliminates the preset gesture system. This invention does not build any gesture template library, does not perform palm state determination, and does not perform micro-gesture matching. It abandons the "collection-comparison-matching" recognition logic that has been used in the industry for many years. All free air operations that conform to common touch screen habits can be recognized and responded to, with no upper limit to the degree of freedom of interaction, which is different from all preset gesture-based interaction solutions.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This invention discloses a digital glasses touch interaction method based on real-time hand image virtualization mapping, belonging to the field of smart wearable XR human-computer interaction technology. Addressing the industry pain points of existing XR smart glasses, which generally employ a fixed interaction mode of preset gesture template matching, fingertip fixed-point cursor control, and virtual touchpad mapping, resulting in high learning costs, low interaction freedom, lack of visual feedback, and disconnect from common touchscreen operating habits, this invention innovatively proposes a virtual-real synchronous interaction paradigm without preset gestures and with full native touch replication. This invention uses multi-dimensional vision and depth sensors to collect real-time, full-domain spatial motion data of the user's hand. After data cleaning and bidirectional calibration of the virtual and real coordinate systems, any free touch actions in the air that conform to touchscreen logic, such as clicking, long pressing, swiping, dragging, and two-finger opening and closing, are mapped and rendered in real-time as a three-dimensional virtual hand model superimposed on the lens interface. Users can simultaneously view their real hand, the virtual display interface, and the synchronously moving virtual hand entity. Interaction commands are analyzed based on the real overlap, touch, and dwell relationships between the virtual hand and interface controls, achieving a WYSIWYG interaction effect in the air that is completely consistent with the physical touchscreens of mobile phones and tablets. This invention completely abandons the traditional XR device-specific interaction rules, has no preset gesture library or template matching mechanism, and its interaction logic, visualization architecture and technical implementation path are all independently original. It has the advantages of zero learning cost, high alignment accuracy, low accidental touch rate, strong versatility and good implementation, and is suitable for various AR and VR digital smart glasses terminal devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of smart wearable devices, AR / VR digital glasses human-computer interaction, machine vision hand recognition, and virtual-real fusion rendering technology, specifically to a digital glasses touch interaction method based on real-time hand image virtualization mapping. Background Technology

[0002] Currently, the gesture interaction solutions for mainstream XR digital smart glasses and immersive head-mounted displays on the market mainly fall into four technical architectures: preset gesture template matching and recognition, fingertip pinch cursor control, biosensor signal recognition, and virtual touchpad planar mapping. All of these existing technologies have inherent and unavoidable technical flaws.

[0003] First, existing technologies all rely on device manufacturers' preset exclusive gestures and palm posture templates, requiring users to memorize and learn exclusive interactive actions. This cannot be adapted to the already widespread physical touch screen operation habits of mobile phones and tablets, resulting in a serious disconnect between human and computer interaction, high learning costs, and great difficulty in popularization.

[0004] Secondly, the existing interaction solutions have extremely low freedom. The device can only respond to limited gestures within the preset template. Users' free and casual touch operations cannot be recognized by the device. The interaction scenarios and operation logic are strictly limited, resulting in extremely poor adaptability.

[0005] Furthermore, traditional air gestures mostly only output a single cursor or fingertip mark, without complete visual feedback of the hand. Users cannot accurately determine the alignment relationship between their hand and the interface controls, which easily leads to problems such as alignment misalignment, operation failure, and mis-touch and misjudgment. The intuitiveness and accuracy of the interaction are poor.

[0006] Finally, existing technologies in the industry are all developed around the fixed ideas of "simplifying hand movements, fixed-point tracking, and template comparison". There is no interactive solution that can completely replicate all physical touch screen operations, has no preset gesture restrictions, and adopts a complete virtual whole hand entity visualization mapping. There is a long-term technological gap, serious industry homogenization, and lack of innovation.

[0007] In summary, existing XR gesture interaction solutions are rigid, unintuitive, have low flexibility, and poor user experience, failing to meet the public's demand for lightweight, learning-free, and highly flexible air-based touch interaction. There is an urgent need for a digital glasses interaction method with a completely new architecture, original logic, and that conforms to native operating habits. Summary of the Invention

[0008] 1. Purpose of the invention The purpose of this invention is to overcome the technical shortcomings of existing XR smart glasses gesture interaction, such as reliance on preset gestures, high learning costs, low degree of interaction freedom, lack of visual feedback, poor alignment accuracy, and fragmented experience. It provides a digital glasses touch interaction method based on real-time virtual mapping of hand images, which completely abandons the traditional XR-specific interaction rules, realizes no preset gestures, fully native touch screen behavior replication, and virtual-real synchronized visual interaction, and greatly improves the naturalness, accuracy and universality of human-computer interaction of digital glasses.

[0009] After a comprehensive search and comparison of existing patents and publicly available technical information related to XR gesture interaction, virtual touch, hand tracking, and virtual-real rendering, it was found that the overall architecture, core principles, interaction logic, and visualization of this invention are all independently original and completely heterogeneous with all existing technologies. It possesses absolute originality and a complete intellectual property barrier. Its core differentiating features are as follows:

[0010] First, it completely eliminates the preset gesture system. This invention does not build any gesture template library, does not perform palm state determination, and does not perform micro-gesture matching. It abandons the "collection-comparison-matching" recognition logic that has been used in the industry for many years. All free air operations that conform to common touch screen habits can be recognized and responded to, with no upper limit to the degree of freedom of interaction, which is different from all preset gesture-based interaction solutions.

[0011] Secondly, it features a unique virtual whole-hand entity mapping architecture. Unlike the planar feedback forms of existing technologies such as single-point cursors, fingertip markers, and virtual touchpads, this invention uses a complete 3D virtual whole-hand model to replicate the real hand in real time. Interaction is achieved through the real touch relationship between the virtual hand entity and interface controls, rather than fixed-point floating positioning. The interaction logic is more in line with the essence of physical touch screens.

[0012] Third, it achieves full migration of native touch behavior. It overturns the exclusive interaction logic of XR devices, replicating all operations of physical touch screens such as click, long press, swipe, drag, and two-finger pinch to zoom, without requiring users to adapt or learn. It eliminates the sense of disconnect between XR interaction and the public's conventional operating habits from the root, and there is no similar technical solution.

[0013] Fourth, the recognition and rendering principles are completely original. It abandons the traditional logic of fingertip point comparison, muscle signal acquisition, predictive touch detection, and trajectory template matching. It takes "full-domain hand posture trajectory acquisition + three-dimensional / two-dimensional coordinate bidirectional real-time mapping + virtual whole hand synchronous rendering + touch behavior feature identification" as the core technical path. The overall technical system has not been disclosed in any existing patents.

[0014] The digital glasses hardware used in this invention is equipped with a front-facing high-definition visual camera module, an infrared imaging module, and a high-precision depth sensor to construct a full-domain recognition field of view covering the interactive space in front of the glasses. It can collect raw data such as the three-dimensional spatial position of the user's hand, palm contour, five-finger posture, fingertip coordinates, and dynamic motion trajectory in real time and continuously.

[0015] The device has a built-in self-developed interactive algorithm that performs real-time filtering and noise reduction, motion stabilization, and spatial coordinate calibration on the raw hand data. This removes invalid interference data such as unconscious micro-movements, random shaking, and hovering drift, and accurately filters out effective touch actions initiated by the user that conform to general touch screen logic.

[0016] The system pre-calibrates the coordinate binding between the 3D hand recognition space and the 2D interface display space, establishing a fixed mapping relationship. It then converts the real hand's 3D spatial coordinates into precise pixel coordinates on the lens interface in real time. Based on the converted coordinate data, a 3D virtual hand model that is completely synchronized with the real hand's movements, postures, and positions is overlaid and rendered in real time on the top layer of the lens display interface.

[0017] When users wear the device, they can simultaneously view their real hand, the virtual display interface, and the virtual whole hand, creating a "what you see is what you get" visual effect that links the virtual and real worlds. The system analyzes the overlap, touch, pause, and displacement states of the virtual whole hand and various control elements on the interface in real time, triggering corresponding touch commands such as click, long press, swipe, drag, and zoom, ultimately achieving free touch interaction in the air, indistinguishable from a physical touchscreen.

[0018] To balance interactive freedom and operational stability, this invention employs an original adaptive dynamic interference filtering mechanism, distinct from traditional single-threshold filtering schemes. Based on a massive dataset of typical touchscreen operation samples, this invention constructs a native touch behavior feature model, comprehensively judging from multiple dimensions including movement speed, trajectory continuity, movement amplitude, and effective interaction space. It only identifies active, continuous, and regular user touch operations, automatically filtering out unconscious, irregular, and transient hand movements and micro-motions. Without restricting user freedom of operation, this significantly reduces the probability of accidental touches and improves interactive stability. Detailed Implementation

[0019] Wearing the digital glasses of this invention, the device automatically completes virtual-real coordinate calibration and sensor module initialization upon startup, collecting real-time dynamic data of the user's hand across the entire field of vision. Following conventional smartphone touchscreen operation habits, the user performs a light tap gesture towards an application icon in mid-air. The device renders a synchronized virtual hand model in real-time, with the virtual fingertip precisely touching the corresponding icon control. The system recognizes the tap and triggers an application opening command. When the user performs a long press gesture on an interface control in mid-air, the virtual hand remains continuously aligned with the corresponding interface position. The system recognizes the stable long-press behavior and triggers extended functions such as pop-up windows and multi-select activation. The entire process requires no memorization of specific gestures or precise point-based operation, completely replicating daily touchscreen usage habits, with stable recognition, intuitive feedback, no delay, and no accidental touches.

[0020] Within the device's effective interactive space, users can perform standard touchscreen actions such as horizontal swiping and vertical dragging in the air. The virtual hand model synchronously follows the hand's continuous movement. The system recognizes the continuous swiping trajectory and performs corresponding operations such as switching application pages, dragging video playback progress, and moving interface icons. In the parameter adjustment interfaces such as volume, screen brightness, and color temperature, users can achieve continuous and smooth adjustments by dragging vertically in the air. The interactive feel and operation logic are completely consistent with the physical touchscreen, allowing for free and unrestricted operation.

[0021] When users browse images, 3D models, and panoramic content, they retain the native touchscreen two-finger zooming habit, freely opening and closing their fingers in the air. The device collects real-time data on the finger posture and spacing changes, synchronously updating the virtual two-finger opening and closing state. Based on the change in the spacing between the virtual two fingers, the system accurately identifies zoom-in and zoom-out commands, completing stepless scaling adjustments of the interface content. This interaction method has no restrictions on movement rules, allowing users to operate freely and adapting to immersive browsing scenarios.

[0022] 1. High originality barrier and independent intellectual property rights. The core combination architecture of this invention, which combines "no preset gestures + virtual whole hand entity visualization + full-scale aerial mapping of native touch screen behavior", has not been disclosed in any publicly available patents or technical documents. The overall interaction paradigm, technical principles, and implementation path are completely independent and original, with no homogenization or infringement risks, and the boundaries of intellectual property rights are clear and solid.

[0023] 2. Zero learning cost and extremely universal applicability. It completely overturns the traditional XR device-specific interaction logic, replicates the touch screen operation habits of mobile phones and tablets, requires no training or memorization of gestures, and can be used proficiently from the start, fundamentally solving the pain points of the fragmented interaction experience and high learning threshold of traditional XR.

[0024] 3. Intuitive and precise interaction, providing an excellent user experience. Abandoning the traditional single-point cursor with vague feedback, it adopts a complete virtual hand entity visualization overlay, allowing users to intuitively observe the touch alignment relationship. What you see is what you get, completely solving the problems of inaccurate alignment, vague feedback, and operation failure in air touch control, significantly improving interaction accuracy.

[0025] 4. High degree of freedom in interaction and good feasibility. There are no preset gestures or fixed action restrictions. It supports all compliant free touch-based hand movements of users and is compatible with interface interactions in all scenarios. The hardware relies on conventional vision and depth sensors, without the need for dedicated custom hardware. The structure is simple, the cost is controllable, and it can be quickly deployed for mass production.

[0026] 5. Strong anti-interference capability and excellent stability. Equipped with a multi-dimensional adaptive anti-mistouch mechanism, it comprehensively identifies effective operations based on trajectory, speed, and spatial range, accurately distinguishing between active operations and unintentional micro-movements and shakes. It balances ultimate interactive freedom with high stability, providing a user experience far superior to traditional restricted gesture interaction solutions.

Claims

1. A digital glasses touch interaction method based on real-time hand image virtualization mapping, characterized in that, Includes the following steps: S100 uses the multi-dimensional vision sensor and depth sensing module mounted on the digital glasses to collect real-time, full-domain dynamic images, three-dimensional spatial coordinates, fingertip point distribution and continuous motion trajectory raw data of the user's hand in front of the glasses. S200 performs filtering, noise reduction, motion stabilization, and spatial coordinate calibration on the collected raw hand data to remove invalid interference data such as unconscious hand micro-movements, random shaking, and hovering drift, and accurately extracts active hand touch actions that conform to the general touch screen operation logic. S300: Pre-construct the binding mapping relationship between the two-dimensional interface coordinate system of the eyeglass lens display interface and the three-dimensional spatial coordinate system of the hand recognition space in front of the eyeglasses, and convert the real three-dimensional spatial coordinates of the hand into two-dimensional interface touch coordinates in real time. S400: Based on real-time coordinate mapping results, a three-dimensional virtual whole hand model is overlaid and rendered on the lens display interface, which is completely synchronized with the real hand posture, position and trajectory. S500 executes corresponding touch control commands based on the spatial overlap, touch coverage, dwell time, and displacement trajectory linkage between the virtual whole hand model and the interface APP control elements to complete the digital glasses interface interaction; This method has no preset gesture library, no gesture template matching, no exclusive customized gesture judgment, and does not rely on single-point cursor tracking and eye-tracking pre-selection assistance. It achieves free interaction entirely based on the native operating habits of the general public's physical touch screen, and is completely heterogeneous with the existing XR preset gesture, fingertip cursor, and virtual touchpad interaction architecture.

2. The digital glasses touch interaction method based on real-time hand image virtualization mapping according to claim 1, characterized in that, The active hand touch actions include single-point clicking in the air, multi-point continuous clicking, long press and hold, straight-line sliding, curved dragging, and two-finger pinch-to-zoom, which users can operate freely without fixed action restrictions.

3. The digital glasses touch interaction method based on real-time hand image virtualization mapping according to claim 1, characterized in that, The three-dimensional virtual whole hand model is a complete three-dimensional solid model that replicates the posture of the five fingers, the outline of the palm, and the spatial position of the real hand. Unlike single-point cursors, fingertip markers, and virtual touchpad layers, the virtual whole hand model synchronizes with the full range of motion details of the real hand in real time, with no deviation in posture and trajectory.

4. The digital glasses touch interaction method based on real-time hand image virtualization mapping according to claim 1, characterized in that, The visualization overlay mechanism in step S400 is as follows: When the user is wearing glasses, they can simultaneously observe the real hand entity, the virtual display interface of the lens, and the virtual whole hand model overlaid on the top layer of the interface. Through the real-time touch state between the virtual whole hand and the interface controls, they can intuitively obtain touch alignment feedback.

5. The digital glasses touch interaction method based on real-time hand image virtualization mapping according to claim 1, characterized in that, The invalid data filtering mechanism in step S200 includes: combining hand movement speed threshold, trajectory continuity features, and effective interaction space range limitation to construct a native touch behavior feature model, responding only to hand movements that conform to the regular active touch screen operation logic, filtering out unconscious invalid movements, and avoiding accidental triggering.

6. The digital glasses touch interaction method based on real-time hand image virtualization mapping according to claim 1, characterized in that, The interactive response logic of step S500 specifically includes: When a real hand clicks in the air, the virtual hand simultaneously completes the fingertip touch control action, triggering interface selection, application opening, and function confirmation commands; When a real hand is detected to hold a long press in the air, the virtual hand continues to fit the corresponding control position, triggering pop-up windows, multi-select activation, and long press extension functions. When a real hand is detected sliding or dragging in the air, the virtual hand moves synchronously to trigger page switching, playback progress adjustment, and interface control repositioning. When the system recognizes the opening and closing motion of two fingers on a real hand, the virtual two fingers simultaneously adjust the spacing, triggering an overall scaling adjustment of the image, model, and interface.