Mixed reality intelligent interaction method and device fusing eye movement tracking and gesture recognition

By integrating multimodal data synchronization and adaptive interaction methods that combine eye tracking and gesture recognition, the problems of insufficient interaction accuracy and user fatigue in existing technologies are solved, achieving an efficient and personalized mixed reality interactive experience.

CN121541780BActive Publication Date: 2026-07-21CHENGDU YUNWANG PERCEPTION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU YUNWANG PERCEPTION TECHNOLOGY CO LTD
Filing Date
2025-11-19
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing mixed reality technologies, eye tracking and gesture recognition have not been deeply integrated, resulting in insufficient interaction accuracy, inability to adapt to different users' operating habits and diverse scenarios, easy occurrence of accidental touch operations, and lack of personalized adaptation, leading to increased user fatigue.

Method used

By collecting basic user interaction data, generating a dedicated configuration file, and combining multimodal data synchronization, we construct object, environment, and user state contexts, calculate adaptive gaze thresholds, use gaze gestures for two-way verification to confirm interaction intent, activate a three-level gesture semantic library, and optimize the interaction process through a multimodal feedback mechanism.

Benefits of technology

It achieves deep collaboration between eye tracking and gesture recognition, improving the accuracy and naturalness of interaction, reducing accidental touches, adapting to individual differences, reducing user fatigue, and providing an efficient and personalized mixed reality interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121541780B_ABST
    Figure CN121541780B_ABST
Patent Text Reader

Abstract

The application discloses a mixed reality intelligent interaction method and device combining eye movement tracking and gesture recognition, relates to the technical field of intelligent interaction, calls an exclusive configuration file, and calculates an adaptive fixation threshold in combination with a current scene; when the duration for which a user fixes a target object is greater than or equal to the adaptive fixation threshold, and a hand enters a preset interaction region, an interaction context is pre-activated; after bidirectional verification and confirmation of an interaction intention through a fixation gesture, a complete interaction context is activated, and a corresponding gesture semantic library is loaded. The application can provide a mixed reality intelligent interaction method and device combining eye movement tracking and gesture recognition, realizes deep cooperation between eye movement tracking and gesture recognition through multi-modal data fusion and three-dimensional context construction, and significantly improves the accuracy and naturalness of mixed reality interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent interaction technology, and in particular to a mixed reality intelligent interaction method and device that integrates eye tracking and gesture recognition. Background Technology

[0002] Mixed reality technology serves as a crucial bridge connecting the virtual and real worlds, with its core value lying in enabling intuitive and natural human-computer interaction. Eye tracking and gesture recognition are two key input methods in mixed reality interaction and are widely used in various mixed reality devices.

[0003] In existing technologies, eye tracking and gesture recognition often operate as independent input channels in parallel, failing to form a deeply integrated interaction logic. These solutions typically use fixed gaze thresholds to determine user interaction intent, which cannot adapt to different user operating habits and diverse scenarios. Gesture semantics are mostly globally fixed settings, not adjusting with changes in the interaction object and environment, resulting in insufficient interaction accuracy.

[0004] Meanwhile, existing solutions lack comprehensive consideration of user status and environmental information, making accidental touch operations easy to occur. Furthermore, the interaction process lacks personalized adaptation, which can easily cause user fatigue after prolonged use, making it difficult to meet the needs of mixed reality applications for efficient, accurate, and natural interaction. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a mixed reality intelligent interaction method and device that integrates eye tracking and gesture recognition. The technical solution adopted is as follows: A mixed reality intelligent interaction method that integrates eye tracking and gesture recognition includes the following steps: Step 1: Collect basic user interaction data to complete personalized calibration and generate a unique configuration file; Simultaneously collect eye-tracking data, gesture data, environmental data, and user status data, and use a timestamp alignment algorithm to achieve multimodal data synchronization and obtain multimodal data; Step 2: Construct and fuse object context, environment context, and user state context based on multimodal data; Step 3: Call the dedicated configuration file and calculate the adaptive gaze threshold based on the current scene; When the duration of a user's gaze at the target object is greater than or equal to the adaptive gaze threshold, and the hand enters the preset interaction area, the interaction context is pre-activated. After confirming the interaction intent through two-way verification of gaze gestures, the complete interaction context is activated and the corresponding gesture semantic library is loaded. Confirming interaction intent through two-way verification of gaze gestures specifically includes: Step 31: After gaze lock, pre-activate the interactive context and monitor pre-gesture actions; Step 32: Collect pre-action features and calculate the matching degree with the available operations of the object; Step 33: When the matching degree meets the standard, the gaze focus accuracy is corrected through pre-action; when the matching degree is insufficient, the user is prompted to adjust. Step 34: After the matching degree is stable and meets the standard, the interaction intent is formally confirmed, the complete interaction context is activated, and the corresponding gesture semantic library is loaded. Step 4: Based on the activated interactive context, load the corresponding basic, extended, and custom three-level gesture semantic library, recognize user gestures and match operation commands, and provide dynamic feedback on the interaction results through visual, tactile, and audio multimodal feedback.

[0006] Optionally, the personalized calibration in step 1 includes gaze point accuracy calibration, gesture habit sampling, and visual sensitivity settings; the dedicated configuration file includes dynamic threshold benchmark values, gesture recognition model parameters, feedback preferences, and user operation history data, and supports cross-device synchronization.

[0007] Optionally, the eye-tracking data includes pupil position, 3D coordinates of the fixation point, gaze movement speed, and blink frequency; The gesture data includes the coordinates of 26 key skeletal points of the hand, hand posture, movement trajectory, movement speed, and electromyographic signals. The environmental data includes a 3D environment map constructed by SLAM, ambient light intensity, spatial dimensions, and obstacle distribution; The user status data includes heart rate data, blink frequency data, electromyography signal data, and operation history data.

[0008] Optionally, step 2, 3D context fusion, specifically includes: Object context: Identifies the type of gaze target, available operations, current state, and valid region parameters; Environmental context: Analyze ambient light intensity levels, available operational space, and obstacle distribution; User state context: Determine the user's fatigue level based on heart rate and blink frequency, determine the user's operation proficiency based on operation history data, and determine the hand force state based on electromyography signals.

[0009] Optionally, the adaptive gaze threshold in step 2 is calculated using the following formula: Where DT is the dynamic threshold, D is the baseline value, and a is the scene coefficient. The scene coefficient is dynamically adjusted according to the ambient light intensity level, user fatigue level, and operation proficiency, with an adjustment range of 0.8-1.5. The preset interaction area is a three-dimensional space range of 30-80cm in front of the user's chest; the interaction area can be adaptively adjusted according to the user's height, range of limb movement and current space size.

[0010] Optionally, step 2, the two-way verification of gaze gestures, specifically includes the following sub-steps: Step 31a: After the pre-activation of the interactive context, start the gesture pre-action monitoring to capture pre-action features, including finger micro-movements and wrist rotations. Step 31b: Collect the finger bending angle, starting direction of the movement trajectory, and electromyographic signal intensity features of the pre-action, compare them with the gesture feature template library corresponding to the available operations of the current object, and calculate the matching degree through a multi-dimensional weighted fusion algorithm; Step 31c: If the matching degree is greater than or equal to 80%, the gaze focus coordinates are corrected by the pre-action trajectory so that the focus falls accurately within the effective area of ​​the object; if the matching degree is greater than or equal to 50% but less than 80%, the user is prompted to adjust the gesture through weak visual feedback; if the matching degree is less than 50%, the pre-activated state is maintained and monitoring continues. Step 31d: When the matching degree remains at or above 80% for more than 50ms, the interaction intent is formally confirmed and the complete interaction context is activated.

[0011] Optionally, the formula for the multi-dimensional weighted fusion algorithm is: Where M is the total matching degree, Weighted matching scores are calculated based on the basic features of the gesture. To annotate the focus correction factor, This refers to the scene adaptation coefficient.

[0012] Optionally, the three-level gesture semantic library mentioned in step 4 includes a basic gesture library, an extended gesture library, and a custom gesture library; The basic gesture library includes globally universal gestures, which include a slight fist clench, thumb and index finger swipe, and palm flip. The gesture library is extended to bind specific object types, including contextual gestures, such as drawing circles with two fingers when gazing at a 3D model and drawing lines with the index finger when gazing at text. The custom gesture library allows users to add exclusive gestures through recording, naming, and binding processes, and the recognition model is automatically optimized after binding.

[0013] Optionally, the multimodal dynamic feedback in step 3 specifically includes visual feedback, tactile feedback, and audio feedback; Visual feedback adaptively adjusts brightness based on ambient light intensity, and displays progress bars or numerical prompts during complex operations; The tactile feedback distinguishes between confirmation feedback, progress feedback, and warning feedback, and the vibration intensity adaptively adjusts according to the user's fatigue level. The audio feedback allows users to select sound effect types, and automatically reduces the volume of the set sound effect when the operation frequency exceeds a set threshold.

[0014] A mixed reality intelligent interaction device integrating eye tracking and gesture recognition is used to realize a mixed reality intelligent interaction method integrating eye tracking and gesture recognition. The mixed reality intelligent interaction device includes a multimodal perception module, an intelligent processing module, and a display feedback module. The intelligent processing module is communicatively connected to the multimodal perception module and controls the execution actions of the display feedback module. The multimodal perception module includes an eye-tracking unit, a gesture environment perception unit, and a user state perception unit; the eye-tracking unit, the gesture environment perception unit, and the user state perception unit are each communicatively connected to the intelligent processing module. The intelligent processing module includes a heterogeneous architecture main processor and an interactive processing chip; the main processor supports multimodal data parallel processing, and the dedicated interactive processing chip integrates three-dimensional context fusion, adaptive threshold calculation, hierarchical gesture recognition, and gaze gesture bidirectional verification algorithms. The display and feedback module includes a gesture environment sensing unit and a user state sensing unit.

[0015] In summary, the present invention has at least one of the following beneficial technical effects: This invention provides a mixed reality intelligent interaction method and device that integrates eye tracking and gesture recognition. Through multimodal data fusion and three-dimensional context construction, it achieves deep collaboration between eye tracking and gesture recognition, significantly improving the accuracy and naturalness of mixed reality interaction.

[0016] Adaptive gaze thresholds and personalized calibration mechanisms cater to different user habits and scenario needs, effectively reducing interaction discomfort caused by individual differences. A two-way gaze gesture verification mechanism significantly reduces accidental touches and improves interaction reliability. A three-level gesture semantic library balances universality and scenario-specific needs, enhancing interaction flexibility.

[0017] Multimodal dynamic feedback makes the interaction results intuitive and perceptible, and the overall solution reduces the user's cognitive load and operational fatigue, creating a mixed reality interactive experience that is more in line with the natural human interaction habits, providing strong support for the large-scale application of mixed reality technology in multiple fields. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the mixed reality intelligent interaction method that integrates eye tracking and gesture recognition according to the present invention. Detailed Implementation

[0019] The present invention will be further described in detail below with reference to the accompanying drawings.

[0020] This invention discloses a mixed reality intelligent interaction method and device that integrates eye tracking and gesture recognition.

[0021] Reference Figure 1Example 1, a mixed reality intelligent interaction method integrating eye tracking and gesture recognition, includes the following steps: Step 1: Collect basic user interaction data to complete personalized calibration and generate a unique configuration file; Simultaneously collect eye-tracking data, gesture data, environmental data, and user status data, and use a timestamp alignment algorithm to achieve multimodal data synchronization and obtain multimodal data; Step 2: Construct and fuse object context, environment context, and user state context based on multimodal data; Step 3: Call the dedicated configuration file and calculate the adaptive gaze threshold based on the current scene; When the duration of a user's gaze at the target object is greater than or equal to the adaptive gaze threshold, and the hand enters the preset interaction area, the interaction context is pre-activated. After confirming the interaction intent through two-way verification of gaze gestures, the complete interaction context is activated and the corresponding gesture semantic library is loaded. Confirming interaction intent through two-way verification of gaze gestures specifically includes: Step 31: After gaze lock, pre-activate the interactive context and monitor pre-gesture actions; Step 32: Collect pre-action features and calculate the matching degree with the available operations of the object; Step 33: When the matching degree meets the standard, the gaze focus accuracy is corrected through pre-action; when the matching degree is insufficient, the user is prompted to adjust. Step 34: After the matching degree is stable and meets the standard, the interaction intent is formally confirmed, the complete interaction context is activated, and the corresponding gesture semantic library is loaded. Step 4: Based on the activated interactive context, load the corresponding basic, extended, and custom three-level gesture semantic library, recognize user gestures and match operation commands, and provide dynamic feedback on the interaction results through visual, tactile, and audio multimodal feedback.

[0022] Example 2: The personalized calibration in step 1 includes gaze point accuracy calibration, gesture habit sampling, and visual sensitivity settings; the dedicated configuration file includes dynamic threshold benchmark values, gesture recognition model parameters, feedback preferences, and user operation history data, and supports cross-device synchronization.

[0023] Example 3, the eye movement data includes pupil position, 3D coordinates of the fixation point, gaze movement speed, and blink frequency; The gesture data includes the coordinates of 26 key skeletal points of the hand, hand posture, movement trajectory, movement speed, and electromyographic signals. The environmental data includes a 3D environment map constructed by SLAM, ambient light intensity, spatial dimensions, and obstacle distribution; The user status data includes heart rate data, blink frequency data, electromyography signal data, and operation history data.

[0024] By adopting the above technical solutions, personalized calibration is based on individual difference adaptation logic. It collects basic data such as user gaze point accuracy, gesture habits, and visual sensitivity to construct a unique configuration file, ensuring that interaction parameters are precisely matched with user physiological characteristics and operating habits, avoiding interaction discomfort caused by generic parameters. Multimodal data synchronization relies on a timestamp alignment algorithm to precisely calibrate the collection time of four types of data: eye movement, gestures, environment, and user state. This ensures data consistency in the time dimension, providing synchronous and complete data support for subsequent context construction and intent judgment, avoiding interaction delays or misjudgments caused by asynchronous data.

[0025] The 3D context construction follows the principle of comprehensive perception and deep fusion. It identifies the gaze target by analyzing eye-tracking data and extracts object context such as target type and available operations; it captures environmental context such as light intensity and spatial dimensions through environmental data; and it determines user state context such as fatigue level and proficiency through user state data. The fusion of these three types of context can comprehensively depict the complete information of the interaction scene. The adaptive gaze threshold is based on the principle of dynamic scene adaptation. Using a baseline value in a dedicated configuration file as a foundation, it dynamically adjusts the scene coefficient by combining ambient light intensity level, user fatigue level, and operation proficiency. This allows the gaze threshold to improve interaction efficiency in focused scenarios and reduce the risk of accidental touches in complex scenarios, achieving optimal interaction response in different scenarios.

[0026] The gaze-gesture two-way verification is based on the "double confirmation of intent" logic. Gaze serves as the initial guide to the interaction intent, while gesture serves as the precise verification of the intent, forming a two-way collaborative mechanism. First, gaze locks onto the target and pre-activates the interaction context, triggering gesture pre-action monitoring to ensure the targeted nature of the interaction. Then, by collecting key features of the pre-action and comparing them with feature templates of available operations on the object, a multi-dimensional weighted fusion algorithm quantifies the matching degree to achieve an initial determination of intent. Subsequently, focus correction or operation prompts are executed based on the matching degree results to optimize the accuracy of the interaction. Finally, a stable matching degree is used as the final condition for intent confirmation, ensuring the clarity of the interaction intent and fundamentally reducing the probability of accidental touches caused by single-modal triggering.

[0027] The layered gesture semantic library is based on a gradient adaptation principle of general-scenario-personalized. The basic gesture library meets global general operation needs, ensuring consistent interaction; the extended gesture library binds exclusive operations to specific object types, improving the efficiency of scenario-based interaction; the custom gesture library allows users to extend it independently, adapting to personalized needs and achieving a balance between versatility and flexibility. Multimodal feedback is based on the principle of multi-sensory collaborative perception. Visual feedback intuitively presents operation results, tactile feedback enhances the sense of operation confirmation, and audio feedback assists in indicating the interaction status. The dynamic adaptation of these three types of feedback, such as adjusting visual brightness according to ambient light intensity and adjusting tactile intensity according to fatigue level, can improve the intuitiveness and comfort of interaction and reduce the cognitive load on users.

[0028] Example 4, step 2 of the three-dimensional context fusion specifically includes: Object context: Identifies the type of gaze target, available operations, current state, and valid region parameters; Environmental context: Analyze ambient light intensity levels, available operational space, and obstacle distribution; User state context: Determine the user's fatigue level based on heart rate and blink frequency, determine the user's operation proficiency based on operation history data, and determine the hand force state based on electromyography signals.

[0029] Example 5: The adaptive gaze threshold in step 2 is calculated using the following formula: Where DT is the dynamic threshold, D is the baseline value, and a is the scene coefficient. The scene coefficient is dynamically adjusted according to the ambient light intensity level, user fatigue level, and operation proficiency, with an adjustment range of 0.8-1.5. The preset interaction area is a three-dimensional space range of 30-80cm in front of the user's chest; the interaction area can be adaptively adjusted according to the user's height, range of limb movement and current space size.

[0030] Example 6, step 2 of the gaze gesture two-way verification specifically includes the following sub-steps: Step 31a: After the pre-activation of the interactive context, start the gesture pre-action monitoring to capture pre-action features, including finger micro-movements and wrist rotations. Step 31b: Collect the finger bending angle, starting direction of the movement trajectory, and electromyographic signal intensity features of the pre-action, compare them with the gesture feature template library corresponding to the available operations of the current object, and calculate the matching degree through a multi-dimensional weighted fusion algorithm; Step 31c: If the matching degree is greater than or equal to 80%, the gaze focus coordinates are corrected by the pre-action trajectory so that the focus falls accurately within the effective area of ​​the object; if the matching degree is greater than or equal to 50% but less than 80%, the user is prompted to adjust the gesture through weak visual feedback; if the matching degree is less than 50%, the pre-activated state is maintained and monitoring continues. Step 31d: When the matching degree remains at or above 80% for more than 50ms, the interaction intent is formally confirmed and the complete interaction context is activated.

[0031] By adopting the above technical solution, the principle of refining the object context is to extract all target interaction attributes. After the system locks the gaze target by parsing eye movement data, it further extracts the target type, available operations, current state and effective area parameters, and clarifies the core attributes and interaction boundaries of the interactive object, providing a precise target benchmark for subsequent gesture semantic matching and focus correction.

[0032] The principle of refining the environmental context is to adapt and perceive key environmental parameters. It focuses on three core parameters: ambient light intensity level, available operating range, and obstacle distribution. This allows the system to clearly understand the environmental conditions and spatial limitations in which it interacts, providing an environmental basis for threshold adjustment, interaction area adaptation, and feedback method optimization.

[0033] The principle behind refining user state context is the quantification and determination of the user's real-time state. Fatigue level assessment criteria are established using heart rate and blink frequency, a proficiency evaluation model is built using historical operation data, and hand force state is captured using electromyography (EMG) signals. This allows the system to dynamically grasp the user's physiological state and operational capabilities, providing user-side support for personalized interaction parameter adjustments. The synergistic integration of these three types of contexts achieves full-dimensional scene perception across the object, environment, and user dimensions, providing complete data support for subsequent interaction decisions.

[0034] The adaptive gaze threshold achieves precise adaptation based on a quantification formula and a multi-factor dynamic weighting principle. The baseline value originates from a user-specific configuration file, representing personalized initial parameters aligned with the user's basic operating habits. The scene coefficient, as the core of dynamic adjustment, quantifies the influence weights of ambient light intensity (threshold lengthens in strong light, shortens in weak light), user fatigue level (threshold lengthens when fatigued, shortens when alert), and operational proficiency (threshold shortens when proficient, lengthens when novice), forming an adjustment range of 0.8-1.5. The formula DT=D×a organically combines the baseline value and the scene coefficient, enabling the gaze threshold to dynamically change according to the real-time scene. This ensures interaction efficiency in focused scenarios while avoiding false triggers in complex scenarios, achieving precise matching between scene, user, and threshold.

[0035] The core principle of the preset interaction area is to balance ease of operation with spatial adaptability. The basic range of 30-80cm is a comfortable operating zone determined by ergonomics, which conforms to the user's natural arm movement radius while avoiding fatigue caused by operating too close or too far. At the same time, the interaction area supports adjustments based on user height: a larger range for taller users and a smaller range for shorter users; and a smaller range for users with limited limb movement. It also adaptively adjusts to the current spatial dimensions to ensure that different users can obtain convenient and non-redundant interaction space in different spatial environments, further reducing the difficulty of operation.

[0036] After pre-activating the interactive context in step 31a, targeted monitoring of pre-action features such as finger micro-movements and wrist rotations is conducted to avoid invalid recognition of irrelevant hand movements, focus on initial actions related to the interaction, and narrow down the scope and improve efficiency for subsequent intent determination.

[0037] Step 31b collects three core features of the pre-action: finger bending angle, starting direction of movement trajectory, and electromyographic signal intensity. These features can accurately reflect the core intent of the gesture. After comparing them with the feature template library of available operations of the object, the intent is transformed into a quantifiable matching degree through a multi-dimensional weighted fusion algorithm, realizing the transformation of "intent from vague to clear" and ensuring the objectivity of the judgment.

[0038] Step 31c performs differentiated operations based on different matching degree ranges: when the matching degree meets the standard, the gaze focus is corrected through the pre-action trajectory, and the accuracy of the gesture is used to compensate for the slight deviation that may exist in the gaze, so as to achieve mutual calibration between gaze and gesture; when the matching degree needs to be optimized, the user is prompted to adjust through weak visual feedback, forming real-time interactive guidance; when the matching degree does not meet the standard, the pre-activated state is maintained to avoid misjudging as no intention, and to balance the flexibility and rigor of the verification.

[0039] The principle of step 31d is to verify the reliability of the interaction intent. The final condition for intent confirmation is that the matching degree is ≥80% and lasts for more than 50ms. This eliminates false triggers caused by instantaneous matching, ensures the stability and clarity of the user's interaction intent, and further reduces the probability of false touches from the process perspective, thereby improving the reliability of the interaction.

[0040] Example 7, the formula for the multi-dimensional weighted fusion algorithm is: Where M is the total matching degree, Weighted matching scores are calculated based on the basic features of the gesture. To annotate the focus correction factor, This refers to the scene adaptation coefficient.

[0041] By adopting the above technical solution, three core features—finder bending angle, starting direction of movement trajectory, and electromyographic signal intensity—are selected because these features most directly reflect the core intent of the gesture, avoiding interference from redundant features in the matching results. Weighting Dynamically assigning weights based on gesture type and interaction scenario, key features receive higher weights to ensure the matching results accurately match the current interaction needs. Feature Score The similarity of features is quantified by accurately comparing user pre-action features with a gesture feature template library. The weighted sum of the two results in a basic quantitative assessment of the degree of fit between the gesture itself and the target operation, providing a core basis for the overall matching degree.

[0042] The scene adaptation coefficient is used to offset the impact of environmental interference and changes in user state on feature recognition. This coefficient is dynamically generated based on 3D context information. When the ambient light intensity is suitable and the user is in a good state (awake and highly skilled), feature recognition accuracy is higher, and the coefficient provides a positive gain, optimizing the matching results. When the ambient light intensity is too strong or too weak, the user is fatigued, or the user's operational skill is low, feature recognition is easily interfered with. The coefficient is adjusted appropriately to correct the deviation in the matching results, avoiding misjudgments caused by scene factors. This correction allows the matching degree calculation to break through the limitations of single feature comparison and better reflect the complex scene changes in actual interaction.

[0043] The calculation principle of the overall matching degree M is "multi-dimensional information complementarity and fusion". First, the matching degree of the basic features of the gesture is quantified by weighted summation. Then, the coordination accuracy of gaze and gesture is calibrated by the gaze focus correction coefficient. Finally, the influence of environment and user state is compensated by the scene adaptation coefficient. The three layers of calculation logic are progressive and complementary. The final overall matching degree not only objectively reflects the degree of matching between gesture and target operation, but also fully considers key factors such as gaze accuracy and scene state, forming a comprehensive and accurate quantitative indicator, providing a scientific and reliable basis for intent determination in gaze-gesture two-way verification.

[0044] Example 8: The three-level gesture semantic library mentioned in step 4 includes a basic gesture library, an extended gesture library, and a custom gesture library; The basic gesture library includes globally universal gestures, which include a slight fist clench, thumb and index finger swipe, and palm flip. The gesture library is extended to bind specific object types, including contextual gestures, such as drawing circles with two fingers when gazing at a 3D model and drawing lines with the index finger when gazing at text. The custom gesture library allows users to add exclusive gestures through recording, naming, and binding processes, and the recognition model is automatically optimized after binding.

[0045] Example 9: In step 3, the multimodal dynamic feedback specifically includes visual feedback, tactile feedback, and audio feedback; Visual feedback adaptively adjusts brightness based on ambient light intensity, and displays progress bars or numerical prompts during complex operations; The tactile feedback distinguishes between confirmation feedback, progress feedback, and warning feedback, and the vibration intensity adaptively adjusts according to the user's fatigue level. The audio feedback allows users to select sound effect types, and automatically reduces the volume of the set sound effect when the operation frequency exceeds a set threshold.

[0046] By adopting the above technical solution, the basic gesture library selects simple and low-cognitive gestures such as micro-clenching of the fist, thumb-index finger sliding, and palm flipping as globally universal commands. These gestures do not depend on specific interaction objects or scenarios, and their semantics are fixed in all interaction contexts. This allows users to complete basic operations without having to learn additional contextualized gestures, lowering the initial usage threshold and ensuring basic consistency and convenience in interaction.

[0047] The extended gesture library strongly associates contextual gestures with specific object types. When the system recognizes the gaze target as a 3D model, it automatically loads gestures adapted to model operations, such as drawing circles with two fingers; when the gaze target is text, it loads gestures adapted to text operations, such as drawing lines with the index finger. This binding allows gesture semantics to dynamically change with the interactive object, avoiding operational redundancy caused by globally fixed semantics, improving the efficiency and accuracy of contextualized interaction, and achieving a natural mapping of "what object you look at, what gesture you use".

[0048] The custom gesture library follows a process that allows users to add their own unique gestures through recording, naming, and binding. The recording phase captures the feature data of the user's personalized gestures, while the binding phase associates the gestures with the target operation. After binding, the system automatically integrates the new gesture features into the recognition model and optimizes model parameters to improve recognition accuracy. This design breaks the limitations of preset gestures, allowing users to customize their operation methods according to their own habits, balancing the flexibility of interaction with personalized needs.

[0049] The core principle of multimodal dynamic feedback is multi-sensory collaborative perception and real-time scene adaptation. Through the dynamic adjustment and synergistic effect of visual, tactile, and audio feedback, users can intuitively and clearly perceive the interaction results, while reducing perceptual load and operational fatigue.

[0050] Example 10: A mixed reality intelligent interaction device integrating eye tracking and gesture recognition, used to implement a mixed reality intelligent interaction method integrating eye tracking and gesture recognition. The mixed reality intelligent interaction device includes a multimodal perception module, an intelligent processing module, and a display feedback module. The intelligent processing module is communicatively connected to the multimodal perception module and controls the execution actions of the display feedback module. The multimodal perception module includes an eye-tracking unit, a gesture environment perception unit, and a user state perception unit; the eye-tracking unit, the gesture environment perception unit, and the user state perception unit are each communicatively connected to the intelligent processing module. The intelligent processing module includes a heterogeneous architecture main processor and an interactive processing chip; the main processor supports multimodal data parallel processing, and the dedicated interactive processing chip integrates three-dimensional context fusion, adaptive threshold calculation, hierarchical gesture recognition, and gaze gesture bidirectional verification algorithms. The display and feedback module includes a gesture environment sensing unit and a user state sensing unit.

[0051] The following specific embodiments illustrate the implementation principle of the present invention: Taking a mixed reality 3D model design and document annotation scenario as an example, when a user wears a mixed reality intelligent interaction device, the device's multimodal perception module, intelligent processing module, and display feedback module start up and complete self-tests. The eye-tracking unit, gesture environment perception unit, and user state perception unit in the multimodal perception module begin to warm up, the intelligent processing module establishes communication connections with each perception unit, and the display feedback module lights up the optical waveguide display, entering the initialization interface.

[0052] Personalized calibration and multimodal data acquisition: 1. Personalized calibration: The system guides users through a personalized calibration process. In the gaze accuracy calibration phase, the user follows a moving virtual marker on the display, rotating their gaze. The eye-tracking unit records the correspondence between the user's pupil position and the gaze coordinates. In the gesture habit sampling phase, the user performs basic gestures such as slightly clenching their fist, sliding their thumb and index finger, and flipping their palm, as prompted. The gesture environment perception unit collects the skeletal key point motion trajectories and electromyographic signals of these gestures. In the visual sensitivity setting phase, the user selects a comfortable visual feedback brightness level through basic gestures. Based on this data, the system generates a personalized configuration file, including dynamic threshold baseline values, gesture recognition model parameters, and feedback preferences, which is automatically synchronized to the user's cloud account and supports cross-device access.

[0053] 2. Synchronous acquisition of multimodal data: After calibration, the user enters the 3D model design and document annotation scenario. The eye-tracking unit collects real-time data on the user's pupil position, gaze point 3D coordinates, gaze movement speed, and blink frequency. The gesture environment perception unit, through an RGB camera, ToF depth camera, and IMU unit, collects coordinates of 26 key skeletal points on the hand, hand posture, movement trajectory, movement speed, and electromyography (EMG) signals. Simultaneously, it constructs a 3D environment map of the current studio using SLAM, detecting ambient light intensity, spatial dimensions, and obstacle distribution. The user status perception unit continuously collects the user's heart rate, blink frequency, and EMG signal data, combining this with the user's past operation records in this scenario to form user status data. All data is synchronized using a timestamp alignment algorithm to ensure consistency in the data's temporal dimension.

[0054] 3D Context Construction and Adaptive Gaze Threshold Calculation: 1. 3D context fusion: The intelligent processing module analyzes the synchronized multimodal data, constructing and fusing three types of context. Regarding object context, the system uses eye-tracking data to pinpoint the user's current gaze target. If the user is gazing at a mechanical 3D model, the model type is identified as a 3D model, with available operations including rotation, scaling, and annotation; the current state is unedited; and the effective area parameter is the model's three-dimensional boundary range. If the user is gazing at a text paragraph in a virtual document, the type is identified as text, with available operations including highlighting, annotation, and page turning; the current state is unselected; and the effective area parameter is the text paragraph's two-dimensional boundary range. Regarding environmental context, the system determines the current ambient light intensity level to be normal, the available operational space to be within three meters of the user (no obstacles), and the obstacle distribution to be tables and chairs in the corner of the studio that do not affect the interaction area. Regarding user state context, the system determines the user's fatigue level as alert based on heart rate and blink frequency, the user's operational proficiency as advanced based on operation history data, and the user's hand force state as stable based on electromyography signals.

[0055] 2. Adaptive gaze threshold calculation: The intelligent processing module calls upon the dynamic threshold baseline value in the dedicated configuration file and calculates the adaptive gaze threshold based on the current 3D context. Since the ambient light intensity is normal, the user is alert, and their operational proficiency is at an advanced level, the scene coefficient is adjusted to 1.0, and the adaptive gaze threshold is the product of the baseline value and 1.0. Simultaneously, the system determines the preset interaction area as a 3D spatial range of 30 to 80 centimeters in front of the user's chest. Taking into account the user's height and the current studio space dimensions, the boundary of the interaction area is automatically fine-tuned to fit the user's natural arm movement range.

[0056] Gaze-Gesture Two-Way Authentication and Interactive Context Activation: 1. Determination of the validity of the gaze: The user focuses their gaze on the mechanical 3D model, and the eye-tracking unit continuously monitors the duration of gaze. When the gaze duration reaches the adaptive gaze threshold, and the gesture environment perception unit detects that the user's hand has entered the preset interaction area, the system pre-activates the interaction context of the 3D model, and a soft halo appears at the edge of the model on the display to indicate that the interaction context has been pre-activated.

[0057] 2. Gaze-Gesture Two-Way Verification: Step 31a: After pre-activation, the system initiates gesture pre-action monitoring. The gesture environment perception unit focuses on capturing pre-action features such as slight finger movements and wrist rotations. When the user's wrist rotates slightly and the fingers show a bending tendency, the pre-action is successfully captured.

[0058] Step 31b: The gesture environment perception unit collects the electromyographic signal intensity features of the finger bending angle and the starting direction of the movement trajectory of the pre-action. The intelligent processing module compares these features with the gesture feature template library corresponding to the available operations of the 3D model and calculates the matching degree through a multi-dimensional weighted fusion algorithm.

[0059] In step 31c, the matching degree calculation result is 85, which meets the standard. The intelligent processing module corrects the gaze focus coordinates through the pre-action trajectory, and corrects the gaze point that originally fell on the edge of the model to the effective area in the center of the model.

[0060] In step 31d, once the matching degree remains stable above 85 for more than 50 milliseconds, the system officially confirms the user's interaction intent, activates the complete interaction context, and loads the extended gesture semantic library corresponding to the 3D model operation.

[0061] Layered gesture execution and multimodal dynamic feedback: 1. Gesture recognition and operation matching: Once the complete interactive context is activated, the system loads the basic extended custom three-level gesture semantic library. When the user needs to rotate the 3D model, they make an extended gesture of drawing a circle with two fingers. The gesture environment perception unit captures the gesture features, and the intelligent processing module quickly identifies and matches the rotation operation command, sending a rotation control signal to the 3D model rendering module.

[0062] Next, the user needs to adjust the model scaling ratio by making a basic thumb-index finger swipe gesture. After the system recognizes the gesture, it matches the parameter adjustment instructions, and the model enlarges according to the swipe direction. The user also needs to add personalized annotations by making a pre-recorded custom gesture of tapping with their index and middle fingers together. The system calls the custom gesture library, matches the annotation instructions, and starts the annotation function.

[0063] 2. Multimodal dynamic feedback: In terms of visual feedback, the model rotates and zooms in real time following the gestures, and virtual annotation boxes appear at the marked positions. Since the ambient light intensity is normal, the feedback brightness remains at a moderate level. Adjusting the scaling ratio is a complex operation, and a scaling ratio progress bar is displayed on the monitor to intuitively show the adjustment status.

[0064] Regarding haptic feedback, when the model rotates to start, the miniature vibration motors on both sides of the device's head emit short vibrations as confirmation feedback; during scaling adjustments, the vibration motors emit continuous vibrations as progress feedback; and when the user's gesture approaches an invalid range, a long vibration is emitted as a warning feedback. Since the user's current fatigue level is conscious, the vibration intensity remains at a standard level.

[0065] In terms of audio feedback, the user selected a crisp sound effect. When the gesture triggers the rotation, zoom, and label operation, the directional speaker emits the corresponding crisp sound effect. Subsequently, if the user performs multiple operations in succession and the operation frequency exceeds the set threshold, the audio volume will be automatically reduced to avoid interference.

[0066] Scene switching and interaction optimization: After the user completes the 3D model operation, they turn their gaze to a text paragraph in the virtual document. The eye-tracking unit detects the change in gaze target, and the intelligent processing module reconstructs the object context, identifies the text type and available operations, and calculates a new adaptive gaze threshold. When the gaze duration reaches the specified value and the hand enters the interaction area, the system pre-activates the text interaction context and loads the corresponding extended gesture semantic library.

[0067] The user makes an extended gesture of drawing a line with their index finger. After the system recognizes this gesture, it matches the highlighted operation command, and the corresponding content in the text paragraph is highlighted. During this process, the user's blinking frequency increases slightly. After the user's state perception unit detects this, the system determines that the fatigue level has changed to mild fatigue. The intensity of the tactile feedback vibration is automatically reduced, and the brightness of the visual feedback is slightly increased.

[0068] During the interaction, the intelligent processing module monitored accidental touches in real time, and no invalid operations with a matching accuracy below 80% were found. All interaction data was recorded in the user's operation history, and the intelligent processing module periodically updated the gesture recognition model parameters in the dedicated configuration file to optimize the recognition accuracy of subsequent interactions.

[0069] Coordination of all modules in the device: Throughout the interaction, the eye-tracking unit of the multimodal perception module continuously outputs gaze point data, the gesture environment perception unit simultaneously collects gesture data and environmental data, and the user state perception unit captures user physiological state data. The heterogeneous architecture main processor of the intelligent processing module is responsible for parallel processing of multimodal data, and the interaction processing chip runs a three-dimensional context fusion adaptive threshold calculation layered gesture recognition and gaze gesture bidirectional verification algorithm to quickly output operation commands. The display feedback module presents visual feedback according to the operation commands and outputs tactile and audio feedback through vibration motors and directional speakers. The three work together to achieve efficient and natural mixed reality interaction.

[0070] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A mixed reality intelligent interaction method integrating eye tracking and gesture recognition, characterized by: Includes the following steps: Step 1: Collect basic user interaction data to complete personalized calibration and generate a unique configuration file; Simultaneously collect eye-tracking data, gesture data, environmental data, and user status data, and achieve time-series synchronization of multimodal data through a timestamp alignment algorithm; Step 2: Construct and fuse object context, environment context, and user state context based on multimodal data; Step 3: Call the dedicated configuration file and calculate the adaptive gaze threshold based on the current scene; When the duration of a user's gaze at the target object is greater than or equal to the adaptive gaze threshold, and the hand enters the preset interaction area, the interaction context is pre-activated. After confirming the interaction intent through two-way verification of gaze gestures, the complete interaction context is activated and the corresponding gesture semantic library is loaded. Confirming interaction intent through two-way verification of gaze gestures specifically includes: Step 31: After pre-activating the interactive context, start the gesture pre-action monitoring to capture pre-action features, including finger micro-movements and wrist rotations. Step 32: Collect the finger bending angle, starting direction of the movement trajectory, and electromyographic signal intensity features of the pre-action, compare them with the gesture feature template library corresponding to the available operations of the current object, and calculate the matching degree through a multi-dimensional weighted fusion algorithm; Step 33: If the matching degree is greater than or equal to 80%, the gaze focus coordinates are corrected by the pre-action trajectory so that the focus falls accurately within the effective area of ​​the object; if the matching degree is greater than or equal to 50% but less than 80%, the user is prompted to adjust the gesture through weak visual feedback; if the matching degree is less than 50%, the pre-activated state is maintained and monitoring continues. Step 34: When the matching degree remains greater than or equal to 80% for more than 50ms, the interaction intent is formally confirmed and the complete interaction context is activated. Step 4: Based on the activated interactive context, load the corresponding basic, extended, and custom three-level gesture semantic library, recognize user gestures and match operation commands, and provide dynamic feedback on interaction results through visual, tactile, and audio multimodal feedback. Step 2, 3D context fusion, specifically includes: Object context: Identifies the type of gaze target, available operations, current state, and valid region parameters; Environmental context: Analyze ambient light intensity levels, available operational space, and obstacle distribution; User status context: Determine the user's fatigue level based on heart rate and blink frequency, determine the user's operation proficiency based on operation history data, and determine the hand force state based on electromyography signals; The formula for the multi-dimensional weighted fusion algorithm is: Where M is the total matching degree, Weighted matching scores are calculated based on the basic features of the gesture. This is the fixation correction factor. For scene adaptation coefficients; It is the weight coefficient of the i-th gesture feature basic feature. is the single-item matching score for the i-th gesture basic feature.

2. The mixed reality intelligent interaction method integrating eye tracking and gesture recognition according to claim 1, characterized in that: The personalized calibration mentioned in step 1 includes gaze point accuracy calibration, gesture habit sampling, and visual sensitivity settings; the dedicated configuration file contains dynamic threshold benchmark values, gesture recognition model parameters, feedback preferences, and user operation history data, and supports cross-device synchronization.

3. The mixed reality intelligent interaction method integrating eye tracking and gesture recognition according to claim 2, characterized in that: The eye-tracking data includes pupil position, 3D coordinates of fixation point, gaze movement speed, and blink frequency; The gesture data includes the coordinates of 26 key skeletal points of the hand, hand posture, movement trajectory, movement speed, and gesture electromyography signals. The environmental data includes a 3D environment map constructed by SLAM, ambient light intensity, spatial dimensions, and obstacle distribution; The user status data includes heart rate data, blink frequency data, gesture electromyography signal data, and operation history data.

4. The mixed reality intelligent interaction method integrating eye tracking and gesture recognition according to claim 3, characterized in that: The adaptive gaze threshold mentioned in step 3 is calculated using the following formula: ; Where DT is the dynamic threshold, D is the baseline value, and a is the scene coefficient. The scene coefficient is dynamically adjusted according to the ambient light intensity level, user fatigue level, and operation proficiency, with an adjustment range of 0.8-1.

5. The preset interaction area is a three-dimensional space range of 30-80cm in front of the user's chest; the interaction area can be adaptively adjusted according to the user's height, range of limb movement and current space size.

5. The mixed reality intelligent interaction method integrating eye tracking and gesture recognition according to claim 4, characterized in that: The three-level gesture semantic library mentioned in step 4 includes a basic gesture library, an extended gesture library, and a custom gesture library; The basic gesture library includes globally universal gestures, which include a slight fist clench, thumb and index finger swipe, and palm flip. The gesture library is extended to bind specific object types, including contextual gestures, such as drawing circles with two fingers when gazing at a 3D model and drawing lines with the index finger when gazing at text. The custom gesture library allows users to add exclusive gestures through recording, naming, and binding processes, and the recognition model is automatically optimized after binding.

6. The mixed reality intelligent interaction method integrating eye tracking and gesture recognition according to claim 5, characterized in that: Step 3, the multimodal dynamic feedback specifically includes visual feedback, tactile feedback, and audio feedback; Visual feedback adaptively adjusts brightness based on ambient light intensity, and displays progress bars or numerical prompts during complex operations; The tactile feedback distinguishes between confirmation feedback, progress feedback, and warning feedback, and the vibration intensity adaptively adjusts according to the user's fatigue level. The audio feedback allows users to select sound effect types, and automatically reduces the volume of the set sound effect when the operation frequency exceeds a set threshold.

7. A mixed reality intelligent interactive device integrating eye tracking and gesture recognition, characterized in that: To implement the mixed reality intelligent interaction method that integrates eye tracking and gesture recognition as described in claim 6, the mixed reality intelligent interaction device includes a multimodal perception module, an intelligent processing module, and a display feedback module. The intelligent processing module is communicatively connected to the multimodal perception module, and the intelligent processing module controls the execution actions of the display feedback module. The multimodal perception module includes an eye-tracking unit, a gesture environment perception unit, and a user state perception unit; the eye-tracking unit, the gesture environment perception unit, and the user state perception unit are each communicatively connected to the intelligent processing module. The intelligent processing module includes a heterogeneous architecture main processor and an interactive processing chip; The main processor supports multimodal data parallel processing, and a dedicated interactive processing chip integrates three-dimensional context fusion, adaptive threshold calculation, hierarchical gesture recognition, and gaze gesture bidirectional verification algorithms. The display feedback module includes a gesture environment perception unit and a user state perception unit.