An interaction design method and system based on gesture recognition
By suppressing global mode switching commands in the locked state and introducing a dual confirmation mechanism, combined with visual and haptic feedback to optimize prompts, the problem of mode switching caused by gesture misrecognition in complex 3D scenes is solved, improving the accuracy and smoothness of interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Filing Date
- 2026-05-07
- Publication Date
- 2026-07-10
Smart Images

Figure CN122363524A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of interactive technology, and more specifically, to an interactive design method and system based on gesture recognition. Background Technology
[0002] In gesture interaction systems designed for complex 3D scenes, complex gestures requiring high precision or multiple steps are often introduced to address the ambiguity of target selection in high-density object environments. However, when users are focused on performing these complex and delicate actions, a certain intermediate or transitional hand gesture formed during the movement can easily be misinterpreted by the system as an independent, global command, such as a command to switch the overall software operating mode. This misinterpretation stems from the system's isolated and instantaneous judgment of gestures, lacking understanding and analysis of the current task context (such as performing a highly difficult selection task). This leads to the system switching modes unconsciously, causing subsequent user interactions to be misinterpreted, resulting in unintended modifications to the 3D scene and severely disrupting the continuity of the workflow and the reliability of the interaction.
[0003] For example, in the 3D animation production process, designers use gesture-based interaction systems for fine-grained operations, such as adjusting tiny connectors in a complex environment with a lot of detail. When a designer attempts to grasp a specific target, because the target object is surrounded by other objects, the projection of their hand movements in 3D space may simultaneously cover multiple objects. When the system interprets the grasping intent, it struggles to accurately determine the designer's true target, often mistakenly selecting objects that are larger or closer. This frequent selection error forces designers to repeatedly undo and retry, significantly slowing down the workflow and fragmenting what should have been a smooth interactive experience.
[0004] A deeper problem lies in the fact that any fully functional 3D editing software includes multiple working modes, such as object transformation mode, viewpoint control mode, material drawing mode, and lighting editing mode. In this gesture interaction system, switching between different modes is also accomplished through specific gesture commands. When the above problems overlap, a critical system flaw is exposed. Imagine a designer intently focused on a complex scene, carefully performing a high-precision selection action, attempting to select a tiny ornament. As he extends his right hand, diligently adjusting his fingers to complete a precise pinching motion, his finger posture, for a brief moment, might unintentionally constitute the system's preset gesture of "switching to material drawing mode." Lacking an understanding of the designer's current task context, the system cannot distinguish whether this gesture is merely an intermediate step in a complex action or an independent instruction. Therefore, the system immediately responds to this unintentional instruction, switching the entire software's operating mode from object transformation to material drawing. The designer, unaware of this, continues with subsequent actions, believing he has successfully selected the target and is preparing to move it to a new location. However, his subsequent movement gesture was interpreted by the system as a brushstroke in the new mode. As a result, an incorrect texture was applied directly to the entire model, causing severe damage to the scene. The designer's workflow was completely disrupted. He had to first undo this disastrous mistake, then execute the correct gesture to switch back to the object transformation mode, and finally restart the tiring process of precise selection. The entire process was filled with frustration and doubts about the system's reliability. Summary of the Invention
[0005] This application discloses an interaction design method and system based on gesture recognition to solve at least one of the technical problems in the prior art.
[0006] Firstly, this application discloses an interaction design method based on gesture recognition, specifically including:
[0007] Obtain user hand movement and posture information;
[0008] Based on hand movement and posture information, determine whether the user has begun to perform a high-precision selection task;
[0009] In response to the determination that the user has started to perform a high-precision selection task, the system's working state is switched to the locked state;
[0010] In the locked state, when the user's hand is detected to form a gesture that constitutes a global mode switching command, the global mode switching command is not executed immediately;
[0011] In the locked state, when the user's dominant hand is detected to form a gesture indicating a global mode switching command, and at the same time the user's non-dominant hand is detected to perform a preset auxiliary confirmation action, the global mode switching command is executed.
[0012] In the locked state, priority is given to responding to local operation commands directly related to high-precision selection tasks; and
[0013] The lock is released when the high-precision selection task is completed or explicitly cancelled.
[0014] Secondly, this application also discloses an interaction design system based on gesture recognition, the system comprising:
[0015] The information acquisition module is used to acquire the user's hand movements and posture information;
[0016] The task judgment module is used to determine whether the user has started to perform a high-precision selection task based on hand movement and posture information;
[0017] The state switching module is used to switch the system's working state to the locked state in response to the user's determination that the user has started to execute a high-precision selection task.
[0018] The instruction suppression module is used to prevent the global mode switching instruction from being executed immediately when the user's hand is detected to form a gesture that constitutes a global mode switching instruction in the locked state.
[0019] The dual confirmation module is used to execute the global mode switching command when the user's dominant hand forms a gesture indicating a global mode switching command and the user's non-dominant hand performs a preset auxiliary confirmation action in the locked state.
[0020] The local priority module is used to prioritize responding to local operation commands directly related to the high-precision selection task while the system is locked; and
[0021] The status release module is used to release the locked state when the high-precision selection task is completed or explicitly canceled.
[0022] Compared with the prior art, this application has at least the following beneficial effects:
[0023] This application provides an interaction design method based on gesture recognition. By introducing a locked state and a dual confirmation mechanism, it effectively solves the problem in the prior art that when users perform high-precision operations in complex 3D scenes, gesture misrecognition is likely to occur, leading to unexpected switching of the global mode and interruption of the workflow.
[0024] Specifically, when the system determines that the user has begun performing a high-precision selection task, it switches the system's operating state to a locked state. In this locked state, even if the system detects that the user's hand has formed a gesture indicating a global mode switching command, it will not immediately execute the command, thus avoiding unintentional mode switching. The system will only execute the global mode switching command when it simultaneously detects that the user's dominant hand has formed a gesture indicating a global mode switching command, and the user's non-dominant hand has performed a preset auxiliary confirmation action. This greatly improves the accuracy of global mode switching and the clarity of the user's intent.
[0025] Furthermore, in the locked state, the system prioritizes responding to local operation commands directly related to the high-precision selection task, ensuring the user's smoothness and focus during precise operations and avoiding efficiency degradation caused by interference from global commands. When the high-precision selection task is completed or explicitly canceled, the system unlocks and returns to normal operating mode. Attached Figure Description
[0026] Figure 1 This is a flowchart illustrating an interaction design method based on gesture recognition provided in this application.
[0027] Figure 2 This is a schematic diagram of the structure of an interaction design system based on gesture recognition provided in this application. Detailed Implementation
[0028] The technical solutions in this application will now be clearly and completely described in conjunction with the accompanying drawings.
[0029] This application proposes an interaction design method based on gesture recognition, such as... Figure 1 As shown, it includes the following steps:
[0030] Obtain user hand movement and posture information;
[0031] Based on hand movement and posture information, determine whether the user has begun to perform a high-precision selection task;
[0032] In response to the determination that the user has started to perform a high-precision selection task, the system's working state is switched to the locked state;
[0033] In the locked state, when the user's hand is detected to form a gesture that constitutes a global mode switching command, the global mode switching command is not executed immediately;
[0034] In the locked state, when the user's dominant hand is detected to form a gesture indicating a global mode switching command, and at the same time the user's non-dominant hand is detected to perform a preset auxiliary confirmation action, the global mode switching command is executed.
[0035] In the locked state, priority is given to responding to local operation commands directly related to high-precision selection tasks; and
[0036] The lock is released when the high-precision selection task is completed or explicitly cancelled.
[0037] This application effectively solves the problem of accidental mode switching caused by gesture misrecognition in high-precision selection tasks by introducing a locked state, a dual confirmation mechanism, and a local priority response strategy. It significantly improves the accuracy, smoothness, and user experience of the interaction, and provides a more reliable and efficient solution for gesture interaction in complex 3D scenes.
[0038] To facilitate a clearer understanding of the technical solution of this application, the following will explain some key terms and implementation environments involved. The "gesture recognition" mentioned in this application refers to capturing information such as the user's hand movement trajectory, finger joint angles, and palm orientation in three-dimensional space using sensors (e.g., depth cameras, infrared sensors, or inertial measurement units, IMUs), and matching this information with a preset gesture model to identify the specific operation command the user intends to execute. For example, a "pinch" gesture might be recognized as a "grab" command, while an "open palm" gesture might be recognized as a "release" command.
[0039] "High-precision selection tasks" refer to tasks in a 3D virtual environment where users need to precisely select one or more specific target objects. These target objects may be small in size, close to other objects, or located in complex and dense scenes, requiring extremely high accuracy. Examples include selecting a tiny screw in 3D modeling software or precisely manipulating a tissue in a virtual surgical simulation.
[0040] "Global mode switching commands" are commands that, once executed, change the operating mode of the entire system or application, thus affecting all subsequent interactions. Examples include switching from "object transformation mode" to "material drawing mode," or from "viewpoint control mode" to "animation editing mode." These commands typically have high priority and a wide range of impact.
[0041] "Locked state" is a system operating state introduced in this application. In this state, the system's response logic to certain specific commands (especially global mode switching commands) will change to prevent accidental operation. Entering the locked state usually means that the user is performing a task that requires a high degree of focus and precise operation.
[0042] The "dominant hand" typically refers to the user's preferred hand for fine motor operations or primary control; for example, a right-handed user usually uses their right hand as their dominant hand. The "non-dominant hand," on the other hand, is typically used for auxiliary operations or to provide additional confirmation information.
[0043] The implementation environment of this application typically includes a three-dimensional virtual reality (VR) or augmented reality (AR) system, which includes gesture recognition sensors, display devices (such as VR headsets or AR glasses), and computing units for processing gesture data and rendering three-dimensional scenes. Users can perform natural gesture interactions in the virtual environment by wearing or using these devices.
[0044] The core of the gesture recognition-based interaction design method in this application lies in ensuring, through a series of carefully designed steps, that users can avoid accidental switching of the global mode due to unintentional gestures during high-precision selection tasks, thereby improving the accuracy and smoothness of the interaction.
[0045] First, the system needs to acquire the user's hand movement and posture information. This can be achieved in several ways. For example, a depth camera mounted on the user's VR headset can be used to capture real-time 3D skeletal data of the user's hands, including the position and rotation information of each finger joint. Another approach is for the user to wear smart gloves with an inertial measurement unit (IMU), whose sensors can provide posture, angular velocity, and acceleration data of the hand and fingers. This data is then transmitted to a processing unit for analysis.
[0046] Next, based on the acquired hand movement and posture information, the system "determines whether the user has begun performing a high-precision selection task." This determination process can be based on various heuristic rules or machine learning models. For example, the system can monitor whether the user's hand remains within a specific area for an extended period and whether the fingers are making fine pinching or grasping movements. Furthermore, it can incorporate information about the user's gaze focus. If the user's gaze is focused on a small or densely packed area of virtual objects for a prolonged period, it may indicate that the user is attempting a high-precision selection. For instance, when a user brings their index finger and thumb together and slowly moves them towards a tiny button in a virtual scene, the system analyzes this hand posture and movement trajectory, combined with whether the user's gaze is focused on the button, to determine whether the user has entered a high-precision selection task.
[0047] In response to the system's determination that a user has begun performing a high-precision selection task, the system switches to a "locked state." Once the system confirms that the user is making a high-precision selection, it enters a special "locked state." In this state, the system's response logic to certain gesture commands changes to prevent accidental operations. For example, the system might display a visual indicator on the user interface to inform the user that it is now in a locked state, or provide a slight vibration cue through a haptic feedback device.
[0048] In the locked state, when the system "recognizes a user's hand forming a gesture indicating a global mode switching command, it does not immediately execute the global mode switching command." This is one of the key innovations of this application. In the locked state, even if the system recognizes a user's hand forming a preset global mode switching command gesture (e.g., a specific gesture for switching to "material drawing mode"), the system will not immediately execute the command. Instead, the command will be temporarily suppressed or suspended. For example, when a user is in the locked state, their dominant hand may unintentionally make an "open palm" gesture, which is typically defined as a global command to switch to "viewpoint control mode." In this case, the system will not immediately switch modes but will wait for further confirmation.
[0049] To ensure the accuracy of user intent, this application introduces a dual confirmation mechanism. Specifically, in the locked state, the global mode switching command is only executed when the system "recognizes that the user's dominant hand forms a posture indicating a global mode switching command, and simultaneously recognizes that the user's non-dominant hand performs a preset auxiliary confirmation action." This means that even if the dominant hand forms a posture indicating a global mode switching command, the system still requires the non-dominant hand to provide an additional, explicit auxiliary confirmation action before the global command can be executed. This auxiliary confirmation action can take various forms, such as the non-dominant hand making a "fist" gesture, lightly touching a virtual button, or performing a specific "click" gesture. For example, when the user's dominant hand forms an "open palm" posture, if the user's non-dominant hand simultaneously makes a "pinch" gesture, the system will confirm that the user indeed wants to switch to "viewpoint control mode" and execute the command. This dual confirmation mechanism greatly reduces the possibility of accidental operation.
[0050] Furthermore, in the locked state, the system prioritizes local operation commands directly related to the high-precision selection task. This means that during the locked state, the system focuses its processing power and response priority on operations directly related to the current high-precision selection task. For example, if a user is trying to select a small object, local gesture commands related to "select," "grab," or "zoom in / out of a local area" will be processed and responded to first, while other unrelated global commands will be suppressed or require double confirmation. For instance, if a user makes a "pinch" motion with their dominant hand while in the locked state, the system will immediately respond and execute the "grab" command without being confused by simultaneously recognizing an intermediate gesture that could be used to switch between global modes.
[0051] Finally, the system unlocks when the high-precision selection task is completed or explicitly cancelled. Once the user successfully completes the high-precision selection task (e.g., successfully selects and manipulates the target object), or the user explicitly cancels the current high-precision selection task through a gesture or voice command, the system exits the locked state and returns to normal system operation mode. For example, when the user successfully moves the selected tiny object to the target location, the system automatically unlocks, allowing the user to freely execute other global or local commands.
[0052] For example, in a 3D animation production scenario, a designer needs to adjust a tiny connector within a complex environment containing numerous details. Traditional systems might inadvertently trigger the global command to "switch to material drawing mode" due to an unintentional intermediate posture formed by the designer during adjustment, leading to workflow interruptions and scene disruptions. However, using the method described in this application, the system enters a locked state when the designer begins such a high-precision selection task. At this point, even if the designer's dominant hand unintentionally forms a "switch to material drawing mode" posture during adjustment, the system will not immediately switch modes. The system will only switch modes when the designer's non-dominant hand simultaneously performs a clear auxiliary confirmation action (e.g., the non-dominant hand makes a "fist" gesture). During this period, the designer can focus on precisely selecting and manipulating the connector, and the system will prioritize responding to local operation commands such as "grab" and "move." This mechanism ensures the accuracy and smoothness of the designer's operations in high-precision tasks, avoiding frustration and work interruptions caused by misoperation.
[0053] Through the above improvements, this application can significantly enhance the reliability and efficiency of gesture interaction in complex 3D scenes, reduce misoperations, ensure the continuity of workflow, and thus provide users with a more intuitive, accurate and efficient interactive experience.
[0054] However, in some complex or high-pressure interactive environments, even with a dual gesture confirmation mechanism, users may still accidentally trigger a global mode switching command due to distraction, environmental interference, or slight deviations in hand movements, thereby interrupting ongoing high-precision selection tasks and affecting user experience and operational efficiency.
[0055] In this regard, this application further proposes that, in the locked state, when the user's dominant hand is detected to be forming a posture indicating a global mode switching command, and simultaneously the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the steps for executing the global mode switching command include:
[0056] In the locked state, it senses the user's hand movements and posture information and determines whether the user has started to perform a high-precision selection task;
[0057] In response to the determination that the user has started to perform a high-precision selection task, the system's working state is switched to the locked state;
[0058] In the locked state, sense whether the dominant hand has formed a posture that triggers a global mode switching command;
[0059] In the locked state, when the dominant hand is perceived to be in a posture that forms a global mode switching command, a visual cue is generated and displayed in a non-critical position in the 3D scene display area.
[0060] When locked, monitor whether the user's gaze moves and stays on the visual cue area;
[0061] In the locked state, when the dominant hand is detected to form a posture that signals a global mode switching command, and the user's non-dominant hand is detected to perform a preset auxiliary confirmation action, and the user's gaze is detected to move and stay on the visual cue area, the global mode switching command is executed.
[0062] Specifically, in the locked state, the system continuously senses the user's hand movements and posture information, and determines whether the user is performing a high-precision selection task. This step aims to ensure that the system is always in a state of accurate understanding of the user's intentions, providing context for subsequent command processing. In response to determining that the user has begun performing a high-precision selection task, the system's operating state is switched to the locked state to ensure that the global mode switching command is not easily triggered accidentally during critical operations.
[0063] In this locked state, the system continuously monitors whether the user's dominant hand is gesturing in a manner in which a global mode-switching command is initiated. When the system detects this gesture, to avoid accidental execution, it does not immediately execute the command but instead generates a visual cue. This cue is designed to appear in a non-critical location within the 3D scene display area, such as a corner or edge of the screen, to avoid interfering with the user's visual focus on the high-precision selection task they are currently performing. This visual cue aims to provide the user with clear, non-intrusive feedback, informing them that the system has detected a potential mode-switching intent.
[0064] Furthermore, the system monitors whether the user's gaze shifts and lingers on the visual cue area. This gaze focus monitoring, achieved through eye-tracking technology, determines whether the user has actively paid attention to and acknowledged the visual cue. The global mode switching command is only executed when the system detects the dominant hand forming a posture indicating a global mode switching instruction, simultaneously detects the user's non-dominant hand performing a preset auxiliary confirmation action, and detects the user's gaze shifting and lingering on the visual cue area. This means that the execution of the global mode switching command requires three confirmation conditions: dominant hand posture, non-dominant hand auxiliary confirmation action, and user gaze focus confirmation.
[0065] This application effectively solves the potential for accidental touches when relying solely on gesture-based dual confirmation in a locked state by introducing visual cues and gaze focus monitoring as additional confirmation mechanisms for global mode switching commands. When the user's dominant hand forms the gesture for the global mode switching command, the system first signals the user of the potential mode switch by displaying a visual cue in a non-critical location, in a way that does not interfere with the current high-precision task. Subsequently, by monitoring whether the user's gaze focus actively moves and stays on the visual cue area, the system can determine whether the user has truly paid attention to and understood the cue, thereby ensuring that the user has a clear intention to switch modes. This multimodal confirmation method, combining physical confirmation through gestures and cognitive confirmation through gaze focus, significantly improves the accuracy and reliability of global mode switching command execution.
[0066] Through the above technical solution, the accidental touch rate of global mode switching commands is significantly reduced during high-precision selection tasks. Even if a user unintentionally makes a gesture similar to a global mode switch while performing fine-grained operations, the system will not respond immediately. Instead, it will guide the user to confirm the gesture through visual cues. The command will only be executed when the user actively shifts their gaze to the cue area and performs the confirmation action, thus avoiding task interruptions and a degraded user experience due to accidental operations. This enhanced confirmation mechanism allows the system to maintain the intuitiveness of gesture interaction while significantly improving the robustness of the interaction and the user's sense of control over the system's behavior.
[0067] Suppose a user is performing a detailed 3D model editing task in a virtual reality (VR) environment, a high-precision selection task. During this process, the user may frequently adjust model details, with their attention highly focused on the model itself. If, at this moment, the user unintentionally makes a dominant hand gesture similar to a global mode-switching command (e.g., "switch tool mode"), and their non-dominant hand happens to perform a pre-defined auxiliary confirmation action (e.g., a slight pinching gesture), the system might immediately switch modes, interrupting the user's editing process. However, according to the solution of this application, when the system detects the dominant hand gesture and the non-dominant hand auxiliary confirmation action, it does not immediately switch modes. Instead, the system displays a semi-transparent visual cue in a non-critical area of the user's field of vision (e.g., the lower left corner), such as "Switch to tool mode?". The system continuously monitors the user's gaze focus. Only when the user actively moves their gaze away from the 3D model and focuses on this visual cue area, confirming their intention, will the system execute the global mode-switching command, switching the tool mode to the new state. If the user does not move their gaze to the prompt area, the system will not switch modes even if the gesture has been completed, thus avoiding accidental mode switching and ensuring the continuity of high-precision tasks and the smoothness of the user experience.
[0068] However, in practical applications, if the display method of visual cues fails to fully consider the dynamic characteristics of the current 3D scene and the user's cognitive load, it may lead to poor cues. For example, the cues may be too abrupt and interfere with the user's ongoing high-precision selection task, or the cues may be too subtle and fail to attract the user's attention in time.
[0069] In response, this application further proposes that, in a locked state, when the dominant hand is perceived to be forming a global mode switching command, a visual cue is generated and displayed at a non-critical location in the 3D scene display area. The steps include:
[0070] In the locked state, when the dominant hand is perceived to be forming a global mode switching command, the visual complexity information of the current 3D scene is obtained.
[0071] Obtain the position information of the user's gaze focus within the 3D scene display area;
[0072] Obtain the distribution density information of interactive elements in a 3D scene;
[0073] Based on the visual complexity information of the 3D scene, the user's gaze focus position information, and the distribution density information of interactive elements, determine the display parameters of the visual cues;
[0074] Based on the determined display parameters, visual cues are generated and displayed at non-critical locations in the 3D scene display area.
[0075] Specifically, acquiring visual complexity information of the current 3D scene refers to the system analyzing the image data of the current 3D scene, such as calculating the image's texture details, color diversity, and edge count, to quantify the visual busyness of the scene. Its purpose is to assess the degree to which the current scene occupies the user's visual attention. Acquiring the location information of the user's gaze focus within the 3D scene display area can be understood as continuously monitoring the user's eye movements through eye-tracking devices, thereby accurately identifying the specific location and dwell time of the user's current gaze within the 3D scene display area. Its purpose is to understand the user's current visual focus and avoid conflicts between visual cues and the key areas the user is focusing on. In practical applications, acquiring the distribution density information of interactive elements in the 3D scene specifically involves the system identifying and counting the number of all interactive objects (such as buttons, menus, selectable targets, etc.) in the current 3D scene and their spatial distribution, calculating their density within specific areas. Its purpose is to assess potential interactive hotspots in the scene and avoid visual cues obscuring or interfering with important interactive elements.
[0076] Furthermore, determining the display parameters of visual cues based on the visual complexity of the 3D scene, the user's focal point, and the distribution density of interactive elements involves the system comprehensively considering these three environmental and user state information, dynamically adjusting various attributes of the visual cues, such as their size, transparency, color, flashing frequency, animation effects, and specific display position. For example, when the scene's visual complexity is high, the user's focal point is concentrated in a key area, and the density of interactive elements is high, the transparency of the visual cues can be increased, their size can be reduced, and they can be placed in an edge area far from the user's focal point. Conversely, when the scene's visual complexity is low and the user's focal point is dispersed, the visual cues can be designed to be more eye-catching, such as increasing their size or decreasing their transparency. The aim is to ensure that the visual cues effectively attract the user's attention without unnecessarily interfering with the user's ongoing high-precision selection task. Thus, based on the determined display parameters, visual cues are generated and displayed in non-critical positions within the 3D scene display area, ensuring that the presentation of the visual cues dynamically adapts to the current environment and user state.
[0077] This application's solution acquires and analyzes information on the visual complexity of the current 3D scene, the user's gaze focus position, and the distribution density of interactive elements, enabling the system to comprehensively perceive the current user environment and cognitive state. It is precisely this comprehensive consideration of information that allows the system to intelligently determine the optimal display parameters for visual cues. For example, when the system detects high visual complexity in the 3D scene, it tends to use simpler, more transparent visual cues to avoid increasing the user's cognitive burden; when the user's gaze is focused on a key area of a high-precision selection task, the visual cues are strategically placed in non-critical areas, and their size or transparency may be adjusted to avoid interfering with the user's core task; simultaneously, by considering the distribution density of interactive elements, the system can prevent visual cues from obscuring important interactive objects. This dynamic and adaptive cue mechanism effectively solves the problems of interference or inconspicuousness that may arise from traditional fixed cues, thus ensuring that auxiliary confirmation cues for global mode switching commands can be effectively perceived by the user without affecting the execution of high-precision tasks.
[0078] Through the above technical solution, this application can dynamically adjust the display parameters of visual cues based on the visual complexity of the current 3D scene, the user's focal point, and the distribution density of interactive elements. This makes the presentation of visual cues more intelligent and user-friendly, significantly improving the user experience. Specifically, this solution avoids interference caused by improper display of visual cues when users perform high-precision selection tasks, while ensuring the effectiveness and perceptibility of the cues. Therefore, when auxiliary confirmation of global mode switching commands is required, users can complete the confirmation operation with lower cognitive load and higher efficiency, thereby improving the smoothness and accuracy of the entire interaction process.
[0079] Suppose a user is performing a high-precision selection task in a complex virtual architectural design scene, such as precisely selecting a tiny architectural component. The system, through image analysis, determines the visual complexity of the current 3D scene to be "high." Eye-tracking technology shows the user's gaze is focused on the tiny architectural component, and scene analysis indicates the density of interactive elements in that area is also "high." When the user's dominant hand forms a gesture indicating a global mode switch command, the system determines the display parameters of the visual cue based on this information. For example, the system might set the visual cue's transparency to 80%, reduce its size to 50% of the default size, and place it in the upper right corner of the 3D scene display area (a non-critical area far from the user's focal point), while using a gentle fade-in animation. This approach ensures that the visual cue neither obscures the architectural component the user is interacting with nor distracts the user from the high-precision task, while still subtly reminding the user of the pending global mode switch command.
[0080] However, relying solely on visual cues in implementation may have limitations in some complex or noisy environments. For example, when users are under high visual load, have distracted attention, or experience poor ambient lighting, a single visual cue may not be sufficient to ensure that users can perceive the cue information in a timely and accurate manner, which may affect the efficiency and reliability of the interaction, or even lead to misoperation.
[0081] To address this, this application further proposes an improved interaction design method. In the aforementioned locked state, when the dominant hand is perceived to be forming a global mode switching command, a visual cue is generated, and the cue is displayed at a non-critical location in the 3D scene display area. The steps include:
[0082] In the locked state, when the dominant hand is perceived to be forming a global mode switching command, the visual complexity information of the current 3D scene is obtained.
[0083] Obtain the position information of the user's gaze focus within the 3D scene display area;
[0084] Obtain the distribution density information of interactive elements in a 3D scene;
[0085] Obtain the noise level of the user's current environment;
[0086] To obtain the vibration feedback capability of the haptic feedback device worn by the user;
[0087] Based on the visual complexity information of the 3D scene, the user's gaze focus position information, the distribution density information of interactive elements, the noise level, and the vibration feedback capability, the display parameters and non-visual feedback parameters of the visual cues are determined.
[0088] Based on the determined display parameters, visual cues are generated and displayed at non-critical locations in the 3D scene display area;
[0089] Based on the determined non-visual feedback parameters, tactile feedback is emitted through a tactile feedback device.
[0090] Specifically, in the aforementioned locked state, when the dominant hand is perceived to be forming a global mode switching command, the system not only acquires information on the visual complexity of the current 3D scene, the position of the user's gaze focus within the 3D scene display area, and the distribution density of interactive elements in the 3D scene, but also further acquires the noise level of the user's current environment and the vibration feedback capability of the haptic feedback device worn by the user. The environmental noise level can be understood as the background noise intensity of the user's environment, which is collected and analyzed in real time through a microphone array to assess the degree of interference to the user's auditory channel. The vibration feedback capability of the haptic feedback device refers to parameters such as vibration intensity, frequency, and mode provided by the device worn by the user (e.g., a smartwatch, bracelet, or dedicated haptic gloves), aiming to understand the available non-visual feedback methods. Based on this multimodal information—the visual complexity of the 3D scene, the position of the user's gaze focus, the distribution density of interactive elements, the noise level, and the vibration feedback capability—the system can comprehensively determine the display parameters of visual cues and the parameters of non-visual feedback. The display parameters for visual cues can include brightness, color, size, flashing frequency, and transparency, while non-visual feedback parameters can include the intensity, duration, and vibration mode (e.g., single vibration, pulse vibration, or continuous vibration) of haptic feedback. Based on the determined display parameters, visual cues are generated and displayed in non-critical locations within the 3D scene display area to avoid interfering with the user's ongoing high-precision selection task. Simultaneously, based on the determined non-visual feedback parameters, haptic feedback is emitted through a haptic feedback device worn by the user. For example, when ambient noise is high or visual complexity is high, the intensity or duration of the haptic feedback may be increased to ensure the user can perceive the cues.
[0091] This application addresses the potential for insufficient perception in complex or noisy environments by incorporating the perception of environmental noise levels and the vibration capability of haptic feedback devices into the determination of feedback parameters. Specifically, when the system senses the dominant hand forming a global mode-switching command posture, in addition to considering visual environmental factors (such as scene complexity, gaze focus, and density of interactive elements) to optimize visual cues, it further considers auditory interference (noise level) and available haptic feedback capabilities in the user's environment. It is precisely this comprehensive consideration of multimodal information that enables the system to dynamically adjust its feedback strategy, not only optimizing the display parameters of visual cues but also generating and issuing non-visual feedback (such as haptic feedback) that matches the current environment and user state. This multi-channel feedback mechanism effectively compensates for the limitations of a single visual channel in specific situations, ensuring that users can reliably receive mode-switching command prompts in various complex environments.
[0092] Through the above technical solution, this application can significantly improve the reliability and efficiency of users perceiving global mode switching command prompts in complex or noisy environments. By combining visual cues and tactile feedback, even under conditions of visual information overload, user distraction, or high environmental noise, users can still perceive the prompts through touch, thereby reducing the risk of misoperation and shortening user response time. This multimodal feedback mechanism makes the interaction design more robust and user-friendly, especially suitable for high-precision selection task scenarios with high requirements for operational accuracy and timeliness.
[0093] However, in practical applications, the lighting conditions in the user's environment may be unstable or less than ideal, which can affect the accuracy of gesture recognition, especially the accuracy of recognition of actions assisted by the non-dominant hand. If the above problems are not addressed, the system may misjudge or miss user intentions under poor lighting conditions, thereby reducing the reliability of the interaction and the user experience.
[0094] In this regard, this application further proposes that, in the locked state, when the user's dominant hand is detected to be forming a posture indicating a global mode switching command, and simultaneously the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the steps for executing the global mode switching command include:
[0095] Obtain image data of the user's current environment, and obtain the ambient light stability index based on the image data of the environment;
[0096] In the locked state, when the user's dominant hand is detected to be forming a global mode switching command posture, and at the same time the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the user's non-dominant hand recognition threshold is adjusted according to the ambient light stability index, and the decision is made whether to execute the global mode switching command based on the adjustment result.
[0097] Specifically, acquiring image data of the user's current environment can be achieved through the system's built-in camera or an externally connected visual sensor. Image data can include, but is not limited to, RGB images, depth images, or infrared images. The ambient light stability index can be understood as an indicator measuring the degree and quality of changes in current ambient light conditions. For example, it can be obtained by real-time analysis of parameters such as brightness, contrast, and color saturation of the acquired image data, or by calculating the light differences between image frames. When light changes drastically or is insufficient, the index value may be low, and vice versa. Adjusting the user's non-dominant hand recognition threshold refers to dynamically changing the confidence or matching threshold required to recognize the non-dominant hand's auxiliary confirmation action based on the ambient light stability index value. For example, when the ambient light stability index is low (indicating poor or unstable lighting conditions), the non-dominant hand recognition threshold can be appropriately increased to reduce the probability of false recognition; when the ambient light stability index is high (indicating good and stable lighting conditions), the non-dominant hand recognition threshold can be appropriately decreased to improve recognition sensitivity. The adjustment result determines whether to execute the global mode switching command. This means that the system will only confirm the effectiveness of the non-dominant hand's auxiliary confirmation action and then execute the global mode switching command when the adjusted recognition threshold is met.
[0098] This application effectively addresses the issue of decreased recognition accuracy of non-dominant hand-assisted confirmation actions under varying lighting conditions by introducing and utilizing the ambient lighting stability index. Specifically, when the system is locked and detects the user's dominant hand forming a global mode switching command gesture, it first acquires and analyzes the current environment's image data to assess the stability of the ambient lighting. Since lighting conditions directly affect the performance of gesture recognition algorithms, the system can quantitatively evaluate the reliability of the current recognition environment by acquiring the ambient lighting stability index in real time. Based on this, the system dynamically adjusts the recognition threshold for non-dominant hand-assisted confirmation actions according to the assessed ambient lighting stability index. For example, in poor or unstable lighting conditions, increasing the recognition threshold allows the system to more cautiously judge the non-dominant hand's actions, effectively avoiding misrecognition caused by environmental interference. Conversely, in good lighting conditions, appropriately lowering the recognition threshold ensures the smoothness and responsiveness of user operations. This adaptive threshold adjustment mechanism enables the system to maintain high recognition accuracy and robustness in complex and changing environments, thus ensuring the reliable execution of global mode switching commands.
[0099] Through the above technical solution, this application can significantly improve the robustness and accuracy of gesture recognition-based interaction design methods in complex environments. Specifically, by sensing the stability of ambient lighting in real time and dynamically adjusting the non-dominant hand recognition threshold accordingly, the system can effectively cope with the impact of lighting changes on gesture recognition accuracy, thereby reducing the occurrence of misoperations and missed operations. This adaptive recognition mechanism allows users to obtain a stable and reliable global mode switching experience when performing high-precision selection tasks under different lighting conditions, greatly enhancing the system's practicality and user satisfaction. Compared with traditional solutions that do not consider environmental factors, this application introduces an ambient lighting stability index as an adjustment parameter, making the recognition of non-dominant hand-assisted confirmation actions more intelligent and environmentally adaptable, thereby improving the overall reliability of the interaction system and the user experience.
[0100] The following is a specific example to illustrate this.
[0101] Suppose a user is performing a high-precision selection task in a virtual reality (VR) environment, such as precisely selecting a tiny part on a 3D model. The system is currently locked. The user wants to switch to another global mode (e.g., from "selection mode" to "movement mode"), so their dominant hand forms a pre-defined global mode switching command gesture. Simultaneously, the user's non-dominant hand performs an auxiliary confirmation action, such as lightly pinching their index finger and thumb. The system first acquires image data of the user's current environment using the VR headset's built-in camera. Suppose the room's lighting suddenly dims or there is strong glare. The system immediately analyzes this image data, calculating the current ambient lighting stability index. If this index indicates unstable or poor lighting conditions, the system accordingly increases the threshold required to recognize the non-dominant hand's auxiliary confirmation action. This means that auxiliary confirmation actions performed by the non-dominant hand need to reach a higher confidence level to be recognized as valid by the system. If the non-dominant hand's movement can still be accurately recognized at the increased threshold, the global mode switching command is executed. Conversely, if the non-dominant hand's movement fails to reach the increased threshold due to poor lighting, the system will not execute the global mode switching command, thus avoiding erroneous operations under uncertain conditions. On the other hand, if the ambient lighting stability index indicates good and stable lighting conditions, the system may appropriately lower the non-dominant hand recognition threshold, allowing users to complete assisted confirmation actions more easily and sensitively, thereby improving the smoothness of the interaction. In this way, the solution of this application can intelligently adjust the recognition strategy according to actual environmental conditions, ensuring the accuracy and reliability of the global mode switching command.
[0102] Specifically, the steps described above for obtaining image data of the user's current environment and obtaining the ambient light stability index based on the image data can be further implemented in the following ways.
[0103] In the locked state, when the user's dominant hand is detected to be forming a gesture indicating a global mode switching command, and simultaneously the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the steps for executing the global mode switching command include:
[0104] The system acquires image data of the user's current environment and performs brightness and contrast analysis on the environmental image data to obtain the ambient light stability index.
[0105] Specifically, acquiring image data of the user's current environment refers to capturing visual information about the user's physical environment in real time using a system-equipped visual sensor (such as a camera). Image data can be a series of continuous frames reflecting dynamic changes in ambient lighting. Brightness and contrast analysis of the environmental image data involves processing the acquired image data to quantify the stability of ambient lighting. Brightness analysis may involve calculating the average pixel intensity or brightness histogram of the image, while contrast analysis may involve calculating the brightness differences between different regions of the image. For example, the global brightness mean and standard deviation of the image, as well as the contrast values of local areas, can be calculated. An ambient lighting stability index can be derived from these analysis results. For example, by setting a threshold or using a specific algorithm, the fluctuation range of the brightness mean and contrast values can be mapped to an index representing stability. When ambient lighting changes drastically, the index value is lower, and vice versa.
[0106] This application accurately quantifies the stability of ambient lighting by analyzing the brightness and contrast of image data of the user's current environment. In gesture recognition interactions, ambient lighting conditions significantly impact recognition accuracy, especially when recognizing auxiliary confirmation actions using the user's non-dominant hand. Unstable lighting can make hand image feature extraction difficult, thus affecting recognition accuracy. By obtaining the ambient lighting stability index, the system can assess the potential interference of the current environment on gesture recognition. When lighting stability is poor, the system can adjust the recognition threshold for the user's non-dominant hand accordingly—for example, increasing the recognition threshold to reduce the false recognition rate, or temporarily disabling certain high-precision recognition functions under extremely poor lighting conditions—thereby ensuring the robustness and reliability of gesture recognition under various environmental conditions.
[0107] Through the above technical solution, the system can dynamically sense and evaluate the lighting conditions of the user's environment and quantify them into an ambient lighting stability index. Obtaining this index provides a precise basis for subsequently adaptively adjusting the user's non-dominant hand recognition threshold based on environmental conditions. Compared to solutions that do not consider changes in ambient lighting, this application effectively avoids misjudgments or omissions in gesture recognition caused by unstable lighting, significantly improving the accuracy and stability of gesture recognition in complex or dynamic lighting environments, thereby optimizing the user's interactive experience in high-precision selection tasks.
[0108] However, in practical applications, even under stable ambient lighting conditions, physiological tremors or non-instructional minor movements of the user's non-dominant hand may interfere with the recognition of assisted confirmation actions, leading to misjudgment or missed judgment by the system.
[0109] In this regard, this application further proposes that, in the locked state, when the user's dominant hand is detected to be forming a posture indicating a global mode switching command, and simultaneously the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the steps for executing the global mode switching command include:
[0110] Continuously track the position data of key joints of the user's non-dominant hand in three-dimensional space when no explicit gesture command is executed;
[0111] Based on the position data in three-dimensional space, calculate the root mean square of the minute displacement of the tracked joint position within a preset time window;
[0112] The root mean square of displacement is compared with a preset displacement threshold to identify physiological micro-tremors and normal resting state of the user's non-dominant hand.
[0113] The stability index of the non-dominant hand is obtained based on the physiological micro-tremor of the non-dominant hand and its normal resting state.
[0114] In the locked state, when the user's dominant hand is detected to be forming a global mode switching command posture, and at the same time the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the user's non-dominant hand recognition threshold is adjusted according to the ambient light stability index and the non-dominant hand stability index, and the decision is made whether to execute the global mode switching command based on the adjustment result.
[0115] Specifically, continuously tracking the position data of key joints in a user's non-dominant hand in three-dimensional space when no explicit gesture command is executed refers to acquiring the motion trajectory and posture information of the user's non-dominant hand in three-dimensional space in real time through gesture recognition devices (such as depth cameras, inertial measurement units, IMUs, etc.). Key joints can include fingertips, finger roots, and other parts sensitive to fine movements. This tracking process is performed when the user's non-dominant hand is not executing any explicit gesture command to capture its subtle movements in a natural state. Key joints of the non-dominant hand can include the fingertips and finger roots of the thumb, index finger, and middle finger.
[0116] Furthermore, calculating the root mean square (RMSD) of the minute displacements of the tracked joint positions within a preset time window refers to the statistical analysis of key joint position data collected within a specific time period (e.g., 0.5 seconds to 2 seconds). The root mean square displacement (RMSD) is an indicator that measures data volatility and can effectively reflect the degree of minute hand tremors or shaking.
[0117] In practical applications, comparing the root mean square displacement with a preset displacement threshold aims to distinguish between physiological tremors and a normal, still hand state. The preset displacement threshold can be calibrated based on extensive user data and experimental results to accommodate individual physiological differences. When the root mean square displacement exceeds this threshold, it indicates a relatively obvious involuntary hand movement, possibly a physiological tremor; conversely, it is considered that the hand is in a relatively still state.
[0118] The non-dominant hand stability index, derived from the physiological micro-tremors and normal resting state of the non-dominant hand, quantifies the stability of the non-dominant hand based on the aforementioned comparison results. For example, the stability index can be set to a value between 0 and 1, where 1 represents complete rest and 0 represents severe tremor. A higher index indicates greater stability of the non-dominant hand.
[0119] Ultimately, in the locked state, when the system detects the user's dominant hand forming a global mode switching command gesture and simultaneously detects the user's non-dominant hand performing a preset auxiliary confirmation action, it comprehensively considers both the ambient light stability index and the newly acquired non-dominant hand stability index. These two indices are used together to adjust the user's non-dominant hand recognition threshold, thereby more accurately determining the effectiveness of the auxiliary confirmation action. For example, in cases of unstable ambient light or a low non-dominant hand stability index (i.e., slight hand tremors), the recognition threshold can be appropriately increased to avoid false judgments; conversely, in cases of stable ambient light and a high non-dominant hand stability index, the recognition threshold can be appropriately decreased to improve recognition sensitivity.
[0120] This application addresses the inaccuracy issues that may arise from relying solely on the ambient light stability index by introducing an assessment of the stability of the user's non-dominant hand. Specifically, by continuously tracking the positional data of key joints in the non-dominant hand in three-dimensional space and calculating the root mean square of their minute displacements within a preset time window, the system can quantify physiological tremors or uninstructed movements of the hand. By comparing this root mean square displacement with a preset displacement threshold, the system can accurately distinguish between physiological tremors and normal stillness, thus obtaining a non-dominant hand stability index that reflects the hand's inherent stability. When the system needs to determine whether to execute a global mode switching command, it considers both the ambient light stability index and the non-dominant hand stability index to jointly adjust the user's non-dominant hand recognition threshold. This dual-consideration mechanism allows the system to more comprehensively evaluate the reliability of assisted confirmation actions. For example, when the user's non-dominant hand exhibits physiological tremors, even with stable ambient light, the system can increase the recognition threshold by using a lower non-dominant hand stability index, thereby avoiding misjudging tremors as assisted confirmation actions; conversely, when the non-dominant hand is very stable, the system can appropriately lower the threshold to improve the response speed and accuracy to genuine assisted confirmation actions.
[0121] Through the above technical solution, this application effectively overcomes the limitations of relying solely on environmental factors to adjust gesture recognition thresholds. By introducing real-time evaluation of the stability of the user's non-dominant hand, the system can more precisely distinguish between user intent and unconscious physiological movements, significantly reducing the risk of misjudging or missing auxiliary confirmation actions in complex or highly individualized scenarios. This not only improves the accuracy and reliability of global mode switching command execution but also greatly enhances the user's interactive experience when performing high-precision selection tasks, enabling the system to provide more stable and reliable gesture interaction services in various practical application environments.
[0122] Suppose a user is performing a high-precision selection task in virtual reality (VR) that requires intense focus, such as precisely selecting a tiny part in a 3D model. The system is currently locked. The user's dominant hand forms a gesture indicating a global mode switching command, but the user's non-dominant hand may experience slight physiological tremors due to prolonged operation or mild fatigue.
[0123] According to the scheme of this application, the system continuously tracks the position data of key joints of the user's non-dominant hand (such as the tip and base of the thumb and index finger) in three-dimensional space. Within a preset 1-second time window, the system calculates the root mean square of the minute displacements of these joint positions. If the calculated root mean square displacement is slightly higher than a preset resting threshold, the system identifies that the non-dominant hand is in a state of physiological microtremor and obtains a low non-dominant hand stability index (e.g., 0.6).
[0124] At the same time, the system also acquires image data of the current environment and analyzes it to obtain the ambient light stability index (e.g., 0.9, indicating stable light).
[0125] When determining whether to execute a global mode switching command, the system considers both indices. Because the non-dominant hand has a lower stability index, the system correspondingly increases the user's non-dominant hand recognition threshold. This means that the non-dominant hand needs to perform a more explicit and significant auxiliary confirmation action to be recognized as a valid command. If the tremor of the user's non-dominant hand is insufficient to reach the increased threshold, the system will not execute the global mode switching command, thus avoiding erroneous operations caused by physiological tremors. The global mode switching command will only be executed when the user's non-dominant hand overcomes its tremors and performs a sufficiently clear auxiliary confirmation action that reaches the new threshold.
[0126] However, in its implementation, considering only the ambient light stability index may be insufficient to fully address the complex and ever-changing real-world application scenarios. For example, the user's physiological state, such as fatigue or physiological tremors, may cause instability in the non-dominant hand when performing auxiliary confirmation actions, thus affecting the accuracy of recognition. If these issues are not addressed, adjusting the threshold solely based on the ambient light stability index may lead to misjudging the auxiliary confirmation action when the user's non-dominant hand exhibits physiological tremors, thereby affecting the accuracy of global mode switching commands and the user experience.
[0127] In response, this application further proposes a more refined threshold adjustment method, which improves the robustness of gesture recognition by comprehensively considering the ambient light stability index and the user's non-dominant hand stability index.
[0128] This application further proposes that, in the locked state, when the user's dominant hand is detected to be forming a posture indicating a global mode switching command, and simultaneously the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the steps for executing the global mode switching command include:
[0129] Continuously track the position data of key joints of the user's non-dominant hand in three-dimensional space when no explicit gesture command is executed;
[0130] Based on the position data in three-dimensional space, calculate the spectral rate of the tracked joint position within a preset time window;
[0131] The spectral rate is compared with a preset spectral rate threshold to identify physiological microtremors and normal resting states in the user's non-dominant hand.
[0132] The stability index of the non-dominant hand is obtained based on the physiological micro-tremor of the non-dominant hand and its normal resting state.
[0133] In the locked state, when the user's dominant hand is detected to be forming a global mode switching command posture, and at the same time the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the user's non-dominant hand recognition threshold is adjusted according to the ambient light stability index and the non-dominant hand stability index, and the decision is made whether to execute the global mode switching command based on the adjustment result.
[0134] Specifically, continuously tracking the position data of key joints of a user's non-dominant hand in three-dimensional space when no explicit gesture command is executed refers to acquiring the position and posture information of the user's non-dominant hand in three-dimensional space in real time through depth sensors, inertial measurement units (IMUs), or other hand tracking devices. Key joints may include, but are not limited to, the fingertips and bases of the thumb, index finger, and middle finger. The minute movements of these joints can effectively reflect hand tremors or shaking.
[0135] The process involves calculating the spectral rate of the tracked joint position within a preset time window based on the three-dimensional spatial location data. This can be understood as performing a Fourier transform on the time-domain signal of the key joint position data over a period of time to analyze its frequency components. The spectral rate can quantify the frequency characteristics of minute hand movements; for example, physiological tremors are typically characterized by energy concentration within a specific frequency range. The length of the preset time window can be set according to the actual application scenario and physiological characteristics, for example, from 0.5 seconds to 2 seconds.
[0136] In practical applications, the spectral rate is compared with a preset spectral rate threshold to identify physiological tremors and normal resting states in the user's non-dominant hand. For example, when the spectral rate exceeds a certain preset threshold, it indicates that the hand has tremors beyond the normal resting range, which can be identified as a physiological tremor state; conversely, it is identified as a normal resting state. The preset spectral rate threshold can be trained and calibrated using a large amount of user data.
[0137] Furthermore, based on the physiological tremors and normal resting state of the non-dominant hand, a non-dominant hand stability index is obtained. The purpose is to quantify the hand's inherent stability into a numerical value. For example, the stability index is higher when identified as a normal resting state and lower when identified as a physiological tremor state. This index can be a continuous value, reflecting the degree of hand stability.
[0138] Therefore, in the locked state, when the system detects a gesture indicating a global mode switching command from the user's dominant hand, and simultaneously detects the user's non-dominant hand performing a preset auxiliary confirmation action, the system adjusts the recognition threshold for the user's non-dominant hand based on the ambient light stability index and the stability index of the non-dominant hand. The system then decides whether to execute the global mode switching command based on the adjustment result. This means the system no longer relies solely on environmental factors, but comprehensively considers both the environment and the user's physiological state, thereby dynamically adjusting the sensitivity of the recognition of auxiliary confirmation actions. For example, when the ambient light is unstable or the stability index of the user's non-dominant hand is low, the recognition threshold can be appropriately increased to avoid false triggering.
[0139] This application overcomes the limitations of relying solely on the ambient light stability index for threshold adjustment by introducing an assessment of the stability of the user's non-dominant hand. Specifically, by continuously tracking the positional data of key joints in the non-dominant hand in three-dimensional space and calculating their spectral rate within a preset time window, the system can objectively quantify the subtle movement characteristics of the hand. By comparing this spectral rate with a preset threshold, it can accurately identify whether the user's non-dominant hand is in a state of physiological micro-tremor or a normal resting state, thus obtaining the non-dominant hand stability index. This stability index, combined with the ambient light stability index, is used to dynamically adjust the user's non-dominant hand recognition threshold. When ambient light conditions are poor or the user's non-dominant hand exhibits physiological micro-tremor, the system will correspondingly increase the recognition threshold, requiring more explicit and stable gestures for the auxiliary confirmation action to be recognized, thereby effectively suppressing misrecognition caused by environmental or physiological factors. Conversely, under favorable environmental conditions and with a stable hand, the threshold can be appropriately lowered to ensure the sensitivity of the interaction.
[0140] Through the above technical solution, this application can significantly improve the robustness and accuracy of gesture recognition-based interaction design methods. By comprehensively considering the ambient light stability index and the non-dominant hand stability index, the system can more intelligently and adaptively adjust the non-dominant hand recognition threshold, effectively avoiding the accidental triggering of global mode switching commands due to misrecognition of auxiliary confirmation actions in complex environments or when the user is in poor physiological condition. This not only improves the operational reliability of users when performing high-precision selection tasks, but also greatly improves the user experience, enabling the gesture interaction system to exhibit higher stability and usability in various practical application scenarios.
[0141] In some implementations, in order to more accurately assess the stability of the user's non-dominant hand, the above-mentioned continuous tracking of the position data of the key joints of the user's non-dominant hand in three-dimensional space when no explicit gesture command is executed, wherein the key joints of the non-dominant hand include the fingertips and bases of the thumb, index finger, and middle finger.
[0142] This application's solution explicitly designates the key joints of the non-dominant hand as the fingertips and bases of the thumb, index finger, and middle finger. This allows the system to focus on the areas most influential on hand stability when continuously tracking the user's non-dominant hand. Because these joints play a central role in fine hand movements and posture maintenance, precise tracking and analysis of their positional data can more effectively capture physiological tremors or subtle fluctuations in the non-dominant hand at rest. This focused tracking method avoids data redundancy and noise interference that might result from indiscriminate tracking of all joints in the entire hand, thereby improving the accuracy and reliability of the non-dominant hand stability index calculation.
[0143] This application also discloses an interaction design system based on gesture recognition, such as Figure 2 As shown, the system includes:
[0144] Information acquisition module 1 is used to acquire user hand movement and posture information;
[0145] Task judgment module 2 is used to determine whether the user has started to perform a high-precision selection task based on hand movement and posture information;
[0146] State switching module 3 is used to switch the system working state to the locked state in response to the judgment that the user has started to execute the high-precision selection task;
[0147] The instruction suppression module 4 is used to prevent the global mode switching instruction from being executed immediately when the user's hand is detected to form a global mode switching instruction in the locked state.
[0148] The dual confirmation module 5 is used to execute the global mode switching command when the user's dominant hand forms a posture that indicates a global mode switching command and the user's non-dominant hand performs a preset auxiliary confirmation action in the locked state.
[0149] Local priority module 6 is used to prioritize responding to local operation commands directly related to the high-precision selection task while in a locked state; and
[0150] The status release module 7 is used to release the locked status when the high-precision selection task is completed or explicitly canceled.
[0151] Through this technical solution, this application introduces a locked state and a dual confirmation mechanism when the user performs a high-precision selection task, which effectively avoids the unexpected switching of the global mode caused by gesture misrecognition. At the same time, it prioritizes the response to local operation commands, ensuring the smooth execution of high-precision tasks and significantly improving the reliability of the interaction and the user experience.
[0152] The above are merely embodiments of this application and are not intended to limit the scope of protection of this application.
Claims
1. An interaction design method based on gesture recognition, characterized in that, include: Obtain user hand movement and posture information; Based on the hand movement and posture information, determine whether the user has started to perform a high-precision selection task; In response to the determination that the user has started to execute the high-precision selection task, the system working state is switched to the locked state; In the locked state, when the user's hand is detected to form a gesture that constitutes a global mode switching command, the global mode switching command is not executed immediately; In the locked state, when the user's dominant hand is detected to be in a posture that forms the global mode switching command, and at the same time the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the global mode switching command is executed. In the locked state, priority is given to responding to local operation commands directly related to the high-precision selection task; and The locking state is released when the high-precision selection task is completed or explicitly cancelled.
2. The interaction design method based on gesture recognition according to claim 1, characterized in that, In the locked state, when the user's dominant hand is detected to be forming a posture indicating the global mode switching command, and simultaneously the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the steps for executing the global mode switching command include: In the locked state, the system senses the user's hand movements and posture information and determines whether the user has started to perform a high-precision selection task. In response to the determination that the user has started to execute the high-precision selection task, the system working state is switched to the locked state; In the locked state, it is sensed whether the dominant hand has formed a posture that triggers a global mode switching command; In the locked state, when the dominant hand is perceived to be in a posture that forms a global mode switching command, a visual cue is generated and displayed in a non-critical position in the 3D scene display area; In the locked state, it is monitored whether the user's gaze focus moves and stays on the visual cue area; In the locked state, when the dominant hand is perceived to form a posture indicating a global mode switching command, and the user's non-dominant hand is simultaneously perceived to perform a preset auxiliary confirmation action, and the user's gaze is detected to move and remain on the visual cue area, the global mode switching command is executed.
3. The interaction design method based on gesture recognition according to claim 2, characterized in that, The step of generating a visual cue and displaying it in a non-critical position in the 3D scene display area when the dominant hand is perceived to be in a posture that triggers a global mode switching command in the locked state includes: In the locked state, when the dominant hand is perceived to be forming a global mode switching command, the visual complexity information of the current three-dimensional scene is obtained. Obtain the position information of the user's gaze focus within the three-dimensional scene display area; Obtain the distribution density information of interactive elements in the three-dimensional scene; The display parameters of the visual cues are determined based on the visual complexity information of the three-dimensional scene, the user's gaze focus position information, and the distribution density information of the interactive elements. Based on the determined display parameters, the visual cues are generated and displayed at non-critical locations in the 3D scene display area.
4. The interaction design method based on gesture recognition according to claim 3, characterized in that, The step of generating a visual cue and displaying it in a non-critical position in the 3D scene display area when the dominant hand is perceived to be in a posture that triggers a global mode switching command in the locked state includes: In the locked state, when the dominant hand is perceived to be forming a global mode switching command, the visual complexity information of the current three-dimensional scene is obtained. Obtain the position information of the user's gaze focus within the three-dimensional scene display area; Obtain the distribution density information of interactive elements in the three-dimensional scene; Obtain the noise level of the user's current environment; To obtain the vibration feedback capability of the haptic feedback device worn by the user; Based on the visual complexity information of the three-dimensional scene, the user's gaze focus position information, the distribution density information of the interactive elements, the noise level, and the vibration feedback capability, the display parameters and non-visual feedback parameters of the visual prompt are determined. Based on the determined display parameters, the visual cues are generated and displayed at non-critical locations in the 3D scene display area; Based on the determined non-visual feedback parameters, tactile feedback is emitted through the tactile feedback device.
5. The interaction design method based on gesture recognition according to claim 1, characterized in that, In the locked state, when the user's dominant hand is detected to be forming a posture indicating the global mode switching command, and simultaneously the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the steps for executing the global mode switching command include: Obtain image data of the user's current environment, and obtain the ambient light stability index based on the image data of the environment; In the locked state, when the user's dominant hand is detected to be in a posture that forms the global mode switching command, and at the same time the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the user's non-dominant hand recognition threshold is adjusted according to the ambient light stability index, and the decision is made whether to execute the global mode switching command based on the adjustment result.
6. The interaction design method based on gesture recognition according to claim 5, characterized in that, The step of acquiring image data of the user's current environment and obtaining the ambient light stability index based on the image data includes: The system acquires image data of the user's current environment and performs brightness and contrast analysis on the image data to obtain the ambient light stability index.
7. The interaction design method based on gesture recognition according to claim 5, characterized in that, In the locked state, when the user's dominant hand is detected to be forming a posture indicating the global mode switching command, and simultaneously the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the steps for executing the global mode switching command include: Continuously track the position data of key joints of the user's non-dominant hand in three-dimensional space when no explicit gesture command is executed; Based on the position data in the three-dimensional space, calculate the root mean square of the minute displacement of the tracked joint position within a preset time window; The root mean square of the displacement is compared with a preset displacement threshold to identify physiological micro-tremors and normal resting states of the user's non-dominant hand. The stability index of the non-dominant hand is obtained based on the physiological micro-tremors and normal resting state of the non-dominant hand. In the locked state, when the user's dominant hand is detected to be in a posture that forms the global mode switching command, and at the same time the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the user's non-dominant hand recognition threshold is adjusted according to the ambient light stability index and the non-dominant hand stability index, and the decision is made whether to execute the global mode switching command based on the adjustment result.
8. The interaction design method based on gesture recognition according to claim 5, characterized in that, In the locked state, when the user's dominant hand is detected to be forming a posture indicating the global mode switching command, and simultaneously the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the steps for executing the global mode switching command include: Continuously track the position data of key joints of the user's non-dominant hand in three-dimensional space when no explicit gesture command is executed; Based on the position data in the three-dimensional space, the spectral rate of the tracked joint position within a preset time window is calculated; The spectral rate is compared with a preset spectral rate threshold to identify physiological microtremors and normal resting states in the user's non-dominant hand. The stability index of the non-dominant hand is obtained based on the physiological micro-tremors and normal resting state of the non-dominant hand. In the locked state, when the user's dominant hand is detected to be in a posture that forms the global mode switching command, and at the same time the user's non-dominant hand is detected to be performing a preset auxiliary confirmation action, the user's non-dominant hand recognition threshold is adjusted according to the ambient light stability index and the non-dominant hand stability index, and the decision is made whether to execute the global mode switching command based on the adjustment result.
9. An interaction design method based on gesture recognition according to claim 7 or 8, characterized in that, The continuous tracking of the position data of key joints of the user's non-dominant hand in three-dimensional space when no explicit gesture command is executed, wherein the key joints of the non-dominant hand include the fingertips and bases of the thumb, index finger, and middle finger.
10. An interaction design system based on gesture recognition, characterized in that, The system includes: The information acquisition module is used to acquire the user's hand movements and posture information; The task judgment module is used to determine whether the user has started to perform a high-precision selection task based on the hand movement and posture information. The state switching module is used to switch the system working state to the locked state in response to the determination that the user has started to execute the high-precision selection task. The instruction suppression module is used to prevent the immediate execution of the global mode switching instruction when the user's hand is detected to form a global mode switching instruction in the locked state. The dual confirmation module is used to execute the global mode switching command when, in the locked state, it is detected that the user's dominant hand forms the posture of the global mode switching command, and at the same time, it is detected that the user's non-dominant hand performs a preset auxiliary confirmation action. A local priority module is configured to, in the locked state, prioritize responding to local operation commands directly related to the high-precision selection task; and The status release module is used to release the locked state when the high-precision selection task is completed or explicitly canceled.