Virtual operation interaction method and system based on intelligent display equipment

By combining gesture and eye-tracking technologies in virtual operation interaction, this method solves the problem of poor user experience in AR/VR, achieving natural and accurate menu operation, and is suitable for augmented reality, virtual reality, and smart home devices.

CN121742718APending Publication Date: 2026-03-27SICHUAN SMART KIDS TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing AR or VR interaction methods suffer from poor user experience. Traditional operation methods, such as voice recognition, are easily affected by environmental noise, and buttons and touchpads are unresponsive and have large errors, which affect immersion and operation efficiency.

Method used

It adopts a virtual operation interaction method based on intelligent display devices, combined with gesture and eye tracking technology. Through gesture image recognition and eye data matching, it calculates the hand-eye relationship to realize natural and intuitive menu operation, supports multiple gestures and left and right hand operations, introduces drag and zoom actions, and designs an anti-accidental touch mechanism.

Benefits of technology

It improves the efficiency and accuracy of interaction, enhances the user experience, and provides intuitive, contactless interaction, suitable for augmented reality, virtual reality, and smart home scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121742718A_ABST
    Figure CN121742718A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of virtual display, and discloses a virtual operation interaction method and system based on intelligent display equipment, and the method comprises the steps: obtaining a gesture image and an eye image in real time; respectively identifying and calculating a corresponding relation between the first hand feature and the first eye feature; if yes, entering a simple menu, and displaying a first-level menu; continuing to judge whether a second hand feature exists or not, and if yes, displaying a second-level menu; judging whether the secondary menu is a complex menu or not, if yes, continuing to judge whether a third hand feature exists or if yes, locking the secondary menu; after the second-level menu is locked, whether a fourth hand feature exists or not is judged, and if yes, the position and / or the size of the second-level menu are / is adjusted according to the fourth hand feature; wherein the fourth hand feature comprises dragging and / or zooming actions. According to the method, errors in the virtual operation interaction process can be reduced, and the interaction experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of virtual display, and in particular to a virtual operation interaction method and system based on an intelligent display device. Background Technology

[0002] In existing AR or VR interaction methods, touchpads, voice, buttons, and touch screens are typically used to operate virtual images, text, videos, and other content.

[0003] Using the existing interaction methods described above may result in a poor user experience. For example, voice recognition may be affected by ambient noise, and button controls and touchpads may be unresponsive or have large errors when operating virtual objects. These traditional operations are relatively slow and complex, and they can also disrupt the immersive experience of AR or VR, all of which will affect the user's operating experience and satisfaction.

[0004] Therefore, a gesture operation interaction method based on image recognition technology is needed to reduce errors in the virtual operation interaction process and improve the interactive experience. Summary of the Invention

[0005] To reduce errors in the virtual operation and interaction process and improve the interactive experience, this application provides a virtual operation and interaction method and system based on a smart display device.

[0006] Firstly, this application provides a virtual operation interaction method based on an intelligent display device, employing the following technical solution:

[0007] A virtual operation interaction method based on an intelligent display device includes the following steps:

[0008] Real-time acquisition of gesture images;

[0009] Gesture data is identified based on the gesture image;

[0010] The first hand feature is matched based on the gesture data;

[0011] Real-time acquisition of eye images;

[0012] Eye data is identified based on the described eye image;

[0013] A first eye feature is matched based on the eye data;

[0014] Calculate the correspondence between the first hand feature and the first eye feature;

[0015] If the correspondence matches a preset correspondence set, then enter the simple menu and display the first-level menu; determine whether a second hand feature is matched in the gesture data, and if so, display the second-level menu;

[0016] Determine whether the secondary menu is a complex menu. If so, determine whether a third hand feature is matched in the gesture data or whether a locking signal is received. If a matching or locking signal is received, attach the secondary menu to the current position and lock the coordinates.

[0017] By adopting the above technical solutions and combining gesture and eye-tracking technologies, interactions such as looking at a specific point while making gestures are more natural and intuitive. Users can quickly access and adjust menus through simple gaze and gesture operations, improving the efficiency and convenience of interaction. Simultaneously, through locking and adjustment mechanisms, users can more precisely control the display and layout of complex menus, enhancing the user experience.

[0018] Optionally, the method further includes:

[0019] After locking the coordinates of the secondary menu, it is determined whether a fourth hand feature is matched in the gesture data. If a fourth hand feature is matched, the position and / or size of the secondary menu is adjusted according to the fourth hand feature. The fourth hand feature includes dragging and / or scaling actions.

[0020] By adopting the above technical solutions, the introduction of drag-and-zoom gestures provides more possibilities for the display of secondary menus. Users can customize the menu layout according to screen size, resolution, or personal preferences to better suit their usage habits. Especially for users with poor eyesight or limited hand dexterity, adjusting the menu position and size through dragging and zooming makes operation easier and more comfortable. Therefore, introducing gesture operations such as dragging and zooming to adjust the position and size of secondary menus not only enhances the user experience but also improves operational efficiency and flexibility.

[0021] Optionally, the method further includes:

[0022] If the correspondence does not match the preset correspondence set, determine whether the current menu is a spatial menu;

[0023] If it is a spatial menu, then continue to acquire gesture images and eye images in real time; wherein, the spatial menu is a secondary menu attached to the current position and locked to the world space, and can be dragged and / or scaled according to the fourth hand feature;

[0024] If it is not a spatial menu, the current menu disappears, and the acquisition of gesture and eye images continues in real time.

[0025] By adopting the above technical solution, when dealing with situations that do not conform to the preset corresponding set, a smoother and more natural interactive experience can be provided to users by distinguishing whether it is a spatial menu. For spatial menus, users can continue to fine-tune the menu through gestures without interrupting the interaction flow. For non-spatial menus, the menu disappears, guiding users to perform the interaction again, thus improving the flexibility and accuracy of the interaction.

[0026] Optionally, the method further includes:

[0027] If the correspondence matches a preset correspondence set, determine whether a menu has already been displayed;

[0028] If a menu is already displayed, determine whether the current menu is a spatial menu; if it is a spatial menu, determine whether the gesture data matches a fifth hand feature; if the fifth hand feature is matched, unlock the current spatial menu, move the spatial menu to the position of the palm, and move with the position of the hand feature; continue to acquire gesture images and eye images in real time; if it is not a spatial menu, the current menu moves with the position of the hand feature.

[0029] If no menu is displayed, the system will enter a simple menu and display the first-level menu.

[0030] By adopting the above technical solutions, not only is the flexibility and responsiveness of the interaction improved, but users are also provided with a more intuitive and natural operating experience. Users can easily switch between different menus, adjust the position and size of menus, and unlock and lock spatial menus as needed, thereby interacting with smart display devices more efficiently.

[0031] Optionally, the method further includes:

[0032] After displaying the second-level menu, it is determined whether the gesture data matches a sixth hand feature. If a sixth hand feature is matched or a return command is received, the first-level menu is displayed.

[0033] In the step of determining whether the secondary menu is a complex menu, if it is not a complex menu, the current menu moves with the position of the hand features; and the gesture image and eye image are acquired in real time.

[0034] By adopting the above technical solutions, the continuity and convenience of the user experience are improved; users can easily switch between different menu levels as needed and quickly return to the primary menu when necessary. At the same time, the hand-tracking positional movement function for non-complex menus also increases the naturalness and intuitiveness of the interaction.

[0035] Optionally, the correspondence between the first hand feature and the first eye feature is calculated; if the correspondence matches a preset correspondence set, the step of entering the simple menu includes the following steps:

[0036] The first hand features include finger position, gesture type, and palm orientation;

[0037] The first eye feature includes eye position and gaze direction;

[0038] Based on the finger position and the eye position, determine whether the finger position and the eye position are both within the corresponding position range;

[0039] If both are true, continue to judge the first hand feature and the first eye feature;

[0040] Determine whether the gesture type is a preset first gesture or a second gesture, and determine whether the angle between the palm orientation and the line of sight is within the corresponding preset angle range. If both are true, then the correspondence is determined to conform to the preset correspondence set.

[0041] By adopting the above technical solution, the system determines whether a user's hand or eye movements are preset based on a combination of these movements. If both are true, the system triggers the action and executes the corresponding instructions based on the user's operation, displaying the corresponding menu to achieve interactive functionality.

[0042] Optionally, the step of recognizing gesture data based on the gesture image and matching a first hand feature based on the gesture data includes:

[0043] Feature information is extracted from the gesture image using image processing techniques, and gesture data is matched using the feature information;

[0044] Alternatively, feature information can be extracted from the gesture image, and mirrored to obtain mirror information; gesture data can then be matched using the mirror information.

[0045] The gesture data is input into a pre-established database of gesture templates. The features corresponding to the gesture data are compared one by one with the gesture features in the database, and all similarity values ​​are calculated.

[0046] The hand gesture image with the highest similarity value is selected as the first hand feature.

[0047] By adopting the above technical solution, the system receives and filters user gesture information in real time, ensuring accurate recognition. The gesture recognition system can distinguish between the user's left and right hands through symmetrical comparison and recognize various gestures, such as palm-up and palm-facing gestures. The system can recognize multiple gestures and supports both left and right-handed operation, increasing its flexibility and adaptability. It solves the problem that existing gesture recognition systems typically only recognize a few gestures and have poor adaptability to both left and right hands. This invention overcomes these limitations through rich gesture recognition and left- and right-handed compatibility.

[0048] Optionally, the method further includes:

[0049] In the step of acquiring gesture images and eye images in real time, multiple gesture images and eye images are acquired within a set time period;

[0050] The correspondence is calculated multiple times. If the correspondence matches the preset correspondence set, the simple menu is entered; otherwise, it is judged as a mis-touch.

[0051] Alternatively, when a hardware trigger signal is received, the user can enter a simple menu.

[0052] By adopting the above technical solution, the existing gesture interaction system is susceptible to environmental interference and unintentional user actions, leading to frequent misoperations. An anti-mistouch mechanism is designed to prevent interference from unintentional user misoperations by repeatedly confirming and matching gestures. Furthermore, the system only responds to gesture operations under specific conditions, such as when the palm is facing upwards and the user is looking at the palm, further improving the accuracy of operation. This invention overcomes this deficiency and improves the system's reliability through strict gesture condition matching. Alternatively, button-triggered menu operations can be used, which can reduce mis-triggers but increases the hardware cost and complexity of the device.

[0053] Secondly, this application provides a virtual operation interaction system based on an intelligent display device, which adopts the following technical solution:

[0054] A virtual operation interaction system based on an intelligent display device includes a processor, wherein the processor executes the steps of the virtual operation interaction method based on an intelligent display device as described in any one of the preceding claims.

[0055] In summary, this application includes at least one of the following beneficial technical effects:

[0056] Gestures and eye contact are natural forms of human communication. Applying them to human-computer interaction makes the entire interaction process more natural and fluid, conforming to human cognitive habits. Utilizing gesture recognition, eye tracking, and interface interaction systems based on these inputs can provide users with intuitive, contactless interactive experiences, especially suitable for scenarios such as augmented reality (AR), virtual reality (VR), or smart homes. Attached Figure Description

[0057] Figure 1 This is a flowchart illustrating the steps of a virtual operation and interaction method based on intelligent display devices.

[0058] Figure 2 This is a flowchart of a virtual operation and interaction method based on intelligent display devices.

[0059] Figure 3 This is a diagram showing the first-level menu.

[0060] Figure 4 This is a diagram showing the display of a second-level menu. Detailed Implementation

[0061] The embodiments of this application are described in detail below, and examples of the embodiments are shown in the accompanying drawings.

[0062] In the description of this specification, the references to "certain embodiments," "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples" refer to specific features, structures, materials, or characteristics described in connection with the described embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0063] This application discloses a virtual operation interaction method based on a smart display device, referring to... Figure 1 and Figure 2 It includes the following steps:

[0064] The camera captures hand gesture images in real time; according to a set shooting frequency, a series of time-series image data are acquired.

[0065] Gesture data is identified from gesture images. Image processing techniques and machine learning algorithms are used to extract specific gesture data from gesture images, such as finger position, shape, and movement trajectory. For example, deep learning models, such as convolutional neural networks (CNNs), can be used to extract features and classify gesture images to identify specific gesture actions.

[0066] The first hand feature is matched based on the gesture data; the recognized gesture data is matched with the preset gesture feature library to determine the specific hand features, such as gestures like palm facing up and facing the user.

[0067] Real-time acquisition of eye images; eye images are captured by a camera.

[0068] Eye data is identified from eye images; eye data can be analyzed using eye tracking algorithms, such as pupil tracking and corneal reflection, to obtain eye movement trajectories, fixation points, and other eye data.

[0069] The system identifies the first eye feature based on eye data; it then matches the eye data with a preset set of eye features to identify the user's gaze direction.

[0070] The system calculates the correspondence between the first hand feature and the first eye feature. Combining the hand and eye features, the algorithm calculates their spatiotemporal correlation and consistency to determine the user's overall operational intent. For example, the corresponding hand and eye features might be: the first hand feature is a palm facing upwards or towards the user, and the first eye feature is the user's gesture of looking at the palm. The correspondence is usually based on the synergistic effect of the user's intent, action, and gaze direction; that is, the simultaneous occurrence of hand and eye movements serves as the triggering condition for judgment.

[0071] If the correspondence matches the preset correspondence set, that is, if it matches the corresponding action of the first hand feature and the first eye feature, then enter the simple menu and display the first-level menu; the first-level menu is as follows: Figure 3 The list of icons shown includes "Navigation Map" and "Product Catalog".

[0072] During the display of the primary menu, it is determined whether a second hand feature matches the gesture data. If a match is found, the secondary menu is displayed. This second hand feature, such as a click action, includes the position of the identified finger clicking, combined with the eye position; ensuring that the clicked icon, the clicking finger, and the eye position are all on a straight line in space. The secondary menu, for example... Figure 4 The image shows the menu interface displayed after clicking the "Product Catalog" icon. The secondary menu also includes a back icon. Similarly, the method also includes:

[0073] After displaying the second-level menu, the system checks if a sixth hand feature is matched in the gesture data. If a sixth hand feature is matched or a return command is received, the system returns to display the first-level menu. The sixth hand feature is the action of clicking the return button. Based on the recognition of the click action and location, the system can return from the second-level menu to the first-level menu. The return command can be received via voice recognition or by recognizing the direction of eye gaze, such as looking at the return button to confirm the return.

[0074] In addition, if no click action is detected during the display of the primary menu, the position of the current primary menu moves according to the coordinates of the hand, and the gesture image and eye image are acquired in real time.

[0075] During the display of the secondary menu, it is determined whether the secondary menu is a complex menu. If it is a complex menu, it is determined whether a third hand feature is matched in the gesture data or whether a locking signal is received. If a match is found or a locking signal is received, the secondary menu is attached to the current position and its coordinates are locked. Among them, the third hand feature, such as a palm flipping action, is recognized by the aforementioned image processing technology combined with algorithms. When a palm flipping action is recognized, the secondary menu is attached to the current position and its coordinates are locked, that is, the current secondary menu is locked in world space, forming a spatial menu.

[0076] While maintaining the coordinate lock of the secondary menu, it is determined whether a fourth hand feature is matched in the gesture data. If a match is found, the position and / or size of the secondary menu are adjusted according to the fourth hand feature; the fourth hand feature includes dragging and / or scaling actions. After dragging and / or scaling are completed, gesture images and eye images continue to be acquired in real time.

[0077] By combining gesture and eye-tracking technology, such as gazing at a point while making a gesture for interaction, the interaction becomes more natural and intuitive. Users can quickly access and adjust menus through simple gaze and gesture operations, improving the efficiency and convenience of interaction. Simultaneously, through locking and adjustment mechanisms, users can more precisely control the display and layout of complex menus, enhancing the user experience.

[0078] During the interaction, since hand and eye features are monitored in real time, the interaction method also includes determining the current menu type before displaying the first-level menu based on these features:

[0079] If the correspondence does not match the preset correspondence set, it means that the hand and eye features are not specific actions, and it is necessary to determine whether the current menu is a spatial menu.

[0080] If it is a spatial menu, then continue to acquire gesture images and eye images in real time; the spatial menu is a secondary menu that is attached to the current position and locked to the world space, and can be dragged and / or scaled according to the fourth hand features.

[0081] If it is not a spatial menu, the current menu disappears, and the acquisition of gesture and eye images continues in real time.

[0082] When dealing with situations that do not conform to the preset set, distinguishing between spatial menus and non-spatial menus provides users with a smoother and more natural interactive experience. For spatial menus, users can continue to fine-tune the menu using gestures without interrupting the interaction flow. For non-spatial menus, the menu disappears, guiding users to perform the interaction again, thus improving the flexibility and accuracy of the interaction.

[0083] To further determine the current menu type, the method also includes:

[0084] If the correspondence matches the preset correspondence set, determine whether a menu has already been displayed;

[0085] If a menu is already displayed, determine if it is a spatial menu. If it is, check if a fifth hand feature is matched in the gesture data. If a fifth hand feature is matched, unlock the current spatial menu, meaning the coordinates of the second-level menu are unlocked. The spatial menu moves to the location of the palm and follows the position of the hand feature. Continue acquiring gesture and eye images in real time. The fifth hand feature is a rapid fist clenching motion. If this motion is detected, it signifies unlocking, and the current menu can be moved. If no fifth hand feature is matched, continue acquiring gesture and eye images in real time.

[0086] If it is not a spatial menu, the current menu moves with the position of the hand feature.

[0087] If no menu is displayed, the system will enter a simple menu and display the first-level menu.

[0088] This not only improves the flexibility and responsiveness of the interaction but also provides users with a more intuitive and natural operating experience. Users can easily switch between different menus, adjust the position and size of menus, and unlock and lock spatial menus as needed, thus interacting with smart display devices more efficiently. It enhances the consistency and convenience of the user experience; users can easily switch between different menu levels and quickly return to the primary menu when needed. At the same time, the hand-tracking positional movement function for non-complex menus also increases the naturalness and intuitiveness of the interaction.

[0089] Calculate the correspondence between the first hand feature and the first eye feature; if the correspondence matches a preset correspondence set, the steps to enter the simple menu include the following:

[0090] The primary hand features include finger position, gesture type, and palm orientation; the primary eye features include eye position and gaze direction.

[0091] Based on the finger position and eye position, it is determined whether the finger position and eye position are both within the corresponding position range. Since the interactive action is a coordinated action based on eye movement and hand movement, the hand is first ensured to be within the field of vision before the judgment is made. Specifically, it is defined based on screen coordinates, image coordinates or the relative position of the user's body.

[0092] If both are true, meaning the hand is within the field of vision, continue to determine the first hand feature and the first eye feature;

[0093] It is determined whether the gesture type is a preset first gesture or a preset second gesture, and whether the angle between the palm's orientation and the gaze direction falls within a corresponding preset angle range. If both are true, then the correspondence is determined to conform to a preset correspondence set. For example, the first gesture is a palm facing upwards, and the second gesture is a palm facing the eye position, i.e., the palm faces the user, and the gaze direction is the direction corresponding to the user's gaze on the palm. The palm is taken as the starting point of the vector corresponding to the first or second gesture, and the direction the palm faces is the direction of the vector.

[0094] By combining specific hand and eye movements, the system determines whether the action is a preset action. If both are true, the action is triggered, and the system executes the corresponding instructions based on the user's operation, displaying the corresponding menu to achieve interactive functionality.

[0095] To recognize multiple gestures and support both left and right hand operations, thereby improving the accuracy of motion recognition, gesture data is identified from gesture images. The steps for matching the first hand feature based on the gesture data include:

[0096] Feature information is extracted from gesture images using image processing techniques, and gesture data is then matched using this feature information.

[0097] Alternatively, feature information can be extracted from the gesture image, and then mirrored to obtain mirrored information; the gesture data can then be matched using the mirrored information. That is, if the left and right hands are symmetrical, they will both be performing the same action.

[0098] The gesture data is input into a pre-established database of gesture templates. The features corresponding to the gesture data are compared one by one with the gesture features in the database, and the similarity values ​​are calculated.

[0099] The hand gesture image with the highest similarity value is selected as the first hand feature.

[0100] By receiving and filtering user gesture information in real time, the system ensures accurate recognition. The gesture recognition system can distinguish between the user's left and right hands through symmetrical comparison and recognize various gestures, such as palm-up and hand-facing gestures. The system can recognize multiple gestures and supports both left and right-handed operation, increasing its flexibility and adaptability. It solves the problem of poor adaptability to both left and right hands in existing gesture recognition systems. This invention overcomes these limitations through rich gesture recognition and left- and right-hand compatible design.

[0101] Because existing gesture interaction systems are easily affected by environmental interference and unintentional user actions, resulting in frequent misoperations, an anti-accidental touch mechanism was designed.

[0102] Method 1: Automatically triggered anti-accidental touch mechanism:

[0103] In the step of acquiring gesture images and eye images in real time, multiple gesture images and eye images are acquired within a set time period;

[0104] The correspondence is calculated multiple times. If the correspondence matches the preset correspondence set, the simple menu is entered; otherwise, it is judged as a mis-touch.

[0105] The anti-accidental touch mechanism avoids interference from unintentional user misoperations by repeatedly confirming and matching gestures. Furthermore, the system only responds to gesture operations under specific conditions, such as when the palm is facing upwards and the user is looking at the palm, further improving operational accuracy. This invention overcomes this deficiency and enhances system reliability through strict gesture condition matching.

[0106] Method 2: Manually triggered anti-accidental touch mechanism:

[0107] When a hardware trigger signal is received, such as a button-triggered menu operation, the user enters a simple menu. Button triggering can reduce false triggers, but it increases the hardware cost and complexity of the device. For example, buttons are used to expand and collapse the menu, arrow keys are used to select focus, and the O key is used to confirm.

[0108] The system incorporates an anti-accidental touch mechanism, employing multiple confirmations and matching of gestures to prevent interference from unintentional user errors. Furthermore, the system only responds to gesture operations under specific conditions, such as when the palm is facing upwards and the user is looking at the palm, further improving operational accuracy. This invention overcomes this deficiency and enhances system reliability through strict gesture condition matching. Alternatively, menu operations can be triggered using buttons, which can reduce accidental triggering but increases the hardware cost and complexity of the device.

[0109] In summary, this application has the following advantages:

[0110] Enhancing User Experience: By combining gesture recognition and eye-tracking technologies, the interaction method becomes closer to natural human communication. Users no longer need complex operating commands; they can interact with smart display devices simply through gestures and gaze, greatly improving ease of use and comfort.

[0111] Reduced accidental operations: By considering both hand and eye characteristics, the system can more effectively verify the user's intent. This dual verification mechanism significantly reduces the likelihood of accidental operations and improves the system's accuracy and reliability.

[0112] Flexible configuration: The preset corresponding set can be flexibly configured according to actual application scenarios to adapt to different interaction needs. This flexibility enables the method of this application to be widely used in various smart display devices, including AR / VR devices, tablets, smartphones, etc.

[0113] Rich gesture recognition: By receiving and filtering user gesture information in real time, the system can recognize a variety of gestures and supports both left- and right-handed operation. This rich gesture recognition capability increases the system's flexibility and adaptability, solving the problem that existing gesture recognition systems can usually only recognize a few gestures and have poor adaptability to both left- and right-handed operation.

[0114] Menu tracking function: Both primary and secondary menus can move according to the user's hand movements, meaning the menu position will move with the user's hand. This design allows the user to maintain natural line of sight and hand movements during operation, improving the smoothness and naturalness of the operation.

[0115] World Space Lock Function: After the secondary menu is expanded, if the gesture no longer meets the conditions or the user wishes to fix the menu position, the system can lock the menu in a specific location in world space. This locking function allows users to have a more stable interactive experience during operation, especially when performing operations such as dragging and zooming.

[0116] Dynamic Adjustment: During the process of locking the coordinates of the secondary menu, if the user adjusts the menu using specific gestures, such as dragging and zooming, the system will dynamically adjust the position and / or size of the menu based on these gesture characteristics. This dynamic adjustment capability allows users to flexibly adjust the menu layout according to their needs, improving the system's usability and flexibility.

[0117] This application also discloses a virtual operation interaction system based on an intelligent display device, including a processor, wherein the processor executes the steps of the virtual operation interaction method based on an intelligent display device as described in any of the above embodiments.

[0118] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A virtual operation interaction method based on an intelligent display device, characterized in that, Includes the following steps: Real-time acquisition of gesture images; Gesture data is identified based on the gesture image; The first hand feature is matched based on the gesture data; Real-time acquisition of eye images; Eye data is identified based on the described eye image; A first eye feature is matched based on the eye data; Calculate the correspondence between the first hand feature and the first eye feature; If the correspondence matches a preset correspondence set, the simple menu is entered and the first-level menu is displayed; it is determined whether a second hand feature is matched in the gesture data. If a second hand feature is matched, the second-level menu is displayed. Determine whether the secondary menu is a complex menu. If it is a complex menu, determine whether a third hand feature is matched in the gesture data or whether a locking signal is received. If a third hand feature is matched or a locking signal is received, attach the secondary menu to the current position and lock the coordinates. If it is not a complex menu, the current menu moves with the position of the hand feature. Continue to acquire gesture images and eye images in real time.

2. The virtual operation interaction method based on an intelligent display device according to claim 1, characterized in that, The method further includes: after locking the coordinates of the secondary menu, determining whether a fourth hand feature is matched in the gesture data; if a fourth hand feature is matched, adjusting the position and / or size of the secondary menu according to the fourth hand feature, wherein the fourth hand feature includes dragging and / or scaling actions.

3. The virtual operation interaction method based on an intelligent display device according to claim 1, characterized in that, The first hand feature includes palm facing upwards or palm facing the user; the second hand feature includes finger clicking actions, including single click, double click, or triple click actions; the third hand feature includes palm flipping actions.

4. The virtual operation interaction method based on an intelligent display device according to claim 1, characterized in that, The method also includes: If the correspondence does not match the preset correspondence set, determine whether the current menu is a spatial menu; If it is a spatial menu, then continue to acquire gesture images and eye images in real time; wherein, the spatial menu is a secondary menu attached to the current position and locked to the world space, and can be dragged and / or scaled according to the fourth hand feature; If it is not a spatial menu, the current menu disappears, and the acquisition of gesture and eye images continues in real time.

5. The virtual operation interaction method based on an intelligent display device according to claim 4, characterized in that, The method also includes: If the correspondence matches a preset correspondence set, determine whether a menu has already been displayed; If a menu is already displayed, determine whether the current menu is a spatial menu; if it is a spatial menu, determine whether the gesture data matches a fifth hand feature; if the fifth hand feature is matched, unlock the current spatial menu, move the spatial menu to the position of the palm, and move with the position of the hand feature; continue to acquire gesture images and eye images in real time; if it is not a spatial menu, the current menu moves with the position of the hand feature. If no menu is displayed, the system will enter a simple menu and display the first-level menu.

6. The virtual operation interaction method based on an intelligent display device according to claim 5, characterized in that, The fifth hand feature includes a rapid fist-clenching motion.

7. The virtual operation interaction method based on an intelligent display device according to claim 5, characterized in that, The method also includes: After displaying the second-level menu, it is determined whether the gesture data matches a sixth hand feature. If a sixth hand feature is matched or a return command is received, the first-level menu is displayed.

8. The virtual operation interaction method based on an intelligent display device according to claim 7, characterized in that, The sixth hand feature includes mobile phone clicking actions, including single click, double click, or triple click actions.

9. The virtual operation interaction method based on an intelligent display device according to claim 1, characterized in that, Calculate the correspondence between the first hand feature and the first eye feature; if the correspondence matches a preset correspondence set, the steps to enter the simple menu include the following: The first hand features include finger position, gesture type, and palm orientation; The first eye feature includes eye position and gaze direction; Based on the finger position and the eye position, determine whether the finger position and the eye position are both within the corresponding position range; If both are true, continue to judge the first hand feature and the first eye feature; Determine whether the gesture type is a preset first gesture or a second gesture, and determine whether the angle between the palm orientation and the line of sight is within the corresponding preset angle range. If both are true, then the correspondence is determined to conform to the preset correspondence set.

10. The virtual operation interaction method based on an intelligent display device according to claim 9, characterized in that, The first gesture includes palm facing upwards, and the second gesture includes palm facing the eyeball.

11. The virtual operation interaction method based on an intelligent display device according to claim 1, characterized in that, The gesture data is identified based on the gesture image; The steps for matching the first hand feature based on the gesture data include: Feature information is extracted from the gesture image using image processing techniques, and gesture data is matched using the feature information; Alternatively, feature information can be extracted from the gesture image, and mirrored to obtain mirror information; gesture data can then be matched using the mirror information. The gesture data is input into a pre-established database of gesture templates. The features corresponding to the gesture data are compared one by one with the gesture features in the database, and all similarity values ​​are calculated. The hand gesture image with the highest similarity value is selected as the first hand feature.

12. The virtual operation interaction method based on an intelligent display device according to claim 1, characterized in that, The method also includes: In the step of acquiring gesture images and eye images in real time, multiple gesture images and eye images are acquired within a set time period; The correspondence is calculated multiple times. If the correspondence matches the preset correspondence set, the simple menu is entered; otherwise, it is judged as a mis-touch. Alternatively, when a hardware trigger signal is received, the user can enter a simple menu.

13. A virtual operation interaction system based on an intelligent display device, characterized in that, Includes a processor, wherein the processor performs the steps of the virtual operation interaction method based on a smart display device as described in any one of claims 1-12.

Citation Information

Cited By

  • Intelligent watch interaction method and system based on gesture action

    CN121979397A