Shape-based graphical indication of interaction events
By explaining the user's virtual element interaction data in the three-dimensional space in the head-mounted device, providing graphic indications that match the shape of the user interface elements, the problem of inconspicuous user interface interaction is solved and interactive visibility and privacy protection are improved.
Patent Information
- Application Number
- CN202510069302.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-01-14
- Filing Date
- 2025-01-15
- Publication Date
- 2025-07-22
AI Technical Summary
When existing systems use head-mounted devices, it is difficult for existing systems to fully display the movement and interaction between the user and the user interface icon, resulting in the inability to obvious enough interaction.
By receiving and interpreting user's virtual element interaction data in three-dimensional space, providing graphical indications to match the determined shape of user interface elements, using sensor data to determine user attention directions, and providing visual effects such as luminescence or highlighting based on the type and temporal aspects of user interface elements.
Improve the visibility and accuracy of user interface elements interaction, reduce violations of user privacy, and simplify the application input recognition process.
Smart Images

Figure CN120353362A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to systems, methods, and devices that enable the evaluation of user interactions to control graphical indications of user interface elements regarding an electronic device. Background Art
[0002] When a user is using a device such as a head-mounted device (HMD), it may be desirable to detect movements and interactions associated with icons of a user interface. However, existing systems may not provide a sufficient display of the interactions when the user is navigating to or selecting an icon of the user interface associated with the user's attention. Summary of the Invention
[0003] Various embodiments disclosed herein include devices, systems, and methods for interpreting user activities when a user interacts with virtual elements (e.g., user interface elements) located within a three-dimensional (3D) space such as an extended reality (XR) environment. Some embodiments utilize an architecture that receives application user interface geometries in a receiving system or a shared simulation area and outputs data (e.g., less than all user activity data) for an application to use in identifying input. An operating system (OS) process may be configured to provide an input support process to support the identification of input intended for one or more separately-executing applications, e.g., by providing some input identification tasks to identify user activities as input for the application, or by converting user activity data into a format that can be more easily, accurately, efficiently, or effectively interpreted by the application and / or in a manner that is conducive to protecting user privacy. The OS process may include a simulation process that utilizes application user interface information to provide 3D information (e.g., 3D world data), and the input support process uses the 3D information to support the identification of input intended for one or more separate (e.g., separately-executing) applications.
[0004] Various embodiments disclosed herein graphically indicate a determined shape of a user interface element on a two-dimensional (2D) web page viewed in an XR environment using a 3D display device (e.g., a wearable device such as a head-mounted device (HMD)) based on a user's intent (attention) for the user interface element. The goal is to match an arbitrary shape of the user interface element with a visual effect that provides illumination or highlighting (as contrasted with a standard visual effect for all user interface elements such as highlighting of a square shape).
[0005] Various embodiments disclosed herein may provide a method of matching a graphical indication with a determined shape of a user interface element rather than using a predefined shape (e.g., a standard rectangular shape around a circular element). In other words, if the user interface element is a star shape, the illumination or highlighting around the user interface element will also display a star shape.
[0006] In some specific implementations, the user's intention can be based on determining the user's attention direction towards a user interface element according to sensor data. For example, the data based on the user can include gaze data, hand data, head data, etc., or other input data (e.g., data from an input controller). The visual effect can be based on the type of user interface element (e.g., photo gallery), time aspect (delay until glowing, or brighter based on gaze length), the identified sub-elements, the type of interaction (gaze, pinch, combination of gaze and pinch, etc.), and so on. In some specific implementations, identifying the shape of the user interface element can be based on image mapping, metadata, image recognition, etc.
[0007] In some specific implementations, providing a graphical indication (e.g., a visual effect such as glowing, highlighting, etc.) corresponding to the determined shape of the user interface element can include matching the boundary / shape of the user interface element, rather than using a predefined shape (e.g., a rectangle around a circular affordance). In some specific implementations, providing the graphical indication can remove the transparent area associated with the user interface element. In some specific implementations, configuring the graphical indication can be based on determining the type of the user interface element. For example, for a photo icon, the shape determination technique for determining the type of the user interface element can use the attributes of the photo icon (such as size, entropy, and / or resolution) to configure the visual effect based on a confidence threshold. In some specific implementations, there can be size constraints. For example, if the user interface element is too large, the visual effect may not be displayed. In some specific implementations, the color of the user interface element can match the glow of the visual effect (e.g., a red glow associated with a user interface element that is normally red).
[0008] In some specific implementations, the method described herein for providing a graphical indication is based on time aspects. For example, the graphical indication (e.g., visual effect) can provide a graphical effect (e.g., glowing) a few milliseconds after the user's attention (e.g., gaze and / or pointing at a specific icon). Additionally or alternatively, in some specific implementations, the graphical indication (e.g., a visual effect such as glowing or highlighting) can become brighter based on the user's gaze on the element for a period of time. Additionally or alternatively, in some specific implementations, the graphical indication can provide different visual effects based on the shape of the user interface element (e.g., a circular shape with a gradient glowing centered at the position where the user is gazing; but if it is a star shape, there can be a matching star-shaped glow).
[0009] User privacy can be preserved by providing some user activity information only to separately executed applications. For example, user activity information that is not associated with an intended user action (such as a user action where the user intends to provide input or certain types of input) can be withheld. In one example, raw hand and / or gaze data can be excluded from the data provided to the application, such that when there is no intended user interface interaction, the application receives limited or no information about where the user is looking or what the user is looking at.
[0010] Generally, an innovative aspect of the subject matter described in this specification can be embodied in a method that includes the following actions: presenting a view of a three-dimensional (3D) environment at a device having a processor, where one or more user interface elements are positioned at 3D locations based on a 3D coordinate system associated with the 3D environment. These actions can also include determining the shape of the one or more user interface elements. These actions can also include receiving data corresponding to user activity in the 3D coordinate system. These actions can also include identifying a user interaction event associated with a first user interface element in the 3D environment based on the data corresponding to the user activity. These actions can also include providing a graphical indication corresponding to the determined shape of the first user interface element according to the identified user interaction event.
[0011] These implementations and other implementations can optionally include one or more of the following features.
[0012] In some aspects, determining the shape based on identifying the user interface element includes that the user interface element includes an interactive item. In some aspects, determining the shape is based on formatting information associated with the one or more user interface elements. In some aspects, determining the shape is based on an image map associated with the one or more user interface elements.
[0013] In some aspects, providing the graphical indication corresponding to the determined shape of the first user interface element includes matching the shape of the graphical indication to the shape of the first user interface element.
[0014] In some aspects, determining the shape of the first user interface element includes: identifying sub-elements of the first user interface element; and determining the shape of the first user interface element based on the identified sub-elements.
[0015] In some aspects, the graphical indication corresponds to the identified sub-elements of the first user interface element. In some aspects, providing the graphical indication includes removing a portion of the view of the first user interface element within the view of the 3D environment. In some aspects, the graphical indication is a highlighting effect corresponding to the user interface element. In some aspects, the graphical indication is a glowing effect corresponding to the user interface element. In some aspects, the graphical indication is based on determining the type of the user interface element.
[0016] In some aspects, the graphical indication is displayed for a first instance based on one or more first attributes, and wherein the graphical indication is displayed for a second instance different from the first instance based on one or more second attributes different from the first attributes.
[0017] In some aspects, the data corresponding to the user activity is obtained via one or more sensors on the device. In some aspects, the data corresponding to the user activity includes gaze data, the gaze data including a stream of gaze vectors corresponding to the gaze direction over time during use of the electronic device. In some aspects, the data corresponding to the user activity includes hand data, the hand data including the hand pose skeleton of multiple joints at each of multiple moments during use of the electronic device. In some aspects, the data corresponding to the user activity includes hand data and gaze data. In some aspects, the data corresponding to the user activity includes controller data and gaze data. In some aspects, the data corresponding to the user activity includes the head pose data of the user.
[0018] In some aspects, the electronic device includes a head-mounted device (HMD).
[0019] According to some embodiments, a device includes one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing to perform any of the methods described herein. According to some embodiments, instructions are stored in a non-transitory computer-readable storage medium, the instructions causing the device to perform or cause to perform any of the methods described herein when executed by one or more processors of the device. According to some embodiments, a device includes: one or more processors, non-transitory memory, and components for performing or causing to perform any of the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Thus, the present disclosure can be understood by those of ordinary skill in the art, and a more detailed description can be referred to in aspects of some illustrative embodiments, some of which are shown in the drawings.
[0021] Figures 1A to 1B Exemplary electronic devices operating in a physical environment according to some embodiments are illustrated.
[0022] Figure 2 Exemplified is a view provided by a device of user interface elements within a 3D physical environment in which a user performs an interaction according to some embodiments Figures 1A to 1B thereof.
[0023] Figure 3 Illustrates a view provided by a device of user interface elements within a 3D physical environment in which a user performs interactions according to some specific implementations. Figures 1A to 1B
[0024] Figure 4 Illustrates an example of tracking the movement of a hand and gaze during an interaction according to some specific implementations.
[0025] Figure 5A and Figure 5B Illustrates a view of an example of interaction recognition of user activities and displaying a graphical indication based on a determined shape of a user interface element according to some specific implementations.
[0026] Figure 6A and Figure 6B Illustrates a view of an example of interaction recognition of user activities and displaying a graphical indication based on a determined shape of a user interface element according to some specific implementations.
[0027] Figures 7A to 7D Illustrates a view of an example of interaction recognition of user activities and displaying a graphical indication based on a determined shape of a user interface element according to some specific implementations.
[0028] Figure 8A and Figure 8B Illustrates a view of an example of interaction recognition of user activities and displaying a graphical indication based on a determined shape of a user interface element associated with a photographic element according to some specific implementations.
[0029] Figure 9A and Figure 9B Illustrates a view of an example of interaction recognition of user activities and displaying a graphical indication based on a determined shape of a user interface element associated with a photographic element according to some specific implementations.
[0030] Figure 10 Illustrates generating interaction data based on hand data, gaze data, and user interface target data using an exemplary input support framework according to some specific implementations.
[0031] Figure 11 Illustrates an example of interaction recognition of user activities and displaying a graphical indication based on a determined shape of a user interface element according to some specific implementations.
[0032] Figure 12 Is a flowchart illustrating a method for providing a graphical indication corresponding to a determined shape of a user interface element based on identifying user interaction events corresponding to user activities according to some specific implementations.
[0033] Figure 13is a block diagram of an electronic device according to some specific implementations.
[0034] Figure 14 is a block diagram of an exemplary head-mounted device according to some specific implementations.
[0035] According to common practice, the various features illustrated in the figures may not be drawn to scale. Thus, for clarity, the dimensions of the various feature portions may be arbitrarily expanded or reduced. Additionally, some of the figures may not depict all of the components of a given system, method, or device. Finally, throughout the specification and the figures, like reference numerals may be used to denote like feature portions. Detailed Implementations
[0036] Numerous details are described to provide a thorough understanding of the example implementations shown in the figures. However, the figures merely illustrate some example aspects of the present disclosure and should not be considered limiting. Those of ordinary skill in the art will understand that other effective aspects and / or variations do not include all of the specific details described herein. Additionally, well-known systems, methods, components, devices, and circuits have not been described in detail so as not to obscure more relevant aspects of the example implementations described herein.
[0037] Figures 1A to 1B Illustrates exemplary electronic devices 105 and 110 operating in a physical environment 100. In Figures 1A to 1B the example, the physical environment 100 is a room that includes a table 120. The electronic devices 105 and 110 may include one or more cameras, microphones, depth sensors, or other sensors that can be used to capture information about and evaluate the physical environment 100 and the objects therein, as well as information about the user 102 of the electronic devices 105 and 110. Information about the physical environment 100 and / or the user 102 can be used to provide visual and audio content, and / or to identify the current location of the physical environment 100 and / or the location of the user within the physical environment 100.
[0038] In some specific implementations, views of an extended reality (XR) environment may be provided to one or more participants (e.g., user 102 and / or other participants not shown) via electronic device 105 (e.g., a wearable device such as an HMD, etc.) and / or 110 (e.g., a handheld device such as a mobile device, a tablet computing device, a laptop computer, etc.). Such an XR environment may include a view of a 3D environment seen through a transparent or translucent display or a 3D environment generated based on camera images and / or depth camera images of the physical environment 100, as well as a representation of user 102 based on camera images and / or depth camera images of user 102. Such an XR environment may include virtual content positioned at a 3D position relative to a 3D coordinate system (e.g., 3D space) associated with the XR environment, and this 3D coordinate system may correspond to the 3D coordinate system of the physical environment 100.
[0039] In some specific implementations, video (e.g., a see-through video depicting the physical environment) is received from an image sensor of a device (e.g., device 105 or device 110). In some specific implementations, the 3D representation of the virtual environment is aligned with the 3D coordinate system of the physical environment. The size of the 3D representation of the virtual environment may be generated especially based on the scale of the physical environment or the positioning of open spaces, floors, walls, etc., such that the 3D representation is configured to be aligned with corresponding features of the physical environment. In some specific implementations, a viewpoint within the 3D coordinate system may be determined based on the position of the electronic device within the physical environment. This viewpoint may be determined especially based on image data, depth sensor data, motion sensor data, etc., which may be retrieved via a virtual inertial odometry (VIO) system, a simultaneous localization and mapping (SLAM) system, etc.
[0040] Figure 2 Illustrates a view provided by a device of user interface elements within a Figures 1A to 1B 3D physical environment in which the user performs an interaction (e.g., a direct interaction). In this example, user 102 makes a gesture with respect to the content presented in views 210a to 210b of the XR environment provided by the device (e.g., device 105 or device 110). Views 210a to 210b of the XR environment include an exemplary user interface 230 of an application (e.g., an example of virtual content) and a representation 220 of a table 120 (e.g., an example of real content). Providing such a view may involve determining 3D attributes of the physical environment 100 and positioning the virtual content, such as user interface 230, in the 3D coordinate system corresponding to that physical environment 100.
[0041] In Figure 2In the example, user interface 230 includes various content user interface elements, including background portion 235 and user interface elements 242, 243, 244, 245, 246, 247. User interface elements 242, 243, 244, 245, 246, 247 can be displayed on a flat two-dimensional (2D) user interface 230. User interface 230 can be a user interface of an application, as illustrated in this example. In some specific implementations, an indication identifier (e.g., a pointer, a highlighting structure, etc.) can be used to indicate an interaction point with any user interface (visual) element (e.g., if using a controller device such as a mouse or other input device). For illustrative purposes, user interface 230 is simplified, and the user interface can actually include any level of complexity, any number of content items, and / or a combination of 2D and / or 3D content. User interface 230 can be provided by various types of operating systems and / or applications, including but not limited to messaging applications, web browser applications, content viewing applications, content creation and editing applications, or any other application that can display, present, or otherwise use visual and / or audio content.
[0042] In this example, the background portion 235 of user interface 230 is flat. In this example, background portion 235 includes all aspects of user interface 230 that are being displayed other than user interface elements 242, 243, 244, 245, 246, 247. Displaying the background portion of the user interface of an operating system or application as a flat surface can provide various advantages. Doing so can provide a user interface of an application that is easy to understand or otherwise easy to access as part of an XR environment. In some specific implementations, multiple user interfaces (e.g., corresponding to multiple different applications) are presented sequentially and / or simultaneously within an XR environment (e.g., within one or more colliders or other such components).
[0043] In some specific implementations, the position and / or orientation of such one or more user interfaces can be determined to facilitate visibility and / or use. In a 3D environment, one or more user interfaces can be in a fixed position and orientation. In such a case, user movement will not affect the position or orientation of the user interface within the 3D environment.
[0044] The position of the user interface within the 3D environment can be based on determining the distance of the user interface from the user (e.g., from the initial or current user position). The position and / or the distance from the user can be determined based on various criteria, including but not limited to criteria that consider application type, application function, content type, content / text size, environment type, environment size, environment complexity, environment lighting, the presence of other people in the environment, the use of an application or content by multiple users, user preferences, user input, and numerous other factors.
[0045] In some specific implementations, one or more user interfaces can be body-locked content, for example, having a certain distance and orientation offset relative to a part of the user's body (e.g., their torso). For example, the body-locked content of the user interface can be 0.5 meters to the left of and at a 45-degree angle to the forward-facing vector of the user's torso. If the user's head rotates while the torso remains stationary, in a 3D environment, the body-locked user interface will appear to remain stationary 2m to the left of and at a 45-degree angle to the forward-facing vector of the torso. However, if the user does rotate their torso (e.g., by rotating in their chair), the body-locked user interface will follow the torso rotation and reposition itself within the 3D environment such that it remains 0.5m to the left of and at a 45-degree angle to the new forward-facing vector of their torso.
[0046] In other specific implementations, the user interface content is defined at a specific distance from the user, where the orientation relative to the user remains stationary (e.g., if initially displayed in a basic direction, it will remain in that basic direction regardless of any head or body movement). In this example, the orientation of the body-locked content is not referenced to any part of the user's body. In this different specific implementation, the body-locked user interface will not reposition itself based on torso rotation. For example, the body-locked user interface can be defined as 2m away, and based on the direction the user is currently facing, it can initially be displayed to the north of the user. If the user rotates their torso 180 degrees to face south, the body-locked user interface will remain 2m to the north of the user, which is now directly behind the user.
[0047] The body-locked user interface can also be configured to always remain aligned with the line of gravity or the horizontal line, such that head and / or body changes in the roll orientation will not cause the body-locked user interface to move within the 3D environment. Translational movement will cause the body-locked content to reposition itself within the 3D environment to maintain the distance offset.
[0048] In Figure 2 the example, user 102 moves their hand from an initial position, as illustrated by the position of reference numeral 222 in view 210a. The hand moves along path 250 to a later position, as illustrated by the position of reference numeral 222 in view 210b. As user 102 moves their hand along this path 250, the finger intersects the user interface 230. Specifically, as the finger moves along path 250, it virtually pierces the user interface element 245 and thus the tip portion of the finger (not shown) is occluded by the user interface 230 in view 210b.
[0049] The specific implementations disclosed herein interpret user movements, such as user 102 moving their hand / finger along a path 250 relative to a user interface element, such as user interface element 245, to recognize user input / interaction. The interpretation of user movements and other user activities can be based on using one or more recognition processes to identify user intent.
[0050] In Figure 2 an example, recognizing the input can involve determining that the gesture is a direct interaction and then using a direct input recognition process to recognize the gesture. For example, such a gesture can be interpreted as a tap input to user interface element 245. When making such a gesture, the actual movement of the user relative to icon 245 can deviate from the ideal movement (e.g., a straight-line path passing through the center of the user interface element in a direction exactly orthogonal to the plane of the user interface element). The actual path can be curved, jagged, or otherwise non-linear and can be angled rather than orthogonal to the plane of the user interface element. The path can have properties that make it similar to other types of input gestures (e.g., swipe, drag, flick, etc.). For example, a non-orthogonal movement can make the gesture similar to a swipe movement, where the user provides input by piercing the user interface element and then moving in a direction along the plane of the user interface.
[0051] Some specific implementations disclosed herein determine that a direct interaction mode is applicable and, based on that direct interaction mode, utilize a direct interaction recognition process to distinguish or otherwise interpret user activities corresponding to direct input, e.g., to recognize the intended user interaction based on whether and how the gesture path intercepts one or more 3D regions of space. Such a recognition process can consider actual human tendencies associated with direct interaction (e.g., natural arching during an intended straight movement, tendency to move based on shoulder or other pivot positions, etc.), human perception issues (e.g., the user cannot see or does not know the exact position of virtual content relative to their hand), and / or other direct interaction-specific issues.
[0052] Note that the movement of the user in the real world (e.g., physical environment 100) corresponds to movement within a 3D space (e.g., an XR environment based on the real world and including virtual content such as a user interface positioned relative to real-world objects including the user). Thus, the user moves their hand in physical environment 100, e.g., through empty space, but the hand (e.g., a drawing or representation of the hand) intersects and / or pierces the user interface of the XR environment based on that physical environment. In this way, the user directly interacts with virtual content.
[0053] Figure 3 illustrates an Figures 1A to 1BExemplary views provided by a device of user interface elements within a 3D physical environment. In this example, user 102 makes a gesture while viewing content presented in view 302 of an XR environment provided by a device (e.g., device 105 or device 110). View 302 of the XR environment includes Figure 2 exemplary user interface 230. In Figure 3 the example, user 102 makes a pointing gesture as illustrated by representation 222 with their hand while gazing in the gaze direction 310 at user interface icon 246 (e.g., a star-shaped application icon or widget). In this example, this user activity (e.g., making a pointing gesture while gazing at a user interface element) corresponds to a user intention to interact with user interface icon 246. For example, the pointing indicates a potential interaction intention, and the gaze (at the time of the pointing) identifies the target of the interaction (e.g., waiting for the system to highlight the icon to indicate the correct target to the user and then initiating the interaction from another user activity such as via a pinch gesture).
[0054] The specific implementations disclosed herein interpret user activities (such as user 102 making a pointing gesture while gazing at a user interface element) to identify the user / interaction. For example, such user activity can be interpreted as a tap input to user interface element 246, e.g., selecting user interface element 246. However, when performing such an action, the user's gaze direction and / or the timing between the gesture and the gaze that the user intends to associate with the gesture cannot be perfectly executed and / or timed.
[0055] Some of the specific implementations disclosed herein determine that an indirect interaction mode is applicable and, based on that indirect interaction mode, utilize an indirect interaction recognition process to identify an intended user interaction based on user activities (e.g., based on whether and how a gesture path intercepts one or more 3D regions of space). Such a recognition process can consider actual human tendencies associated with indirect interaction (e.g., eye saccades, eye fixations, and other natural human gazing behaviors, hand arching movements, retractions that do not correspond to the intended insertion direction, etc.), human perception issues (e.g., the user cannot see or does not know the exact position of virtual content relative to their hand), and / or other indirect interaction-specific issues.
[0056] Some specific implementations determine an interaction mode, e.g., a direct interaction mode or an indirect interaction mode, such that user behavior can be interpreted by a specialized (or otherwise separate) recognition process for the appropriate interaction type, e.g., using a direct interaction recognition process for direct interaction and an indirect interaction recognition process for indirect interaction. Such a specialized (or otherwise separate) process can be more efficient, more accurate, or provide other benefits compared to using a single recognition process configured to identify multiple types (e.g., both direct and indirect) of interactions.
[0057] Figure 2 and Figure 3 illustrates example interaction patterns based on user activities in a 3D environment. Additionally or alternatively, other types or modes of interaction may be used, including but not limited to user activities via input devices such as keyboards, touchpads, mice, handheld controllers, etc. In one example, a user uses an input device such as a keyboard, touchpad, mouse, or handheld controller to provide an interaction intent via an activity (e.g., performing an action such as tapping a button or the surface of a touchpad), and identifies a user interface target based on the user's gaze direction when input is made on the input device. Similarly, user activities may involve voice commands. In one example, a user uses an input device such as a keyboard, touchpad, mouse, or handheld controller to provide an interaction intent via an activity (e.g., performing an action such as tapping a button or the surface of a touchpad), and identifies a user interface target based on the user's gaze direction when a voice command is made. In another example, user activities identify the intent of an interaction (e.g., via pinching, gestures, voice commands, input device input, etc.), and determine user interface elements based on a non-gaze-based direction (e.g., based on where the user is pointing within the 3D environment). For example, a user may pinch with one hand to provide input indicating an interaction intent, while pointing a finger of the other hand at a user interface button. As another example, a user may manipulate the orientation of a handheld device within the 3D environment to control a controller direction (e.g., a virtual line extending from a controller within the 3D environment), and may identify the user interface element that the user is interacting with based on that controller direction (e.g., based on identifying what user interface element the controller direction intersects when an input indicating an interaction intent is received).
[0058] The various specific implementations disclosed herein provide an input support process, such as an OS process separate from the execution of an application, that processes user activity data (e.g., regarding gaze, gestures, other 3D activities, HID input, etc.) to produce data for an application that the application can interpret as user input. The application may not need to have 3D input recognition capabilities because the data provided to the application can be in a format that the application can recognize using 2D input recognition capabilities, such as those used within applications developed for use on 2D touchscreens and / or 2D cursor-based platforms. Thus, at least some aspects of interpreting user activities for an application can be performed by a process external to the application. Doing so can simplify or reduce the complexity, requirements, etc. of the application's own input recognition process, ensure unified and consistent input recognition across multiple different applications, protect private usage data from application access, and numerous other benefits as described herein.
[0059] Figure 4Exemplary interactions are illustrated for tracking the movement of the two hands 422, 424 of the user 102 and the gaze along the path 410 when the user 102 virtually interacts with the user interface element 415 of the user interface 400. Specifically, Figure 4 Interactions with the user interface 400 are illustrated when the user faces the user interface 400. In this example, the user 102 is using the device 105 to view and interact with the XR environment including the user interface 400. The interaction recognition process (e.g., direct or indirect interaction) can use sensor data and / or user interface information to determine, for example, which user interface element the user's hand is virtually touching, which user interface element the user intends to interact with, and / or where on the user interface element the interaction occurs. Direct interaction can additionally (or alternatively) involve evaluating user activity to determine the user's intent, e.g., whether the user intends to perform a direct tap gesture through the user interface element or a slide / scroll movement along the user interface element. Additionally, the recognition of the user's intent can utilize information about the user interface element. For example, determining the user's intent regarding a user interface element can include the position, size, and type of the element, the types of interactions that can be performed on the element, the types of interactions enabled on the element, which of the set of potential target elements for the user activity accepts which types of interactions, etc.
[0060] Various two - hand gestures can be enabled based on interpreting hand positions and / or movements using sensor data (e.g., images or other sensor data captured by outward - facing sensors on an HMD such as the device 105). For example, a pan gesture can be performed by pinching the hands together and then moving the hands in the same direction, e.g., holding the hands outward at a fixed distance apart from each other and moving both of them an equal amount to the right to provide an input for panning to the right. In another example, a zoom gesture can be performed by holding the hands outward and moving one hand or both hands to change the distance between the hands (e.g., moving the hands closer to each other to zoom in and farther apart from each other to zoom out).
[0061] Additionally or alternatively, in some implementations, the identification of such interaction of both hands may be based on functions performed both via system processes and via application processes. For example, an input support process of the OS may interpret hand data from a sensor of the device to identify an interaction event and provide limited or interpreted information about the interaction event to an application providing user interface 400. For example, rather than providing detailed hand information (e.g., identifying the 3D positions of multiple joints of a hand model representing the configuration of hand 422 and hand 424), the OS input support process may simply identify the 2D point on user interface element 415 within 2D user interface 400 where the interaction (e.g., interaction pose) occurred. The application process may then interpret the 2D point information (e.g., interpreting the information as a selection, mouse click, touch screen tap, or other input received at that point) and provide a response, e.g., modifying its user interface accordingly.
[0062] In some implementations, hand motion / position may be tracked using a varying shoulder-based pivot position that is assumed to be at a location based on a fixed offset from the current location of the device 105. The fixed offset may be determined using an expected fixed spatial relationship between the device and the pivot point / shoulder. For example, given the current location of the device 105, the shoulder / pivot point may be determined to be at location X given the fixed offset. This may involve updating the shoulder position over time (e.g., every frame) based on changes in the location of the device over time. The fixed offset may be determined as a fixed distance between a determined location at the top of the center of the user's 102 head and the shoulder joint.
[0063] Figure 5A , Figure 5B , Figure 6A , Figure 6B , Figures 7A to 7D , Figure 8A , Figure 8B , Figure 9A and Figure 9B Different examples of tracking user activity (e.g., movement of a hand, gaze, etc.) during an interaction in which the user is attempting to perform a gesture (e.g., the user's intent (attention) to a user interface element) in order to provide a graphical indication (e.g., a visual effect) to highlight a user interface element that matches the shape of the user interface element that the user is focusing on are illustrated. For example, each figure illustrates identifying the location of an object (e.g., a user interface element) based on tracking a portion of the user (e.g., the user's hand) using a sensor (e.g., an outward-facing image sensor) on a head-mounted device (e.g., device 105) as the user moves through and interacts with the environment (e.g., an XR environment). For example, the user may be looking at an XR environment such as Figure 2 The XR environment 205 illustrated in Figure 3the XR environment 305 illustrated therein and interact with elements within the application window of the user interface (e.g., user interface 230) while the device (e.g., device 105) tracks the hand movements and / or gaze of the user 102. Then, the user activity tracking system can determine whether the user is attempting to interact with a specific user interface element. A user interface object can be virtual content that the application window allows the user to interact with, and the user activity tracking system can determine whether the user is interacting with any specific element or performing a specific movement in the 3D coordinate space, such as performing a pinch gesture. For example, when the user views the user interface or object at a first instance in time and performs a user interaction event, hand representation 222 represents the left hand of user 102, and hand representation 224 represents the right hand of user 102. When the user activity indicates an interaction with a specific object (e.g., user interface element) at a second instance in time, the application can initiate the identified action (e.g., provide a graphical indication for the determined expected user interface element). In some embodiments, when the user moves his or her hand (e.g., hand representations 222, 224), the user activity tracking system can track the movement of the hand based on the movement of one or more points, and the application can perform actions (e.g., zoom, rotate, move, pan, etc.) based on the detected movement of both hands.
[0064] Figure 5A and Figure 5B illustrates an example of the interaction recognition of user activity of (e.g., Figures 1A to 1B user 102) with a user interface element and the display of a graphical indication based on the determined shape of the user interface element. Figure 5A and Figure 5B are presented in view 510A and view 510B respectively of the XR environment provided by Figures 1A to 1B the electronic device 105 and / or the electronic device 110. Views 510A to 510B of the XR environment 505 include views of representation 220 when the user 102 interacts with the user interface element 246 (e.g., the icon of an application in the shape of a star) of the user interface 230. Specifically, Figure 5A view 510A illustrates the intention (attention) of user 102 at a first instance in time, which is directed at the user interface element 246, as illustrated by the left finger pointing (e.g., hand representation 222) and the gaze along path 502. Figure 5B view 510B illustrates the graphical indication 520 at a second instance in time, which matches the shape of the user interface element 246 and is based on according to Figure 5AThe intent of user 102 determined by the user activities exemplified therein (e.g., the intent or attention to user interface element 246). The graphical indication 520 provides a visual effect or glow that matches the shape of user interface element 246, rather than displaying a highlight of a general or standard shape (such as a square). Thus, the highlight of the star shape (graphical indication 520) presents a glow or some other type of visual effect behind the star-shaped user interface element 246 to indicate to user 102 that the system has recognized the user intent to interact with the application associated with user interface element 246.
[0065] Figure 6A and Figure 6B exemplifies the interaction recognition of user activities of user 102 (e.g., Figures 1A to 1B of user 102) with user interface elements and the display of graphical indications based on the determined shape of the user interface elements. Figure 6A and Figure 6B present XR environments provided by Figures 1A to 1B electronic device 105 and / or electronic device 110 in views 610A and 610B respectively. Views 610A to 610B of XR environment 605 include views of representation 220 when user 102 interacts with user interface element 245 (e.g., an icon of an application in the shape of a building) of user interface 230. Specifically, Figure 6A View 610A exemplifies the intent (attention) of user 102 at a first time instance, which intent (attention) is directed to user interface element 246, as exemplified by the left finger pointing (e.g., hand representation 222) and the gaze along path 602. Figure 6B View 610B exemplifies graphical indication 620 at a second time instance, which graphical indication matches the shape of user interface element 245 and is based on the intent of user 102 determined by the user activities exemplified in Figure 6A therein (e.g., the intent or attention to user interface element 245). The graphical indication 620 provides a visual effect or glow that matches the shape of user interface element 245, rather than displaying a highlight of a general or standard shape (such as a square). In other words, the contour shape of the object of user interface element 245 (building) is matched, and the graphical indication 620 is extended and presented to look like a separate 2D window or layer, which separate 2D window or layer looks larger and is behind user interface element 245. Thus, the highlight of the building shape (graphical indication 620) presents a glow or some other type of visual effect behind user interface element 246 to indicate to user 102 that the system has recognized the user intent to interact with the application associated with user interface element 245.
[0066] Figures 7A to 7DIllustrates an example of interaction recognition of user activities of a user 102 (e.g., Figures 1A to 1B with user interface elements and display of a graphical indication based on a determined shape of the user interface element. Figures 7A to 7D Views 710A to 710D respectively present an XR environment provided by an Figures 1A to 1B electronic device 105 and / or an electronic device 110. Views 710A to 710D of the XR environment 705 include views of a representation 220 when the user 102 interacts with a user interface element 247 (e.g., an icon of an application in a concentric circle shape) of the user interface 230.
[0067] Specifically, Figure 7A View 710A illustrates the intention (attention) of the user 102 at a first time instance, which is directed at the user interface element 247, as illustrated by a gaze along a path 702 towards the user interface element 247 (e.g., the user 102 first looks at the icon). Figure 7B View 710B illustrates a second time instance, as illustrated by a left finger pointing (e.g., hand representation 222) combined with a gaze along the path 702 (e.g., the user 102 continues to focus and now points at the icon). Figure 7B Also illustrated is a graphical indication 720 that matches the outer circular shape of the user interface element 247 and is based on the intention of the user 102 determined according to the Figure 7A user activities illustrated in (e.g., gazing at the user interface element 245) and according to the Figure 7B user activities illustrated in (e.g., gazing and finger pointing at the user interface element 245). The graphical indication 720 provides a visual effect or glow that matches the shape of the user interface element 247. In other words, the contour shape (outer circle of the concentric circles) of the object of the user interface element 247 is matched, and the graphical indication 720 is expanded and presented to look like a separate 2D window or layer that looks larger and is behind the user interface element 247. Therefore, the highlighted circular shape (graphical indication 720) presents a glow or some other type of visual effect behind the user interface element 247 to indicate to the user 102 that the system has recognized the user's intention to interact with the application associated with the user interface element 247.
[0068] Figure 7C View 710C illustrates a third time instance, as illustrated by a left hand pinch (e.g., hand representation 222) combined with a gaze along the path 702. For example, the user 102 continues to focus and now performs an action (such as a pinch) on the icon, which may trigger an action by the application associated with the user interface element 247, as Figure 7D shown. Figure 7CA graphical indication 722 is also illustrated, which matches the outer circular shape of the user interface element 247, but is larger than Figure 7B the graphical indication 720. In some embodiments, the appearance or other properties of the graphical indication may change over time (e.g., grow larger, blink, fluctuate in size, change color, etc.). Thus, as Figure 7B illustrated between the second time instance in Figure 7C and the third time instance in Figure 7D (e.g., from a few milliseconds up to one or two seconds or longer), the size of the first graphical indication 720 increases to the second graphical indication 722. In other words, when the user's intention remains focused on the same item (e.g., the user interface element 247) over a period of time, the visual effect will reflect this change to indicate to the user which element the system has determined the user intention to interact with. Figure 7C The view 710D at the fourth time instance is illustrated, such as a left hand pinch (e.g., hand representation 222) and a right hand pinch (e.g., hand representation 224) and interaction with the application 730 associated with the user interface element 247. In other words,
[0069] Figure 8A and Figure 8B illustrate an example of interaction recognition of user activities with a user interface element associated with a photographic element and display of a graphical indication based on the determined shape of the photographic element, according to some embodiments (e.g., Figures 1A to 1B the user 102). Figure 8A and Figure 8B present XR environments provided by Figures 1A to 1B the electronic device 105 and / or the electronic device 110 in view 810A and view 810B, respectively. The views 810A to 810B of the XR environment 805 include a representation 220 of a table 120 and a view of a user interface 830 (e.g., a photo gallery application), which includes a series of digital images, e.g., user interface elements 812, 814, 816, 818, each digital image having a different quality (e.g., resolution, entropy, size, data size, etc.). Additionally, the user 102 is interacting (e.g., viewing and pointing) with a specific photo, the user interface element 816 (e.g., a higher resolution digital photographic image of someone) of the user interface 830. Specifically, Figure 8AView 810A illustrates the intent (attention) of user 102 at a first instance in time, which is directed at user interface element 816, as illustrated by the left finger pointing (e.g., hand gesture 222) and the gaze along path 802. Figure 8B View 810B illustrates graphical indicators 820 and 822 at a second instance in time, which are based on matching the shape of user interface element 816 and on the intent of user 102 determined in response to Figure 8A the user activity illustrated in. For example, graphical indicator 820 provides an initial visual effect or glow around the person's face, and graphical indicator 822 provides a visual effect or glow that applies one or more photographic visual effects that match the shape of the person's face. Determining to provide a photographic visual effect (e.g., graphical indicator 822) can be based on determining that the element is a photo and that user interface element 816 meets or exceeds one or more confidence thresholds associated with the photo (e.g., resolution threshold, entropy threshold, data size threshold, size threshold relative to the current view, etc.). In other words, based on a higher quality image, a higher quality visual effect is applied (e.g., graphical indicator 822 for the person's face). Additionally, an oval-shaped highlighted photo indicator (graphical indicator 820) presents a glow around the person's head image or some other type of visual effect behind user interface element 816 to indicate to user 102 that the system has recognized the user intent to interact with the application associated with user interface element 816 (e.g., increase the size of the image, move to another location within environment 905, etc.).
[0070] Figure 9A and Figure 9B illustrates an example of the interaction recognition of the user activity of user 102 (e.g., Figures 1A to 1B of) with a user interface element associated with a photographic element and the display of graphical indicators based on the determined shape of the photographic element. Figure 9A and Figure 9B are presented in view 910A and view 910B, respectively, by Figures 1A to 1BThe XR environment provided by the electronic device 105 and / or the electronic device 110. The views 910A to 910B of the XR environment 905 include a representation 220 of the table 120 and a view of the user interface 830 (e.g., a photo gallery application), the user interface including a series of digital images, e.g., user interface elements 812, 814, 816, 818, each digital image having a different quality (e.g., resolution, entropy, size, data size, etc.). Additionally, the user 102 is interacting (e.g., viewing and pointing) with a specific photo of the user interface 830, the user interface element 818 (e.g., a lower-resolution digital photographic image of a person and a house). Specifically, Figure 9A View 910A illustrates the intention (attention) of the user 102 at a first time instance, the intention (attention) being directed at the user interface element 818, as illustrated by the left finger pointing (e.g., hand representation 222) and the gaze along the path 902. Figure 9B View 910B illustrates the graphical indication 920 at a second time instance, the graphical indication matching the shape of the user interface element 818 and being based on the response to Figure 9A the user activity (e.g., intention or attention to the user interface element 818) illustrated in. For example, compared to the graphical indications 820 and 822 of Figure 8B , the graphical indication 920 provides another example of a photographic visual effect or glow that matches the external shape of the person and the house in the photo. Determining to provide a photographic visual effect (e.g., graphical indication 920) may be based on determining that the element is a photo and the user interface element 818 is less than one or more confidence thresholds associated with the photo (e.g., resolution threshold, entropy threshold, data size threshold, size threshold relative to the current view, etc.). In other words, based on a lower-quality image, a lower-quality visual effect is applied. For example, the graphical indication 920 matches the external shape of the image of the user interface element 818 to indicate to the user 102 the user intention of the system to recognize an interaction with the photo associated with the user interface element 816 (e.g., increasing the size of the image, moving to another location within the XR environment 905, etc.).
[0071] In some specific implementations, determining the shape may be based on combining elements of a single interactive item (e.g., different layers / elements corresponding to content that the user perceives as a single element). For example, the user interface element 818 is displayed as a photo of a house and a person. However, in some specific implementations, the user interface element 818 may be two overlapping 2D layered photos (e.g., the person overlaps the house, but not completely). In this case, the shape determination process may consider the external shapes of the two elements combined into one element, so the graphical indication 920 goes further from the person in the lower left to under the house, even if the user may only be focusing on the house.
[0072] Figure 10 Illustrated is the use of the exemplary input support framework 1040 to generate interaction data 1050 based on hand data 1010, gaze data 1020, and user interface target data 1030, which can be provided to one or more applications and / or used by system processes to provide a desired user experience. In some embodiments, the input support process 1040 is configured to understand the intent of user interactions, generate input signals and events to create a reliable and consistent user experience across multiple applications, detect out-of-process inputs, and responsibly route them through the system. The input support process 1040 can, for example, arbitrate which applications, processes, and / or user interface elements should receive user input based on identifying which application or user interface element is the intended target of the user's activity. The input support process 1040 can keep sensitive user data (e.g., gaze, hand / body registration data, etc.) private; sharing only abstracted or high-level information with applications.
[0073] The input support process can obtain hand data 1010, gaze data 1020, and user interface target data 1030 and determine the user interaction state. In some embodiments, it does so in a user environment where multiple input modalities are available to the user (e.g., an environment where the user can interact directly as shown in Figure 2 or indirectly as shown in Figure 3 to achieve the same interaction with user interface elements). For example, the input support process can determine that the user's right hand is performing an intentional pinch and gaze interaction with a user interface element, the left hand is directly tapping the user interface element, or the left hand is trembling and thus idle / not doing anything related to the user interface. In some embodiments, the user interface target data 1030 includes information associated with user interface elements, such as scalable vector graphics (SVG) information for vector graphics (e.g., may have some information from basic shapes, paths, or may contain masks or clip paths) and / or other image data (e.g., RGB data for bitmap images or image metadata).
[0074] Based on determining the user interaction intention, the input support framework 1040 can generate interaction data 1050 (e.g., including interaction pose, manipulator pose, and / or interaction state). The input support framework can generate input signals and events that can be consumed by an application without the need for in-process customization or 3D input recognition algorithms. In some specific implementations, the input support framework provides the interaction data 1050 in a format that an application can consume as a touch event on a touchscreen or as a touchpad tap with a 2D cursor at a specific location. Doing so enables the same application (with little or no additional input recognition process) to interpret interactions across different environments, including new environments that the application was not initially created for and / or using new and different input modalities. Additionally, the application response to the input can be more reliable and consistent across applications within a given environment and across different environments. For example, a consistent user interface response is achieved for 2D interactions with applications on a tablet, mobile device, laptop, etc., and for 3D interactions with applications on an HMD and / or other 3D / XR devices.
[0075] The input support framework can also manage user activity data such that different applications are unaware of user activities related to other applications. For example, one application will not receive user activity information while the user types a password into another application. This can involve the input support framework accurately identifying which application the user's activity corresponds to and then routing the interaction data 1050 only to the correct application. Applications may utilize multiple processes to host different user interface elements (e.g., using an out-of-process photo picker) for various reasons (e.g., privacy). The input support framework can accurately identify which process the user's activity corresponds to and route the interaction data 1050 only to the correct process. The input support framework can use details about the UIs of multiple potential target applications and / or processes to disambiguate the input.
[0076] Figure 11 An example of interaction recognition of user activity and displaying a graphical indication based on the determined shape of user interface elements is illustrated. In this example, sensor data and / or user interface information on the device 105 is used to identify user interactions made by the user 102, for example, based on outward-facing image sensor data, depth sensor data, eye sensor data, motion sensor data, etc., and / or information available to the application providing the user interface. The sensor data can be monitored to detect user activities corresponding to engagement conditions corresponding to the start of a user interaction.
[0077] In this example, at block 1110, the process presents a 3D environment (e.g., an XR environment) including a view of user interface 1100, which includes virtual elements / objects (e.g., user interface element 1115). At block 1120, the process determines the shape of the user interface element (e.g., user interface element 1115). In this example, the process determines that user interface element 1115 is an interactive element and matches the shape of the outer portion or edge of user interface element 1115, which in this example is a star-shaped element. In some embodiments, determining the shape can be based on information associated with the web page or user interface 1100, such as SVG information for vector graphics and / or other image data. For example, there may be some information from basic shapes, paths, or a mask or clipping path that includes user interface element 1115. Additionally or alternatively, other image data associated with user interface element 1115 can include RGB data for bitmap images or image metadata. In some embodiments, determining the shape can be based on image mapping or other image recognition techniques for identifying photo items.
[0078] In some embodiments, determining the shape can be based on combining elements of a single interactive item (e.g., different layers / elements corresponding to content that the user perceives as a single element). For example, the concentric circles of element 247 can be different layers of multiple elements, which can then be treated as a combined single 2D element to determine the outer shape of the combined element. Another example is user interface element 818 of a house and a person. If user interface element 818 is two overlapping elements (e.g., the person overlaps the house but not completely), the shape determination process considers the outer shape of the two elements combined into one element, so the graphical indication 920 is further into the house below the person in the lower left.
[0079] At block 1130, the process receives user activity data, such as hand data and / or gaze information (e.g., the gaze direction 1105 of user 102). In an exemplary embodiment, the process can detect that user 102 has positioned hand 1122 within the field of view of an outward-facing image sensor. In some embodiments, the process can detect one or more specific single-handed or two-handed configurations (e.g., claw, pinch, point, flat hand, steady hand in any configuration, etc.) as an indication of hand engagement, or can simply detect the presence of the hand within the sensor's field of view to initiate the process.
[0080] At block 1140, the process identifies a user interaction event with a user interface element. In this example, the process identifies that the gaze direction 1105 of user 102 and the pointing direction of hand 1122 are on user interface element 1115 (or any other object within the view of the XR environment). However, the process can identify an object (e.g., user interface element 1115) based on only gaze or only hand activity.
[0081] At block 1150, the process displays a graphical indication based on the shape of the identified user interface element. In other words, the process matches graphical indication 1117 to the determined shape of user interface element 1115, rather than using a predefined shape (e.g., a rectangle around a circular element). In this example, graphical indication 1117 (e.g., a visual effect providing feedback), a star-shaped glow or highlight graphically differentiates user interface element 1115 to indicate that user interface element 1115 now has a different state (e.g., a "hover" state similar to the state of a traditional user interface icon when a cursor is over an item without a click / tap).
[0082] Additionally, at block 1150, the process can identify a gesture to be associated with the identified object and can update the object (or launch an application associated with the object) based on the pose of hand 1122. In this example, user 102 is gazing at user interface element 1115 while making a pinching gesture with hand 1122, and this pinching gesture can be interpreted as initiating an action on user interface element 1115, e.g., causing a "click" event similar to that of a traditional user interface icon (during which the cursor is positioned over the icon and a trigger such as a mouse click or touchpad tap is received) or a selection action similar to a touchscreen "tap" event.
[0083] In some specific implementations, the application that provides the user interface information does not need to be notified of the hover state and the associated feedback provided by the graphical indication 1117. Instead, hand involvement, object identification, and the display of feedback can be processed, for example, by an operating system process outside of the process (e.g., outside of the application process). For example, such a process can be provided via the input support process of the operating system. Doing so can reduce or minimize potentially sensitive user information (e.g., such as a constant gaze direction vector or a hand movement direction vector), which otherwise might be provided to the application to enable the application to process these functions within the application process. Whether and how to display the feedback can be specified by the application, even if it is executed within the process. For example, the application can define that an element should display hover feedback or highlight feedback, and define how the hover or highlight will occur, such that the out-of-process aspect (e.g., the operating system) can provide the hover or highlight according to the defined appearance. Alternatively, the feedback can be defined out of the process (e.g., defined only by the OS) or defined to use a default appearance / animation in case the application does not specify an appearance.
[0084] The recognition of such interactions with user interface elements can be based on functions performed both via the system process and via the application process. For example, the input process of the OS can interpret hand data and (optionally, also) gaze data from the sensors of the device to identify interaction events and provide limited or interpreted / abstracted information about the interaction event to the application that provides the user interface 1100. For example, the OS input support process can identify a 2D point on the user interface element 1115 within the 2D user interface 1100, e.g., an interaction pose, rather than providing gaze direction information that identifies the gaze direction 1105. Then, the application process can interpret the 2D point information (e.g., interpret the information as a selection, a mouse click, a touchscreen tap, or other input received at that point) and provide a response, e.g., modify its user interface accordingly.
[0085] Figure 11 Examples of identifying indirect user interactions to determine whether to display a graphical indication as feedback (e.g., hover) are illustrated. Numerous other types of indirect interactions can be identified, for example, based on one or more user actions that identify user interface elements and / or one or more user actions that provide input (e.g., no-action / hover type input, selection type input, input with direction, path, speed, acceleration, etc.). Inputs in 3D space similar to inputs on a 2D interface can be identified, e.g., inputs similar to mouse movement, mouse button click, touchscreen touch events, touchpad events, joystick events, game controller events, etc.
[0086] Some specific implementations utilize an input support framework outside of the process (e.g., outside of the application process) to facilitate accurate, consistent, and efficient input recognition in a manner that protects private user information. For example, aspects of the input recognition process can be performed outside of the process such that the application has little or no access to information about where the user is looking, e.g., the gaze direction. In some specific implementations, application access to some user activity information (e.g., data based on gaze direction) is limited to specific types of user activities, e.g., activities that meet specific criteria. For example, the application can be restricted to receiving only information associated with deliberate or intentional user activities (e.g., deliberate or intentional actions that indicate an intention to interact with a user interface element (e.g., select, activate, move, etc. a user interface element)).
[0087] Some specific implementations use functional elements that are executed via both the application process and system processes outside of the application process to identify input. Thus, contrary to frameworks where all (or most) input recognition functionality is managed within the application process, some algorithms involved in input recognition can be moved outside of the process, e.g., outside of the application process. For example, this can involve moving algorithms for detecting gaze input and intent outside of the application's process such that the application does not have access to user activity data corresponding to where the user is looking, or in some cases only has access to such information, e.g., only for specific instances where the user exhibits an intention to interact with a user interface element.
[0088] Some specific implementations use a model in which the application declares or otherwise provides information about its user interface elements such that system processes outside of the application process can better facilitate input recognition. For example, the application can declare the locations and / or user interface behaviors / capabilities of its buttons, scroll bars, menus, objects, and other user interface elements. Such declarations can identify how the user interface should behave given different types of user activities, e.g., the button should (or should not) exhibit hover feedback when the user is looking at it.
[0089] System processes (e.g., outside of application processes) can use such information to provide desired user interface behavior (e.g., provide hover feedback via graphical indications in the context of appropriate user activity). For example, a system process can trigger a graphical indication (hover feedback) of a user interface element based on a claim from an application that indicates that the user interface of the application includes the element and that the user interface should display hover feedback, for example, when gazed at. The system process can provide such hover feedback based on identifying a triggering user activity (e.g., gazing at a user interface object), and can do so without revealing to the application details of the user activity associated with triggering the hover, the occurrence of the user activity that triggers the hover feedback, and / or the provision of the hover feedback. The application may not know the user's gaze direction and / or that hover feedback is provided for a user interface element.
[0090] Some aspects of input recognition can be handled by the application itself (e.g., in-process). However, system processes can filter, abstract, or otherwise manage the information available to the application to identify input to the application. The system process can do so in a way that promotes efficient, accurate, and consistent (within the application and across multiple applications) input recognition and allows the application to potentially use input recognition processes that are easier to implement and / or traditional input recognition processes, such as input recognition processes developed for different systems or input environments (e.g., using touchscreen input processes used in traditional mobile applications).
[0091] Some embodiments use system processes to provide interaction event data to applications so that these applications can recognize inputs. The interaction event data can be restricted such that all user activity data is not available to these applications. Providing only limited user activity information can help protect user privacy. The interaction event data can be configured to correspond to events that can be recognized by applications using general or traditional recognition processes. For example, a system process can interpret 3D user activity data to provide interaction event data that the application can recognize in the same way that the application would recognize a touch event on a touchscreen. In some embodiments, an application receives interaction event data that corresponds only to certain types of user activity (e.g., intentional or deliberate actions on user interface objects), and may not receive information about other types of user activity (e.g., only gazing activity, the user moving their hand in a way not associated with interacting with the user interface, the user moving closer to or further away from the user interface, etc.). In one example, over a period of time (e.g., one minute, 10 minutes, etc.), a user gazes around a 3D XR environment, including gazing at certain user interface text, buttons, and other user interface elements, and eventually performs an intentional user interface interaction, e.g., by making an intentional pinch gesture while gazing at button X. The system process can process all user interface feedback during the time of gazing around the various user interface elements without providing application information about those gazes. On the other hand, the system process can provide interaction event data to the application based on the intentional pinch gesture when gazing at button X. However, even this interaction event data can provide limited information to the application, e.g., providing an interaction location or pose that identifies the interaction point on button X without providing information about the actual gaze direction. The application can then interpret the interaction point as an interaction with button X and respond accordingly. Thus, user behavior not associated with intentional interaction of the user with user interface elements (e.g., only gazing, hovering, menu expansion, reading, etc.) is processed outside the process where the application does not have access to user data, and information about intentional user interface element interactions is restricted such that the information does not include all user activity details.
[0092] Figure 12FIG. 1200 is a flow diagram illustrating a method 1200 for providing a graphical indication of a determined shape corresponding to a user interface element based on identifying user interaction events corresponding to user activities. In some embodiments, a device such as electronic device 105 or electronic device 110 executes method 1200. In some embodiments, method 1200 is executed on a mobile device, a desktop computer, a laptop computer, an HMD, or a server device. Method 1200 is executed by processing logic (including hardware, firmware, software, or combinations thereof). In some embodiments, method 1200 is executed on a processor executing code stored in a non-transitory computer-readable medium (e.g., memory).
[0093] Various embodiments of method 1200 disclosed herein are based on a user's intent (attention) with respect to a user interface element to graphically indicate a determined shape of a user interface element (e.g., on a 2D web page) viewed in an XR environment using a 3D display device (e.g., a wearable device such as HMD (device 105)). The goal is to match an arbitrary shape of a user interface element with a visual effect that provides illumination or highlighting. Various embodiments of method 1200 disclosed herein can match a graphical indication to a determined shape of a user interface element rather than using a predefined shape (e.g., a standard rectangular shape around a circular element). In other words, if the user interface element is a star shape, the illumination or highlighting around the user interface element will also display a star shape.
[0094] At block 1202, method 1200 presents a view of a 3D environment in which one or more user interface elements are positioned at 3D locations based on a 3D coordinate system associated with the 3D environment. For example, as Figure 2 shown, a 2D web page (e.g., user interface 230) can be viewed using a 3D device (e.g., device 105). In some embodiments, during an input support process, the process includes obtaining data corresponding to the positioning of user interface elements of an application within a 3D coordinate system. The data can correspond to the positioning of the user interface elements based at least in part on data provided by the application (e.g., the position / shape of 2D elements intended for a 2D window area), e.g., such as user interface information provided from the application to an operating system process. In some embodiments, the operating system manages information regarding virtual and / or real content located within a 3D coordinate system. Such a 3D coordinate system can correspond to an XR environment representing a physical environment and / or virtual content corresponding to content from one or more applications. An executing application can provide information regarding the positioning of its user interface elements via a hierarchical tree (e.g., a declarative hierarchical tree), where some layers are identified for remote (i.e., outside of the application process) input effects.
[0095] At block 1204, method 1200 determines the shape of the one or more user interface elements. For example, the process identifies the 2D shape or 3D shape or geometric representation of the user interface element (e.g., star shape, circular shape, building shape, etc.). In some embodiments, determining the shape may be based on identifying / detecting the user interface element as an interactive item (e.g., an interactive icon associated with an application that a user can click).
[0096] In some embodiments, determining the shape is based on formatting information associated with the one or more user interface elements. For example, the formatting information may be based on information associated with a web page or user interface, such as SVG information for vector graphics and / or other image data. For example, there may be some information from basic shapes, paths, or may include a mask or clipping path of the user interface element. Additionally or alternatively, other image data associated with the user interface element may include RGB data for bitmap images or image metadata. In some embodiments, determining the shape may be based on image mapping or other image recognition techniques for identifying photo items. For example, detecting user interface elements based on image mapping to provide a simpler way to link various parts of an image without dividing the image into separate image files.
[0097] At block 1206, method 1200 receives data corresponding to user activity in the 3D coordinate system. For example, the user activity data may include hand data and gaze data, or data corresponding to other input modalities (e.g., an input controller). As described with respect to Figure 11 such data may include, but is not limited to, hand data, gaze data, and / or human interface device (HID) data. Various combinations of a single type of data or two or more different types of data may be received, such as hand data and gaze data, controller data and gaze data, hand data and controller data, voice data and gaze data, voice data and hand data, etc. Different combinations of sensor / HID data may correspond to different input modalities. In one exemplary embodiment, the data includes both hand data (e.g., a hand pose skeleton that identifies 20+ joint positions) and gaze data (e.g., a gaze vector stream), and both the hand data and the gaze data may be related to identifying inputs via direct touch input modalities and indirect touch input modalities.
[0098] At block 1208, method 1200 identifies a user interaction event associated with a first user interface element in the 3D environment based on the data corresponding to the user activity. For example, the user interaction event may be determined based on gaze and / or pinch data using the direction based on eye gaze, head, hand, arm, etc. to determine whether the user is focusing (paying attention) on a particular object (e.g., a user interface element). In some implementations, identifying the user interaction event may be based on determining that the pupil response corresponds to directing attention to an area associated with the user interface element. In some implementations, identifying the user interaction event may be based on finger pointing and hand movement gestures. In some implementations, the user interaction event is based on the direction of the user's gaze or face relative to the user interface. The direction of the user's face relative to the user interface can be determined by extending a ray from a position on the user's face and determining the intersection of the ray with a visual element on the user interface.
[0099] At block 1210, in accordance with identifying the user interaction event, method 1200 provides a graphical indication (e.g., visual effect glow, highlight, etc.) corresponding to the shape of the determined first user interface element. For example, after identifying that a particular user interface element is associated with an interaction event (user attention), the graphical indication matches the boundary / shape of the item, rather than using a predefined shape (e.g., a rectangle around a circular affordance). For example, as Figure 5B , Figure 6B and Figure 7B illustrate, graphical indications 520, 620, 720 respectively match the identified shapes of the associated user interface elements (e.g., the star-shaped graphical indication 520 matches the star-shaped user interface element 246). In some implementations, the graphical indication may be configured based on the determined item type, such as configuring the effect (e.g., based on a confidence threshold) for a photo item using attributes such as size, entropy, and / or resolution. In some implementations, for example, there may be size constraints for the graphical indication if the user interface element is too large. In some implementations, the color of the user interface element may match the visual appearance of the graphical indication. In some implementations, the shape of the glow effect may vary based on the determined shape of the object (e.g., a circular shape glow with a gradient centered at the position the user is gazing; but if it is a star shape, there may be a matching star-shaped glow).
[0100] In some specific implementations, determining the shape may be based on combining elements of a single interactive item (e.g., different layers / elements corresponding to content that a user perceives as a single element). In an exemplary specific implementation, determining the shape of the first user interface element includes: identifying sub-elements of the first user interface element; and determining the shape of the first user interface element based on the identified sub-elements. In some specific implementations, the graphical indication corresponds to the identified sub-elements of the first user interface element. For example, the figure indication 920 further into the house below by the figure of the lower left of the user interface element 818 of Figure 9B enters further below the house.
[0101] In some specific implementations, the graphical indication visual effect may remove the transparent area. In an exemplary specific implementation, providing the graphical indication includes removing a part of the view of the first user interface element within the view of the 3D environment (e.g., removing the transparent area of the user interface element).
[0102] In some specific implementations, there are time aspects for determining when to display the graphical indication. In an exemplary specific implementation, the graphical indication is displayed for a first instance based on one or more first attributes, and the graphical indication is displayed for a second instance different from the first instance based on one or more second attributes different from the first attributes. For example, the graphical indication may glow a few milliseconds after the user's attention, and / or the visual effect (glowing) may become brighter based on the user's gaze on the element for a specific period of time (e.g., greater than one second or two seconds). For example, when the user has continued to focus on the user interface element 247 for a period of time, as Figure 7C shown, the graphical indication 722 is greater than Figure 7B the graphical indication 720.
[0103] In some specific implementations, the interaction event data may include interaction pose (e.g., 6DOF data of points on the user interface of the application), manipulator pose (e.g., 3D position of the stable hand center or pinch centroid), interaction state (e.g., direct, indirect, hovering, pinching, etc.) and / or identifying which user interface element is being interacted with. In some specific implementations, the interaction data may exclude data associated with user activities occurring between intentional events. The interaction event data may exclude detailed sensor / HID data, such as hand bone data. The interaction event data may extract detailed sensor / HID data to avoid providing the application with data that is unnecessary for the application to recognize the input and may be private to the user.
[0104] In some specific implementations, method 900 may display a view of an extended reality (XR) environment corresponding to a (3D) coordinate system, where user interface elements of an application are displayed in the view of the XR environment. Such an XR environment may include user interface elements from multiple application processes corresponding to multiple applications, and an input support process may identify interaction event data for the multiple applications and route the interaction event data only to the appropriate application, e.g., the application that the user intends to interact with. Accurately routing the data only to the intended application can help ensure that one application does not misuse input data intended for another application (e.g., one application does not track a user entering a password into another application).
[0105] In some specific implementations, data corresponding to user activities may have various formats and be based on or include (but are not limited to being based on or including) sensor data (e.g., hand data, gaze data, head pose data, etc.) or HID data. In some specific implementations, data corresponding to user activities includes gaze data that includes a stream of gaze vectors corresponding to the gaze direction over time during the use of an electronic device. Data corresponding to user activities may include hand data that includes the hand pose skeletons of multiple joints at each of multiple moments during the use of the electronic device. Data corresponding to user activities may include both hand data and gaze data. Data corresponding to user activities may include controller data and gaze data. Data corresponding to user activities may include, but is not limited to, any combination of one or more types of data associated with one or more sensors or one or more sensor types, associated with one or more input modalities, associated with one or more parts of a user (e.g., eyes, nose, cheeks, mouth, hands, fingers, arms, torso, etc.) or the entire user, and / or associated with one or more items worn or held by the user (e.g., a mobile device, a tablet, a laptop computer, a laser pointer, a handheld controller, a wand, a ring, a watch, a bracelet, a necklace, etc.).
[0106] In some specific implementations, method 900 further includes identifying interaction event data for an application and may involve identifying only certain types of activities within user activities to be included in the interaction event data. In some specific implementations, activities within user activities that are determined to correspond to unintentional events rather than intentional user interface element inputs (e.g., the type of activity) are excluded from the interaction event data. In some specific implementations, passive only gaze activities within user activities are excluded from the interaction event data. Such passive only gaze behavior (non-intentional input) is different from intentional only gaze interactions (e.g., gaze dwell, or gaze to perform an air gesture until a gaze HUD is invoked / dismissed, etc.).
[0107] Identifying interaction event data for an application can involve identifying only specific attributes of the data corresponding to user activity to include in the interaction event data, such as including the hand center instead of the positions of all joints used to model the hand, including a single gaze direction or a single HID pointing direction for a given interaction event. In another example, the starting position of the gaze direction / HID pointing direction is altered or suppressed, such as to obscure data indicating how far the user is from the user interface or where the user is in a 3D environment. In some implementations, the data corresponding to user activity includes hand data representing the positions of multiple joints of a hand, and the interaction event data includes a single hand pose provided in place of the hand data.
[0108] In some implementations, method 900 is performed by an electronic device that is a head-mounted device (HMD), and / or the XR environment is a virtual reality environment or an augmented reality environment.
[0109] Figure 13 is a block diagram of electronic device 1300. Device 1300 illustrates an exemplary device configuration of electronic device 110 or electronic device 105. Although certain specific features are illustrated, those skilled in the art will recognize from this disclosure that, for the sake of brevity and to not obscure more relevant aspects of the specific implementations disclosed herein, various other features are not illustrated. To that end, as a non-limiting example, in some implementations, device 1300 includes one or more processing units 1302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 1306, one or more communication interfaces 1308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE, SPI, I2C, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 1310, one or more output devices 1312, one or more internal and / or external-facing image sensor systems 1314, memory 1320, and one or more communication buses 1304 for interconnecting these components and various other components.
[0110] In some specific embodiments, one or more communication buses 1304 include circuitry that interconnects and controls communication between system components. In some specific embodiments, one or more I / O devices and sensors 1306 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, or one or more depth sensors (e.g., structured light, time-of-flight, etc.), etc.
[0111] In some specific embodiments, one or more output devices 1312 include one or more displays configured to present a view of a 3D environment to a user. In some specific embodiments, one or more output devices 1312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some specific embodiments, one or more displays correspond to diffractive, reflective, polarization, holographic, etc. waveguide displays. In one example, device 1300 includes a single display. In another example, device 1300 includes a display for each eye of the user.
[0112] In some specific embodiments, one or more output devices 1312 include one or more audio generating devices. In some specific embodiments, one or more output devices 1312 include one or more speakers, surround sound speakers, a speaker array, or headphones for generating spatialized sound such as 3D audio effects. Such devices can virtually place sound sources in a 3D environment, including behind, above, or below one or more listeners. Generating spatialized sound can involve transforming sound waves (e.g., using head-related transfer functions (HRTF), reverberation, or cancellation techniques) to simulate natural sound waves (including reflections from walls and floors) that emanate from one or more points in the 3D environment. The spatialized sound can induce the listener's brain to interpret the sound as if it were occurring at one or more points in the 3D environment (e.g., from one or more specific sound sources), even though the actual sound may be generated by speakers in other locations. One or more output devices 1312 can additionally or alternatively be configured to generate haptic sensations.
[0113] In some specific implementations, one or more image sensor systems 1314 are configured to obtain image data corresponding to at least a portion of a physical environment. For example, one or more image sensor systems 1314 may include one or more RGB cameras (e.g., having a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), monochrome cameras, IR cameras, depth cameras, event-based cameras, etc. In various specific implementations, one or more image sensor systems 1314 further include an illumination source that emits light, such as a flash. In various specific implementations, one or more image sensor systems 1314 further include an on-camera image signal processor (ISP) that is configured to perform a plurality of processing operations on the image data.
[0114] Memory 1320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some specific implementations, memory 1320 includes non-volatile memory, such as one or more disk storage devices, optical disc storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 1320 optionally includes one or more storage devices that are remotely located from one or more processing units 1302. Memory 1320 includes a non-transitory computer-readable storage medium.
[0115] In some specific implementations, memory 1320 or the non-transitory computer-readable storage medium of memory 1320 stores an optional operating system 1330 and one or more instruction sets 1340. The operating system 1330 includes procedures for handling various basic system services and for performing hardware-related tasks. In some specific implementations, the instruction set 1340 includes executable software defined by binary information stored in the form of charges. In some specific implementations, the instruction set 1340 is software that can be executed by one or more processing units 1302 to implement one or more of the techniques described herein.
[0116] The instruction set 1340 includes a user interaction instruction set 1342 that is configured to identify and / or interpret user gestures and other user activities as described herein when executed. The instruction set 1340 includes an application instruction set 1344 for one or more applications. In some specific implementations, each of these applications is provided as a separately executable code set, e.g., capable of being executed via an application process. The instruction set 1340 may be embodied as a single software executable file or multiple software executable files.
[0117] Although instruction set 1340 is shown as residing on a single device, it should be understood that in other embodiments, any combination of elements may be located in separate computing devices. Additionally, the figures are more of a functional description of the various features that exist in a particular embodiment, as opposed to a structural schematic of the embodiments described herein. As would be recognized by one of ordinary skill in the art, items shown separately may be combined and some items may be separated. The actual number of instruction sets and how the features are allocated therein will vary depending on the embodiment and may depend in part on the particular combination of hardware, software, and / or firmware selected for a particular embodiment.
[0118] Figure 14 A block diagram of an exemplary head-mounted device 1400 in accordance with some embodiments is illustrated. The head-mounted device 1400 includes a housing 1401 (or enclosure) that houses various components of the head-mounted device 1400. The housing 1401 includes (or is coupled to) an eye pad (not shown) disposed at a proximal (to user 102) end of the housing 1401. In various embodiments, the eye pad is a plastic or rubber piece that comfortably and snugly holds the head-mounted device 1400 in place on the face of user 102 (e.g., around the eyes of user 102).
[0119] The housing 1401 houses a display 1410 that displays an image, emits light towards or into the eyes of user 102. In various embodiments, the display 1410 emits light through an eyepiece having one or more optical elements 1405 that refract the light emitted by the display 1410 such that the display appears to user 102 to be at a virtual distance that is farther than the actual distance from the eyes to the display 1410. For example, the optical elements 1405 may include one or more lenses, waveguides, other diffractive optical elements (DOEs), etc. To enable user 102 to focus on the display 1410, in various embodiments, the virtual distance is at least greater than the minimum focal length of the eyes (e.g., 7 cm). Additionally, to provide a better user experience, in various embodiments, the virtual distance is greater than 1 meter.
[0120] The housing 1401 also houses a tracking system that includes one or more light sources 1422, cameras 1424, camera 1432, camera 1434, camera 1436, and a controller 1480. One or more light sources 1422 emit light onto the eyes of the user 102, and the light is reflected as a light pattern (e.g., a flash ring) that can be detected by the camera 1424. Based on the light pattern, the controller 1480 can determine the eye tracking characteristics of the user 102. For example, the controller 1480 can determine the gaze direction and / or blink state (open or closed eyes) of the user 102. As another example, the controller 1480 can determine the pupil center, pupil size, or point of regard. Thus, in various embodiments, light is emitted by one or more light sources 1422, reflected from the eyes of the user 102, and detected by the camera 1424. In various embodiments, light from the eyes of the user 102 is reflected from a hot mirror or passes through an eyepiece before reaching the camera 1424.
[0121] The display 1410 emits light in a first wavelength range, and one or more light sources 1422 emit light in a second wavelength range. Similarly, the camera 1424 detects light in the second wavelength range. In various embodiments, the first wavelength range is the visible wavelength range (e.g., a wavelength range of approximately 400 nm - 700 nm within the visible spectrum), and the second wavelength range is the near-infrared wavelength range (e.g., a wavelength range of approximately 700 nm - 1400 nm within the near-infrared spectrum).
[0122] In various embodiments, eye tracking (or specifically, the determined gaze direction) is used to enable the user to interact (e.g., the user 102 selects an option on the display 1410 by looking at it), provide perforated rendering (e.g., presenting a higher resolution in the area of the display 1410 that the user 102 is looking at and a lower resolution elsewhere on the display 1410), or correct for distortion (e.g., for an image to be provided on the display 1410).
[0123] In various embodiments, one or more light sources 1422 emit light toward the eyes of the user 102, and the light is reflected in the form of multiple flashes.
[0124] In various embodiments, the camera 1424 is a frame / shutter-based camera that generates an image of the eyes of the user 102 at a frame rate at a particular time point or multiple time points. Each image includes a matrix of pixel values corresponding to the pixels of the image, which correspond to the positioning of the light sensor matrix of the camera. In an embodiment, each image is used to measure or track pupil dilation by measuring changes in the pixel intensity associated with one or both of the user's pupils.
[0125] In various embodiments, camera 1424 is an event camera that includes a plurality of light sensors (e.g., a light sensor matrix) at a plurality of respective locations, and the event camera generates an event message indicating a specific location of a specific light sensor in response to a change in light intensity detected by the specific light sensor.
[0126] In various embodiments, cameras 1432, 1434, and 1436 are frame / shutter-based cameras that can generate an image of the face of user 102 or capture the external physical environment at a frame rate at a particular point in time or multiple points in time. For example, camera 1432 captures an image of the user's face below the eyes, camera 1434 captures an image of the user's face above the eyes, and camera 1436 captures the user's external environment (e.g., environment 100 of FIG. 1). The images captured by cameras 1432, 1434, and 1436 may include light intensity images (e.g., RGB) and / or depth image data (e.g., time-of-flight, infrared, etc.).
[0127] It should be understood that the embodiments described above are cited by way of example, and the present disclosure is not limited to what has been particularly shown and described above. Instead, the scope includes both combinations and sub-combinations of the various features described above, as well as variations and modifications of the various features that will occur to those skilled in the art upon reading the foregoing description and that are not disclosed in the prior art.
[0128] As described above, one aspect of the present technology is to collect and use sensor data that may include user data to improve the user experience of an electronic device. The present disclosure contemplates that, in some cases, the collected data may include personal information data that uniquely identifies a specific person or can be used to identify the interests, characteristics, or tendencies of a specific person. Such personal information data may include motion data, physiological data, demographic data, location-based data, phone numbers, email addresses, home addresses, device characteristics of personal devices, or any other personal information.
[0129] The present disclosure recognizes that the use of such personal information data in the technology of the present invention can be used to benefit the user. For example, personal information data can be used to improve the content viewing experience. Thus, the use of such personal information data may enable planned control of the electronic device. In addition, the present disclosure also anticipates other uses of personal information data that are beneficial to the user.
[0130] The present disclosure also contemplates that entities responsible for the collection, analysis, disclosure, transmission, storage, or other use of such personal information and / or physiological data will comply with established privacy policies and / or privacy practices. Specifically, such entities should implement and adhere to privacy policies and measures that are recognized as meeting or exceeding industry or government requirements for maintaining the privacy and security of personal information data. For example, personal information from users should be collected for legitimate and reasonable uses of the entity and not shared or sold outside of those legitimate uses. Additionally, such collection should only occur after the user's informed consent. Additionally, such entities should take any steps necessary to safeguard and protect access to such personal information data and ensure that others who have access to the personal information data comply with their privacy policies and procedures. Additionally, such an entity may subject itself to third-party assessments to demonstrate its compliance with widely accepted privacy policies and practices.
[0131] Regardless of the foregoing, the present disclosure also contemplates ways in which users can selectively block the use or access of personal information data. That is, the present disclosure anticipates that hardware or software elements may be provided to prevent or block access to such personal information data. For example, in the case of a content delivery service customized for a user, the techniques of the present invention can be configured to allow the user to select to "opt-in" or "opt-out" of participating in the collection of personal information data during the registration for the service. In another example, the user may choose not to provide personal information data for a targeted content delivery service. In yet another example, the user may choose not to provide personal information but allow the transmission of anonymized information for use in improving the functionality of the device.
[0132] Accordingly, while the present disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, the present disclosure also anticipates that various embodiments may also be implemented without access to such personal information data. That is, the various embodiments of the techniques of the present invention will not be rendered inoperative due to the lack of all or a portion of such personal information data. For example, preferences or settings can be inferred by relying on non-personal information data or an absolute minimum amount of personal information such as the content requested by a device associated with the user, other non-personal information available to the content delivery service, or publicly available information, and content can be selected and delivered to the user based on those inferences.
[0133] In some embodiments, a public key / private key system that only allows the owner of the data to decrypt the stored data is used to store the data. In some other specific implementations, the data may be stored anonymously (e.g., without identification and / or without personal information about the user, such as legal name, username, time, and location data, etc.). In this way, other users, hackers, or third parties cannot determine the identity of the user associated with the stored data. In some specific implementations, a user may access their stored data from a user device different from the user device used to upload the stored data. In these cases, the user may need to provide login credentials to access their stored data.
[0134] Numerous specific details are set forth herein to provide a thorough understanding of the claimed subject matter. However, those skilled in the art will understand that the claimed subject matter may be practiced without these specific details. In other instances, well-known methods, devices, or systems have not been described in detail so as not to obscure the claimed subject matter.
[0135] Unless otherwise specifically stated, it should be understood that throughout the specification, discussions using terms such as "processing", "computing", "computed", "determining", and "identifying" refer to actions or processes of a computing device, such as one or more computers or similar electronic computing devices, that manipulate or transform data represented as physical electronic or magnetic quantities within a memory, register, or other information storage device, transmission device, or display device of a computing platform.
[0136] One or more of the systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement that provides results conditioned on one or more inputs. Suitable computing devices include computer systems based on general-purpose microprocessors that access stored software that programs or configures the computing system from a general-purpose computing device into a special-purpose computing device that implements one or more specific implementations of the subject matter of the present invention. Any suitable programming, scripting, or other type of language or combination of languages may be used to implement the teachings contained herein in the software used to program or configure the computing device.
[0137] Specific implementations of the methods disclosed herein may be performed in the operation of such computing devices. The order of the blocks presented in the examples above may be varied, e.g., the blocks may be reordered, combined, and / or divided into sub-blocks. Certain blocks or processes may be performed in parallel.
[0138] The use of "is applicable to" or "is configured to" in this document means open and inclusive language that does not exclude devices that are applicable to or configured to perform additional tasks or steps. Additionally, the use of "based on" means open and inclusive because a process, step, calculation, or other action "based on" one or more of the stated conditions or values can in practice be based on additional conditions or values beyond those stated. The headings, lists, and numbers included in this document are for ease of explanation only and are not intended to be restrictive.
[0139] It will also be understood that although terms such as "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first node may be referred to as a second node, and similarly, a second node may be referred to as a first node, which changes the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. The first node and the second node are both nodes, but they are not the same node.
[0140] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of this particular embodiment and the appended claims, the singular forms "a", "an", and "the" are intended to also cover the plural forms unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will also be understood that the term "comprising", as used in this specification, specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0141] As used herein, the term "if" can be interpreted to mean "when the precondition is true" or "while the precondition is true" or "in response to determining" or "in accordance with determining" or "in response to detecting" the precondition is true, depending on the context. Similarly, the phrases "if it is determined [that the precondition is true]" or "if [the precondition is true]" or "when [the precondition is true]" are interpreted to mean "when it is determined that the precondition is true" or "in response to determining" or "in accordance with determining" the precondition is true or "when it is detected that the precondition is true" or "in response to detecting" the precondition is true, depending on the context.
[0142] The foregoing description and summary of the invention are to be understood as illustrative and exemplary in every respect and not restrictive, and the scope of the invention disclosed herein is determined not only by the detailed description of the illustrative embodiments but by the full breadth permitted by patent law. It should be understood that the specific embodiments shown and described herein are merely illustrative of the principles of the invention and that various modifications can be effected by those skilled in the art without departing from the scope and spirit of the invention.
Claims
1. A method, the method comprising: At an electronic device having a processor: Present a view of a three-dimensional (3D) environment, wherein one or more user interface elements are positioned at 3D positions based on a 3D coordinate system associated with the 3D environment; Determine the shape of the one or more user interface elements; Receive data corresponding to user activity in the 3D coordinate system; Identify a user interaction event associated with a first user interface element in the 3D environment based on the data corresponding to the user activity; And Provide a graphical indication corresponding to the determined shape of the first user interface element based on identifying the user interaction event.
2. The method according to claim 1, wherein determining the shape based on identifying user interface elements includes interactive items.
3. The method according to claim 1, wherein determining the shape is based on formatting information associated with the one or more user interface elements.
4. The method according to claim 1, wherein determining the shape is based on an image map associated with the one or more user interface elements.
5. The method according to claim 1, wherein providing the graphical indication corresponding to the determined shape of the first user interface element includes matching the shape of the graphical indication to the shape of the first user interface element.
6. The method according to claim 1, wherein determining the shape of the first user interface element includes: Identifying sub-elements of the first user interface element; And Determining the shape of the first user interface element based on the identified sub-elements, wherein the graphical indication corresponds to the identified sub-elements of the first user interface element.
7. The method according to claim 1, wherein providing the graphical indication includes removing a portion of the view of the first user interface element within the view of the 3D environment.
8. The method according to claim 1, wherein the graphical indication is a highlighting effect or a glowing effect corresponding to the user interface element.
9. The method according to claim 1, wherein the graphical indication is based on determining the type of the user interface element.
10. The method according to claim 1, wherein the graphical indication is displayed for a first instance based on one or more first attributes, and wherein the graphical indication is displayed for a second instance different from the first instance based on one or more second attributes different from the first attributes.
11. The method according to claim 1, wherein the data corresponding to the user activity is obtained via one or more sensors on the device.
12. The method according to claim 1, wherein the data corresponding to the user activity includes gaze data, the gaze data including a stream of gaze vectors corresponding to the gaze direction over time during use of the electronic device.
13. The method according to claim 1, wherein the data corresponding to the user activity includes hand data, the hand data including hand pose skeletons of multiple joints at each of multiple moments during use of the electronic device.
14. The method according to claim 1, wherein the data corresponding to the user activity includes hand data and gaze data.
15. The method according to claim 1, wherein the data corresponding to the user activity includes controller data and gaze data.
16. The method according to claim 1, wherein the data corresponding to the user activity includes head pose data of the user.
17. The method according to claim 1, wherein the electronic device includes a head-mounted device (HMD).
18. An apparatus, the apparatus comprising: A non-transitory computer-readable storage medium; And One or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium includes program instructions that, when executed on the one or more processors, cause the one or more processors to perform operations, the operations including: Presenting a view of a three-dimensional (3D) environment, wherein one or more user interface elements are positioned at 3D locations based on a 3D coordinate system associated with the 3D environment; Determining the shape of the one or more user interface elements; Receiving data corresponding to a user activity in the 3D coordinate system; Identifying a user interaction event associated with a first user interface element in the 3D environment based on the data corresponding to the user activity; and Providing a graphical indication corresponding to the determined shape of the first user interface element according to the identified user interaction event.
19. The apparatus according to claim 18, wherein determining the shape is based on: Identifying that the user interface element includes an interactive item; Formatting information associated with the one or more user interface elements; or An image map associated with the one or more user interface elements.
20. A non-transitory computer-readable storage medium storing program instructions that can be executed on a device to perform operations, the operations including: Presenting a view of a three-dimensional (3D) environment, wherein one or more user interface elements are positioned at 3D locations based on a 3D coordinate system associated with the 3D environment; Determining the shape of the one or more user interface elements; Receiving data corresponding to a user activity in the 3D coordinate system; Identifying a user interaction event associated with a first user interface element in the 3D environment based on the data corresponding to the user activity; And Providing a graphical indication corresponding to the determined shape of the first user interface element according to the identified user interaction event.