Method and apparatus for managing an attention accumulator
By detecting changes in gaze direction and adjusting the attention accumulator value, the selection of user interface elements is automatically switched, which solves the noise and inaccuracy problems of eye tracking in the user interface and improves the continuity of the user interface and the user experience.
Patent Information
- Application Number
- CN202310079157.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-01-19
- Filing Date
- 2023-01-17
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2043-01-17
AI Technical Summary
Eye tracking can be noisy and inaccurate in user interface selection, leading to user experience issues such as discontinuity or jumpiness.
By detecting changes in the user's gaze direction, the attention accumulator value of user interface elements is adjusted, and the selection of user interface elements is automatically switched to reduce discontinuities in user interaction.
It improves the continuity of user interface selection and user experience, and reduces discontinuity issues caused by eye-tracking noise and inaccuracy.
Smart Images

Figure CN116458881B_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 300,941, filed January 19, 2022, which is hereby incorporated by reference in its entirety. TECHNICAL FIELD
[0003] The present disclosure relates generally to selecting user interface (UI) elements, and in particular to systems, devices, and methods for selecting UI elements with an eye tracking based attention accumulator. BACKGROUND
[0004] Generally, eye tracking can be noisy and / or inaccurate. A dwell time timer can be used to select UI elements with eye tracking input. However, using a dwell time timer to switch between UI elements can result in additional user experience (UX) issues, such as discontinuity or jank. BRIEF DESCRIPTION OF DRAWINGS
[0005] Accordingly, the present disclosure can be understood with reference to some illustrative aspects of specific implementations, some of which are illustrated in the attached drawings.
[0006] Figure 1 is a block diagram of an example operational architecture according to some implementations.
[0007] Figure 2 is a block diagram of an example controller according to some implementations.
[0008] Figure 3 is a block diagram of an example electronic device according to some implementations.
[0009] Figure 4A is a block diagram of a first portion of an example content delivery architecture according to some implementations.
[0010] Figure 4B is shown an example data structure according to some implementations.
[0011] Figure 4C is a block diagram of a second portion of an example content delivery architecture according to some implementations.
[0012] Figures 5A to 5H is shown a sequence of an example of a content delivery scenario according to some implementations.
[0013] Figures 6A to 6H is shown another sequence of an example of a content delivery scenario according to some implementations.
[0014] Figures 7A to 7CA flowchart representation of a method of selecting a UI element with an eye tracking based attention accumulator is shown in accordance with some implementations.
[0015] In accordance with common practice the various features described with reference to the drawings can not be drawn to scale. In the interest of clarity, not all components of the systems, methods, and devices are necessarily shown in the accompanying drawings and figures. Finally, elements of similar structure or function are generally designated with like reference numerals throughout the entire disclosure. SUMMARY
[0016] Various implementations disclosed herein include devices, systems, and methods for selecting a UI element with an eye tracking based attention accumulator. According to some implementations, the method is performed at a computing system that includes one or more processors and non-transitory memory, where the computing system is communicatively coupled to a display device and one or more input devices. The method includes: while a first user interface (UI) element is currently selected, detecting a first gaze direction that is directed to a second UI element that is different from the first UI element; in response to detecting the first gaze direction that is directed to the second UI element, decreasing a first attention accumulator value associated with the first UI element and increasing a second attention accumulator value associated with the second UI element based on a length of time that the first gaze direction is directed to the second UI element; in accordance with a determination that the second attention accumulator value associated with the second UI element exceeds the first attention accumulator value associated with the first UI element, deselecting the first UI element and selecting the second UI element; and in accordance with a determination that the second attention accumulator value associated with the second UI element does not exceed the first attention accumulator value associated with the first UI element, maintaining selection of the first UI element.
[0017] According to some implementations, an electronic device includes one or more displays, one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing performing any of the methods described herein. According to some implementations, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors of a device, cause the device to perform or cause performing any of the methods described herein. According to some implementations, a device includes one or more displays, one or more processors, non-transitory memory, and means for performing or causing performing any of the methods described herein.
[0018] According to some implementations, a computing system includes one or more processors, non-transitory memory, an interface to communicate with a display device and one or more input devices, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions to perform or cause performance of the operations of any of the methods described herein. According to some implementations, a non-transitory computer-readable storage medium has stored therein instructions that, when executed by one or more processors of a computing system having an interface to communicate with a display device and one or more input devices, cause the computing system to perform or cause performance of the operations of any of the methods described herein. According to some implementations, a computing system includes one or more processors, non-transitory memory, an interface to communicate with a display device and one or more input devices, and means for performing or causing performance of the operations of any of the methods described herein. DETAILED DESCRIPTION
[0019] Many details are described to provide a thorough understanding of the example implementations shown in the illustrations. However, the illustrations merely illustrate some example aspects of the present disclosure and should not be considered to be limiting. One of ordinary skill in the art would understand that other effective aspects and / or variants not including all of the specific details described herein are possible. Moreover, well-known systems, methods, components, devices and circuits have not been described in exhaustive detail so as to not obscure the more pertinent aspects of the example implementations described herein.
[0020] A physical environment refers to the physical world that people are able to sense and / or interact with without the aid of electronic devices. A physical environment can include physical features such as physical surfaces or physical objects. For example, a physical environment corresponds to a physical park that includes physical trees, physical buildings, and physical people. People are able to directly sense and / or interact with a physical environment such as through sight, touch, hearing, taste, and smell. In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic devices. For example, an XR environment can include augmented reality (AR) content, mixed reality (MR) content, virtual reality (VR) content, etc. In the case of an XR system, a subset of a person’s physical motions, or representations thereof, are tracked, and in response, one or more characteristics of one or more virtual objects simulated in the XR system are adjusted in a manner consistent with at least one physical law. For example, an XR system can detect head movements and, in response, adjust graphical content and a sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. As another example, an XR system can detect movements of an electronic device (e.g., a mobile phone, a tablet, a laptop, etc.) that presents an XR environment and, in response, adjust graphical content and a sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. In some cases (e.g., for accessibility reasons), an XR system can adjust characteristics of graphical content in an XR environment in response to representations of physical motions (e.g., voice commands).
[0021] There are many different types of electronic systems that enable a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields having integrated display capability, windows having integrated display capability, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system can have an integrated opaque display and one or more speakers. Alternatively, a head-mounted system can be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system can incorporate one or more imaging sensors to capture images or video of the physical environment, and / or one or more microphones to capture audio of the physical environment. Rather than an opaque display, a head-mounted system can have a transparent or translucent display. The transparent or translucent display can have a medium through which light representative of images is directed to a person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, hologram medium, optical combiner, optical reflector, or any combination thereof. In some implementations, the transparent or translucent display can be configured to selectively become opaque. Projection-based systems can employ retinal projection technology that projects graphical images onto a person's retinas. Projection systems can also be configured to project virtual objects into the physical environment, for example, as a hologram or on a physical surface.
[0022] Figure 1 is a block diagram of an example operational architecture 100 in accordance with some implementations. While pertinent part is shown, those of ordinary skill in the art will appreciate from the disclosure herein that various other features can be included, as desired, in accordance with the example implementations disclosed herein. As such, operational architecture 100 includes, by way of non-limiting example, an optional controller 110 and an electronic device 120 (e.g., a tablet, a mobile phone, a laptop, a near-eye system, a wearable computing device, etc.), in accordance with some implementations.
[0023] In some implementations, controller 110 is configured to manage and coordinate XR experiences (sometimes also referred to herein as "XR environments" or "virtual environments" or "graphical environments") of user 150 and, optionally, other users. In some implementations, controller 110 includes a suitable combination of software, firmware, and / or hardware. Reference is made to Figure 2Controller 110 is described in greater detail. In some implementations, controller 110 is a computing device located locally or remotely with respect to physical environment 105. For example, controller 110 is a local server located within physical environment 105. In another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside of physical environment 105. In some implementations, controller 110 is communicatively coupled with electronic device 120 via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802. l lx, IEEE 802.16x, IEEE 802.3x, etc.). In some implementations, the functionality of controller 110 is provided by electronic device 120. As such, in some implementations, the components of controller 110 are integrated into electronic device 120.
[0024] In some implementations, electronic device 120 is configured to present audio and / or video (A / V) content to user 150. In some implementations, electronic device 120 is configured to present a user interface (UI) and / or XR environment 128 to user 150. In some implementations, electronic device 120 includes a suitable combination of software, firmware, and / or hardware. Reference is made to FIG. 2 for a description of the components that make up electronic device 120. Figure 3 Electronic device 120 is described in greater detail.
[0025] According to some implementations, when user 150 is physically present within physical environment 105, which includes table 107 and figure 523 within a field of view (FOV) 111 of electronic device 120, electronic device 120 presents an XR experience to user 150. As such, in some implementations, user 150 holds electronic device 120 in one or both of his / her hands. In some implementations, while presenting the XR experience, electronic device 120 is configured to present XR content (also sometimes referred to herein as “graphical content” or “virtual content”), including XR cylinder 109, and enable video pass-through of physical environment 105 (e.g., including table 107 and figure 523 (or a representation thereof)) on display 122. For example, XR environment 128, including XR cylinder 109, is stereoscopic or three-dimensional (3D).
[0026] In one example, the XR cylinder 109 corresponds to head / display-locked content such that the XR cylinder 109 remains displayed at the same location on the display 122 as the FOV 111 changes due to translational and / or rotational movement of the electronic device 120. As another example, the XR cylinder 109 corresponds to world / object-locked content such that the XR cylinder 109 remains displayed at its original location as the FOV 111 changes due to translational and / or rotational movement of the electronic device 120. Thus, in this example, the XR environment 128 displayed will not include the XR cylinder 109 if the FOV 111 does not include the original location. As another example, the XR cylinder 109 corresponds to body-locked content such that it remains at a positional and rotational offset from the body of the user 150. In some examples, the electronic device 120 corresponds to a near-eye system, a mobile phone, a tablet, a laptop, a wearable computing device, etc.
[0027] In some implementations, the display 122 corresponds to an additive display that enables optical pass-through of the physical environment 105 (including the table 107 and the person 523). For example, the display 122 corresponds to a transparent lens and the electronic device 120 corresponds to a pair of glasses worn by the user 150. Thus, in some implementations, the electronic device 120 presents the user interface by projecting XR content (e.g., the XR cylinder 109) onto the additive display, which in turn is superimposed on the physical environment 105 from the perspective of the user 150. In some implementations, the electronic device 120 presents the user interface by displaying XR content (e.g., the XR cylinder 109) on the additive display, which in turn is superimposed on the physical environment 105 from the perspective of the user 150.
[0028] In some implementations, the user 150 wears the electronic device 120, such as a near-eye system. Thus, the electronic device 120 includes one or more displays (e.g., a single display or a display for each eye) that are provided to display XR content. For example, the electronic device 120 encloses the FOV of the user 150. In such implementations, the electronic device 120 presents the XR environment 128 by displaying data corresponding to the XR environment 128 on the one or more displays or by projecting data corresponding to the XR environment 128 onto the retina of the user 150.
[0029] In some implementations, the electronic device 120 includes an integrated display (e.g., a built-in display) that displays the XR environment 128. In some implementations, the electronic device 120 includes a head-mountable housing. In various implementations, the head-mountable housing includes an attachment region to which another device having a display can be attached. For example, in some implementations, the electronic device 120 can be attached to a head-mountable housing. In various implementations, the head-mountable housing is shaped to form a receptacle for receiving another device (e.g., the electronic device 120) that includes a display. For example, in some implementations, the electronic device 120 is slid / snapped into or otherwise attached to a head-mountable housing. In some implementations, the display of the device attached to the head-mountable housing presents (e.g., displays) the XR environment 128. In some implementations, the electronic device 120 is replaced with an XR room, housing, or chamber configured to present XR content in which the user 150 does not wear the electronic device 120.
[0030] In some implementations, the controller 110 and / or the electronic device 120 causes the XR representation of the user 150 to move within the XR environment 128 based on movement information (e.g., body pose data, eye tracking data, hand / limb / finger / endpoint tracking data, etc.) from the electronic device 120 and / or optional remote input devices within the physical environment 105. In some implementations, the optional remote input devices correspond to fixed or movable sensory devices (e.g., image sensors, depth sensors, infrared (IR) sensors, event cameras, microphones, etc.) within the physical environment 105. In some implementations, each remote input device is configured to collect / capture input data and provide the input data to the controller 110 and / or the electronic device 120 while the user 150 is physically within the physical environment 105. In some implementations, the remote input devices include microphones and the input data includes audio data (e.g., speech samples) associated with the user 150. In some implementations, the remote input devices include image sensors (e.g., cameras) and the input data includes images of the user 150. In some implementations, the input data characterizes body poses of the user 150 at different times. In some implementations, the input data characterizes head poses of the user 150 at different times. In some implementations, the input data characterizes hand tracking information associated with the hands of the user 150 at different times. In some implementations, the input data characterizes velocity and / or acceleration of body parts of the user 150, such as his / her hands. In some implementations, the input data indicates joint positions and / or joint orientations of the user 150. In some implementations, the remote input devices include feedback devices, such as speakers, lights, etc.
[0031] Figure 2is a block diagram of an example of a controller 110 according to some implementations. While certain specific features are illustrated, one of ordinary skill in the art will appreciate from the disclosure herein that various other features have not been illustrated for the sake of brevity and so as not to obscure more pertinent aspects of the implementations disclosed herein. To that end, as a non-limiting example, in some implementations the controller 110 includes one or more processing units 202 (e.g., microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, and / or the like), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., universal serial bus (USB), IEEE 802.3x, IEEE 802.1 lx, IEEE 802.16x, global system for mobile communications (GSM), code division multiple access (CDMA), time division multiple access (TDMA), global positioning system (GPS), infrared (IR), Bluetooth, ZIGBEE, and / or the like), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.
[0032] In some implementations, the one or more communication buses 204 include circuitry that interconnects a system’s components and / or controls communication among system components. In some implementations, the one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, touchscreen, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0033] The memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid state memory devices. In some implementations, the memory 220 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. The memory 220 optionally includes one or more storage devices remotely located from the one or more processing units 202. The memory 220 comprises a non-transitory computer readable storage medium. In some implementations, the memory 220, or the non-transitory computer readable storage medium of the memory 220, stores the following described programs, modules, and data structures, or a subset thereof. Figure 2 The memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid state memory devices. In some implementations, the memory 220 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. The memory 220 optionally includes one or more storage devices remotely located from the one or more processing units 202. The memory 220 comprises a non-transitory computer readable storage medium. In some implementations, the memory 220, or the non-transitory computer readable storage medium of the memory 220, stores the following described programs, modules, and data structures, or a subset thereof.
[0034] The operating system 230 includes procedures for handling various basic system services and for performing hardware dependent tasks.
[0035] In some implementations, the data fetcher 242 is configured to fetch data (e.g., captured image frames of the physical environment 105, presentation data, input data, user interaction data, camera pose tracking information, eye tracking information, head / body pose tracking information, hand / limb / finger / limb tracking information, sensor data, position data, etc.) from at least one of the I / O devices 206 of the controller 110, the I / O devices and sensors 306 of the electronic device 120, and the optional remote input devices. To this end, in various implementations, the data fetcher 242 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0036] In some implementations, the mapper and localizer engine 244 is configured to map the physical environment 105 and to track at least a localization / position of the electronic device 120 or the user 150 relative to the physical environment 105. To this end, in various implementations, the mapper and localizer engine 244 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0037] In some implementations, the data transmitter 246 is configured to transmit data (e.g., presentation data such as rendered image frames associated with an XR environment, position data, etc.) to at least the electronic device 120 and optionally one or more other devices. To this end, in various implementations, the data transmitter 246 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0038] In some implementations, the privacy architecture 408 is configured to ingest data and filter user information and / or identify information within the data based on one or more privacy filters. Reference is made below to FIGS. 5A-5C for a more detailed description of the privacy architecture 408. Figure 4A The privacy architecture 408 is described in more detail. To this end, in various implementations, the privacy architecture 408 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0039] In some implementations, the motion state estimator 410 is configured to fetch (e.g., receive, retrieve, or determine / generate) a motion state vector 411 associated with the electronic device 120 (and the user 150) based on the input data (e.g., including a current motion state associated with the electronic device 120) and to update the motion state vector 411 over time. For example, as Figure 4BAs shown, the motion state vector 411 includes motion state descriptors 472 for the electronic device 120 (e.g., stationary, in motion, walking, running, cycling, operating or riding in a car, operating or riding in a boat, operating or riding in a bus, operating or riding in a train, operating or riding in an airplane, etc.), translational movement values 474 associated with the electronic device 120 (e.g., heading, velocity value, acceleration value, etc.), angular movement values 476 associated with the electronic device 120 (e.g., angular velocity value, angular acceleration value, etc. for each of the pitch, roll, and yaw dimensions), etc. See below for reference. Figure 4A The motion state estimator 410 is described in more detail. To this end, in various specific implementations, the motion state estimator 410 includes instructions and / or logic components for the instructions, as well as heuristics and metadata for the heuristics.
[0040] In some specific implementations, the eye-tracking engine 412 is configured to acquire (e.g., receive, retrieve, or determine / generate) based on input data, such as... Figure 4B The eye-tracking vector 413 shown is (e.g., having a gaze direction) and is updated over time. For example, the gaze direction indicates a point in the physical environment 105 that user 150 is currently viewing (e.g., associated with x, y, and z coordinates relative to physical environment 105 or the world), a physical object, or a region of interest (ROI). As another example, the gaze direction indicates a point in the XR environment 128 that user 150 is currently viewing (e.g., associated with x, y, and z coordinates relative to XR environment 128), an XR object, or an ROI. See below for further details. Figure 4A The eye-tracking engine 412 is described in more detail. For this purpose, in various specific implementations, the eye-tracking engine 412 includes instructions and / or logic components for those instructions, as well as heuristics and metadata for those heuristics.
[0041] In some specific implementations, the body / head pose tracking engine 414 is configured to acquire (e.g., receive, retrieve, or determine / generate) a pose representation vector 415 based on input data and update the pose representation vector 415 over time. For example, as Figure 4B As shown, the posture representation vector 415 includes a head posture descriptor 492A (e.g., up, down, neutral, etc.), head posture translation values 492B, head posture rotation values 492C, body posture descriptor 494A (e.g., standing, sitting, prone, etc.), body part / limb / limb / joint translation values 494B, body part / limb / limb / joint rotation values 494C, etc. See below for reference. Figure 4AThe body / head pose tracking engine 414 is described in greater detail. To this end, in various implementations, the body / head pose tracking engine 414 includes instructions and / or logic therefor, as well as heuristics and metadata therefor. In some implementations, the motion state estimator 410, the eye tracking engine 412, and the body / head pose tracking engine 414 can be located on the electronic device 120 in addition to or instead of the controller 110.
[0042] In some implementations, the content selector 422 is configured to select XR content (sometimes also referred to as “graphic content” or “virtual content”) from the content library 425 based on one or more user requests and / or inputs (e.g., voice commands, selections from a user interface (UI) menu of an XR content item or virtual agent (VA), etc.). Reference is made below to Figure 4A The content selector 422 is described in greater detail. To this end, in various implementations, the content selector 422 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0043] In some implementations, the content library 425 includes a plurality of content items, such as audio / visual (A / V) content, virtual agents (VAs), and / or XR content, objects, items, scenes, etc. As one example, the XR content includes 3D reconstructions of user-captured videos, movies, TV episodes, and / or other XR content. In some implementations, the content library 425 is pre-populated or hand-curated by the user 150. In some implementations, the content library 425 is located locally with respect to the controller 110. In some implementations, the content library 425 is located remotely from the controller 110 (e.g., at a remote server, a cloud server, etc.).
[0044] In some implementations, the characterization engine 442 is configured to determine / generate the characterization vector 443 based on at least one of the motion state vector 411, the eye tracking vector 413, and the pose characterization vector 415 as Figure 4A shown. In some implementations, the characterization engine 442 is further configured to update the pose characterization vector 443 over time. As Figure 4B shown, the characterization vector 443 includes motion state information 4102, gaze direction information 4104, head pose information 4106A, body pose information 4106B, limb tracking information 4106C, position information 4108, etc. Reference is made below to Figure 4A The characterization engine 442 is described in greater detail. To this end, in various implementations, the characterization engine 442 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0045] In some implementations, the content manager 430 is configured to manage and update the layout, settings, structure, etc. of the XR environment 128, including one or more of the VA, XR content, one or more user interface (UI) elements associated with the XR content, etc. Reference is made below to FIGS. 4A-4C for further details. Figure 4C The content manager 430 is described in further detail. To this end, in various implementations, the content manager 430 includes instructions and / or logic therefor, as well as heuristics and metadata therefor. In some implementations, the content manager 430 includes a frame buffer 434, a content updater 436, and a feedback engine 438. In some implementations, the frame buffer 434 includes XR content for one or more past instances and / or frames, rendered image frames, etc.
[0046] In some implementations, the content updater 436 is configured to modify the XR environment 105 over time based on translational or rotational movement of the electronic device 120 or a physical object within the physical environment 128, user input (e.g., changes in context, hand / limb tracking input, eye tracking input, touch input, voice commands, modification / manipulation input to a physical object, etc.), etc. To this end, in various implementations, the content updater 436 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0047] In some implementations, the feedback engine 438 is configured to generate sensory feedback (e.g., visual feedback such as text or lighting changes, audio feedback, haptic feedback, etc.) associated with the XR environment 128. To this end, in various implementations, the feedback engine 438 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0048] In some implementations, the rendering engine 450 is configured to render the XR environment 128 (sometimes also referred to as a “graphics environment” or “virtual environment”) or image frames associated therewith, as well as the VA, XR content, one or more UI elements associated with the XR content, etc. To this end, in various implementations, the rendering engine 450 includes instructions and / or logic therefor, as well as heuristics and metadata therefor. In some implementations, the rendering engine 450 includes a pose determiner 452, a renderer 454, an optional image processing architecture 462, and an optional compositor 464. Those of ordinary skill in the art will appreciate that for a video pass-through configuration, there can be the optional image processing architecture 462 and the optional compositor 464, but for a full VR or optical pass-through configuration, the optional image processing architecture and the optional compositor can be removed.
[0049] In some implementations, the pose determiner 452 is configured to determine a current camera pose of the electronic device 120 and / or the user 150 relative to the A / V content and / or the XR content. Reference is made below. Figure 4A The pose determiner 452 is described in more detail. To this end, in various implementations, the pose determiner 452 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0050] In some implementations, the renderer 454 is configured to render the A / V content and / or the XR content in accordance with the current camera pose associated therewith. Reference is made below. Figure 4A The renderer 454 is described in more detail. To this end, in various implementations, the renderer 454 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0051] In some implementations, the image processing architecture 462 is configured to obtain (e.g., receive, retrieve, or capture) an image stream comprising one or more images of the physical environment 105 from the current camera pose of the electronic device 120 and / or the user 150. In some implementations, the image processing architecture 462 is further configured to perform one or more image processing operations on the image stream, such as warping, color correction, gamma correction, sharpening, noise reduction, white balancing, and the like. Reference is made below. Figure 4A The image processing architecture 462 is described in more detail. To this end, in various implementations, the image processing architecture 462 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0052] In some implementations, the compositor 464 is configured to composite the rendered A / V content and / or the XR content with the processed image stream of the physical environment 105 from the image processing architecture 462 to produce a rendered image frame of the XR environment 128 for display. Reference is made below. Figure 4A The compositor 464 is described in more detail. To this end, in various implementations, the compositor 464 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0053] Although the data acquirer 242, mapper and localizer engine 244, data transmitter 246, privacy architecture 408, motion state estimator 410, eye tracking engine 412, body / head pose tracking engine 414, content selector 422, content manager 430, operational modality manager 440, and rendering engine 450 are shown as residing on a single device (e.g., the controller 110), it will be appreciated that, in other implementations, any combination of the data acquirer 242, mapper and localizer engine 244, data transmitter 246, privacy architecture 408, motion state estimator 410, eye tracking engine 412, body / head pose tracking engine 414, content selector 422, content manager 430, operational modality manager 440, and rendering engine 450 can be located in separate computing devices.
[0054] In some implementations, the functionality and / or components of the controller 110 are combined with or provided by the electronic device 120 shown below in FIG. 4. Figure 3 In addition, Figure 2 More generally, the functions of the various features that can be present in particular implementations are described more fully below in the context of the electronic device 120 shown in FIG. 4, and the functions of the electronic device 120 are described more fully below in the context of the controller 110 shown in FIG. 3. As will be appreciated by those of ordinary skill in the art, the items shown separately can be combined, and some items can be separated. For example, some of the functional modules shown separately in FIG. 3 and FIG. 4 can be implemented in a single module, and various functions of a single functional block can be implemented by one or more functional blocks in various implementations. The actual number of modules and the division of particular functions between them, and how features are allocated among them, will vary from implementation to implementation, and in some implementations, depend in part on the particular combination of hardware, software, and / or firmware chosen for a particular implementation. Figure 2
[0055] Figure 3 is a block diagram of an example of some particular implementations of an electronic device 120 (e.g., a mobile phone, a tablet computer, a laptop computer, a near-eye system, a wearable computing device, etc.). While some specific features are illustrated, one of skill in the art will appreciate from the disclosure herein that various other features have not been illustrated for the sake of brevity and so as not to obscure more pertinent aspects of the implementations disclosed herein. For instance, in some implementations, the electronic device 120 includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BLUETOOTH, ZIGBEE, and / or the like types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more displays 312, image capture devices 370 (one or more optional inward- and / or outward-facing image sensors), memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0056] In some implementations, the one or more communication buses 304 include circuitry that interconnects and controls communications between system components. In some implementations, the one or more I / O devices and sensors 306 include at least one of an inertial measurement unit (IMU), an accelerometer, a gyroscope, a magnetometer, a thermometer, one or more physiological sensors (e.g., blood pressure monitor, heart rate monitor, blood oxygen saturation monitor, blood glucose monitor, etc.), one or more microphones, one or more speakers, a haptics engine, a heating and / or cooling unit, a skin shear engine, one or more depth sensors (e.g., structured light, time-of-flight, LiDAR, etc.), a positioning and mapping engine, an eye tracking engine, a body / head pose tracking engine, a hand / limb / finger / limb tracking engine, a camera pose tracking engine, etc.
[0057] In some implementations, the one or more displays 312 are configured to present an XR environment to a user. In some implementations, the one or more displays 312 are also configured to present planar video content to a user (e.g., two-dimensional or “flat” AVI, FLV, WMV, MOV, MP4, etc. files associated with a television show or movie, or live video pass-through of the physical environment 105). In some implementations, the one or more displays 312 correspond to a touchscreen display. In some implementations, the one or more displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transitory (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), micro-electro-mechanical system (MEMS), and / or similar display types. In some implementations, the one or more displays 312 correspond to diffractive, reflective, polarized, holographic, etc. waveguide displays. For example, the electronic device 120 includes a single display. As another example, the electronic device 120 includes a display for each eye of a user. In some implementations, the one or more displays 312 are capable of presenting AR and VR content. In some implementations, the one or more displays 312 are capable of presenting AR or VR content.
[0058] In some implementations, the image capture devices 370 correspond to one or more RGB cameras (e.g., with a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), IR image sensors, event-based cameras, etc. In some implementations, the image capture devices 370 include a lens assembly, photodiodes, and a front-end architecture. In some implementations, the image capture devices 370 include outward-facing and / or inward-facing image sensors.
[0059] The memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM or other random access solid state memory devices. In some implementations, the memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. The memory 320 optionally includes one or more storage devices remotely located from the one or more processing units 302. The memory 320 comprises a non-transitory computer readable storage medium. In some implementations, the memory 320, or the non-transitory computer readable storage medium of the memory 320, stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 330 and a presentation engine 340.
[0060] The operating system 330 includes procedures for handling various basic system services and for performing hardware dependent tasks. In some implementations, the presentation engine 340 is configured to present media items and / or XR content to a user via the one or more displays 312. To this end, in various implementations, the presentation engine 340 includes a data acquirer 342, an interaction handler 420, a presenter 470, and a data transmitter 350.
[0061] In some implementations, the data acquirer 342 is configured to acquire data (e.g., presentation data, such as rendered image frames associated with a user interface or XR environment, input data, user interaction data, head tracking information, camera pose tracking information, eye tracking information, hand / limb / finger / limb tracking information, sensor data, position data, etc.) from at least one of the I / O devices and sensors 306 of the electronic device 120, the controller 110, and a remote input device. To this end, in various implementations, the data acquirer 342 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0062] In some implementations, the interaction handler 420 is configured to detect user interactions with presented A / V content and / or XR content (e.g., gesture inputs detected via hand / limb tracking, eye gaze inputs detected via eye tracking, voice commands, etc.). To this end, in various implementations, the interaction handler 420 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0063] In some implementations, the presenter 470 is configured to present and update A / V content and / or XR content (e.g., rendered image frames associated with a user interface or XR environment 128, including VA, XR content, one or more UI elements associated with the XR content, etc.) via the one or more displays 312. To this end, in various implementations, the presenter 470 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0064] In some implementations, the data transmitter 350 is configured to transmit data (e.g., presentation data, position data, user interaction data, head tracking information, camera pose tracking information, eye tracking information, hand / limb / finger / limb tracking information, etc.) to at least the controller 110. To this end, in various implementations, the data transmitter 350 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0065] Although the data obtainer 342, the interaction handler 420, the presenter 470, and the data transmitter 350 are shown as residing on a single device (e.g., the electronic device 120), it will be appreciated that, in other implementations, any combination of the data obtainer 342, the interaction handler 420, the presenter 470, and the data transmitter 350 can be located in separate computing devices.
[0066] Furthermore, Figure 3 More functionally describe various features that can be present in particular implementations, rather than structural schematics of the implementations described herein. As will be appreciated by one of ordinary skill in the art, items shown separately can be combined, and items shown separately can be divided. For example, Figure 3 Some of the functional modules shown separately in can be implemented in a single module, and various functions of a single functional block can be implemented by one or more functional blocks in various implementations. The actual number of modules and the division of particular functions between them, as well as how the features are allocated among them, will vary from one implementation to another and, in some implementations, depend partly on the particular combination of hardware, software, and / or firmware chosen to implement the particular implementation.
[0067] Figure 4A is a block diagram of a first portion 400A of an example content delivery architecture in accordance with some implementations. While relevant features are shown, one of ordinary skill in the art will appreciate from the disclosure herein that various other features have not been shown for brevity and so as to not obscure more relevant aspects of the example implementations disclosed herein. To that end, the content delivery architecture is included in a computing system, such as Figure 1 and Figure 2 the controller 110 shown in Figure 1 and Figure 3 the electronic device 120 shown in; and / or suitable combinations thereof.
[0068] As Figure 4A shown, one or more local sensors 402 of the controller 110, the electronic device 120, and / or combinations thereof obtain local sensor data 403 associated with the physical environment 105. For example, the local sensor data 403 includes images or streams thereof of the physical environment 105, simultaneous localization and mapping (SLAM) information of the physical environment 105, and a location of the electronic device 120 or the user 150 relative to the physical environment 105, environmental lighting information of the physical environment 105, environmental audio information of the physical environment 105, acoustic information of the physical environment 105, dimensional information of the physical environment 105, semantic labels of objects within the physical environment 105, and the like. In some implementations, the local sensor data 403 includes unprocessed or post-processed information.
[0069] Similarly, asFigure 4A The one or more remote sensors 404 associated with optional remote input devices within the physical environment 105 acquire remote sensor data 405 associated with the physical environment 105, as shown. For example, the remote sensor data 405 includes images or streams thereof of the physical environment 105, SLAM information of the physical environment 105, and a location of the electronic device 120 or the user 150 relative to the physical environment 105, ambient lighting information of the physical environment 105, ambient audio information of the physical environment 105, acoustic information of the physical environment 105, dimensional information of the physical environment 105, semantic labels of objects within the physical environment 105, and the like. In some implementations, the remote sensor data 405 includes unprocessed or post-processed information.
[0070] According to some implementations, the privacy architecture 408 ingests the local sensor data 403 and the remote sensor data 405. In some implementations, the privacy architecture 408 includes one or more privacy filters associated with user information and / or identifying information. In some implementations, the privacy architecture 408 includes an opt-in feature in which the electronic device 120 notifies the user 150 of which user information and / or identifying information is being monitored and how the user information and / or identifying information will be used. In some implementations, the privacy architecture 408 selectively prevents and / or limits the content delivery architecture 400A / 400B or portions thereof from acquiring and / or transmitting user information. To do so, the privacy architecture 408 receives user preferences and / or selections from the user 150 in response to prompting the user 150 for the user preferences and / or selections. In some implementations, the privacy architecture 408 prevents the content delivery architecture 400A / 400B from acquiring and / or transmitting user information unless and until the privacy architecture 408 acquires informed consent from the user 150. In some implementations, the privacy architecture 408 anonymizes (e.g., scrambles, obfuscates, encrypts, etc.) certain types of user information. For example, the privacy architecture 408 receives user input specifying which types of user information the privacy architecture 408 is to anonymize. As another example, the privacy architecture 408 independently of user specification (e.g., automatically) anonymizes certain types of user information that can include sensitive and / or identifying information.
[0071] According to some implementations, the motion state estimator 410 acquires the local sensor data 403 and the remote sensor data 405 after being subjected to the privacy architecture 408. In some implementations, the motion state estimator 410 acquires (e.g., receives, retrieves, or determines / generates) and updates over time a motion state vector 411 based on the input data.
[0072] Figure 4B An example data structure of the motion state vector 411 is shown according to some implementations. As shown, the motion state vector 411 includes a position vector 412, a velocity vector 414, and an orientation vector 416. Figure 4BAs shown, the motion state vector 411 may correspond to an N-tuple representation vector or representation tensor, which includes a timestamp 471 (e.g., the most recent time the motion state vector 411 was updated), a motion state descriptor 472 for the electronic device 120 (e.g., stationary, in motion, car, ship, bus, train, airplane, etc.), translational movement values 474 associated with the electronic device 120 (e.g., heading, displacement, velocity, acceleration, jerk, etc.), angular movement values 476 associated with the electronic device 120 (e.g., angular displacement, angular velocity, angular acceleration, angular jerk, etc. for each of the pitch, roll, and yaw dimensions), and / or miscellaneous information 478. Those skilled in the art will understand that... Figure 4B The data structure of motion state vector 411 in the example is only an example, and it may include different information parts and be constructed in various other implementations in various other implementations.
[0073] In some implementations, eye-tracking engine 412 acquires local sensor data 403 and remote sensor data 405 after undergoing privacy architecture 408. In some implementations, eye-tracking engine 412 acquires (e.g., receives, retrieves, or determines / generates) eye-tracking vector 413 based on input data and updates eye-tracking vector 413 over time.
[0074] Figure 4B An exemplary data structure for eye-tracking vector 413 according to some specific implementation is shown. For example... Figure 4B As shown, the eye-tracking vector 413 may correspond to an N-tuple representation vector or representation tensor, which includes a timestamp 481 (e.g., the most recent time the eye-tracking vector 413 was updated), one or more angle values 482 of the current gaze direction (e.g., tilt, pitch, and yaw values), one or more translation values 484 of the current gaze direction (e.g., x, y, and z values relative to the physical environment 105, the entire world, etc.), and / or miscellaneous information 486. Those skilled in the art will understand that... Figure 4B The data structure of the eye tracking vector 413 in the example is only an example, and it may include different information parts and be constructed in various other implementations in various other implementations.
[0075] For example, the gaze direction indicates a point in the physical environment 105 that user 150 is currently viewing (e.g., associated with x, y, and z coordinates relative to physical environment 105 or the world), a physical object, or a region of interest (ROI). As another example, the gaze direction indicates a point in the XR environment 128 that user 150 is currently viewing (e.g., associated with x, y, and z coordinates relative to XR environment 128), an XR object, or a region of interest (ROI).
[0076] In some implementations, the body / head pose tracking engine 414 acquires local sensor data 403 and remote sensor data 405 after undergoing a privacy architecture 408. In some implementations, the body / head pose tracking engine 414 acquires (e.g., receives, retrieves, or determines / generates) a pose representation vector 415 based on the input data and updates the pose representation vector 415 over time.
[0077] Figure 4B An exemplary data structure for the pose representation vector 415 according to some specific implementation is shown. For example... Figure 4B As shown, the pose representation vector 415 may correspond to an N-tuple representation vector or representation tensor, which includes a timestamp 491 (e.g., the most recent update time of the pose representation vector 415), a head pose descriptor 492A (e.g., up, down, neutral, etc.), a translation value 492B for the head pose, a rotation value 492C for the head pose, a body pose descriptor 494A (e.g., standing, sitting, prone, etc.), a translation value 494B for body parts / limbs / joints, a rotation value 494C for body parts / limbs / joints, and / or miscellaneous information 496. In some specific implementations, the pose representation vector 415 may also include information associated with finger / hand / limb tracking. Those skilled in the art will understand that... Figure 4B The data structure of the pose representation vector 415 in the example is merely illustrative; it may include different information components and be constructed in various ways in various other implementations. According to some implementations, the motion state vector 411, the eye tracking vector 413, and the pose representation vector 415 are collectively referred to as the input vector 419.
[0078] According to some specific implementations, the representation engine 442 acquires the motion state vector 411, the eye tracking vector 413, and the pose representation vector 415. In some specific implementations, the representation engine 442 acquires (e.g., receives, retrieves, or determines / generates) the representation vector 443 based on the motion state vector 411, the eye tracking vector 413, and the pose representation vector 415.
[0079] Figure 4B An exemplary data structure for representation vector 443 according to some specific implementation is shown. For example... Figure 4CAs shown, the representation vector 443 can correspond to an N-tuple representation vector or representation tensor that includes a timestamp 4101 (e.g., a most recent time at which the representation vector 443 was updated), motion state information 4102 (e.g., a motion state descriptor 472), gaze direction information 4104 (e.g., a function of one or more angle values 482 and one or more translation values 484 within the eye tracking vector 413), head pose information 4106A (e.g., a head pose descriptor 492A), body pose information 4106B (e.g., a function of the body pose descriptor 494A within the body pose representation vector 415), limb tracking information 4106C (e.g., a function of the body pose descriptor 494A within the body pose representation vector 415 associated with a limb of the user 150 tracked by the controller 110, the electronic device 120, and / or a combination thereof), location information 4108 (e.g., a home location (such as a kitchen or living room), a vehicle location (such as a car, an airplane, etc.), and / or the like), and / or miscellaneous information 4109.
[0080] Figure 1 is a block diagram of a second portion 400B of an example content delivery architecture in accordance with some implementations. While pertinent part is shown, those of ordinary skill in the art will appreciate from the disclosure herein that various other features can also be included, as for example, in a computing system such as Figure 2 and Figure 1 the controller 110 as shown; Figure 3 and Figure 4C the electronic device 120 as shown; and / or suitable combinations thereof. Figure 4A Similar to and adapted from Figure 4A Thus, Figure 4C and Figure 4A use similar reference numbers. Accordingly, for the sake of brevity only the differences between Figure 4C and Figures 5A to 5H will be described below.
[0081] According to some implementations, the interaction handler 420 obtains (e.g., receives, retrieves, or detects) one or more user inputs 421 provided by the user 150 that are associated with selecting A / V content, one or more VA and / or XR content for presentation. For example, the one or more user inputs 421 correspond to gesture inputs that select XR content from a UI menu detected via hand / limb tracking, eye gaze inputs that select XR content from a UI menu detected via eye tracking, voice commands that select XR content from a UI menu detected via a microphone, and the like. In some implementations, the content selector 422 selects XR content 427 from the content library 425 based on the one or more user inputs 421 (e.g., voice commands, selections from a menu of XR content items, and the like).
[0082] In various implementations, the content manager 430 manages and updates the layout, settings, structure, and the like of the XR environment 128 based on the characterization vector 443, (optionally) the user inputs 421, and the like, which includes one or more of VA, XR content, one or more UI elements associated with the XR content, and the like. To this end, the content manager 430 includes a frame buffer 434, a content updater 436, and a feedback engine 438.
[0083] In some implementations, the frame buffer 434 includes XR content, rendered image frames, and the like for one or more past instances and / or frames. In some implementations, the content updater 436 modifies the XR environment 128 over time based on the characterization vector 443, user inputs 421 associated with modifying and / or manipulating XR content or VA, translational or rotational movement of objects within the physical environment 105, translational or rotational movement of the electronic device 120 (or the user 150), and the like. In some implementations, the feedback engine 438 generates sensory feedback (e.g., visual feedback such as text or lighting changes, audio feedback, haptic feedback, and the like) associated with the XR environment 128.
[0084] According to some implementations, the pose determiner 452 determines a current camera pose of the electronic device 120 and / or the user 150 relative to the XR environment 128 and / or the physical environment 105 based at least in part on the pose characterization vector 415. In some implementations, the renderer 454 renders the VA, the XR content 427, one or more UI elements associated with the XR content, and the like in accordance with the current camera pose relative thereto.
[0085] According to some implementations, the optional image processing architecture 462 obtains an image stream from the image capture device 370, which includes one or more images of the physical environment 105 from the current camera pose of the electronic device 120 and / or the user 150. In some implementations, the image processing architecture 462 also performs one or more image processing operations on the image stream, such as warping, color correction, gamma correction, sharpening, noise reduction, white balancing, and the like. In some implementations, the optional compositor 464 composites the rendered XR content with the processed image stream of the physical environment 105 from the image processing architecture 462 to produce a rendered image frame of the XR environment 128. In various implementations, the Tenderer 470 presents the rendered image frame of the XR environment 128 to the user 150 via the one or more displays 312. Those of ordinary skill in the art will understand that the optional image processing architecture 462 and the optional compositor 464 can not be applicable to fully virtual environments (or optical pass-through scenarios).
[0086] Figure 1 A sequence of examples 500-570 of content delivery scenarios is shown in accordance with some implementations. While some specific features are shown, those of ordinary skill in the art will appreciate from the present disclosure that various other features have not been shown in the interest of brevity and so as not to obscure more relevant aspects of the implementations disclosed herein. To that end, by way of non-limiting example, the sequence of examples 500-570 is rendered and presented by a computing system such as Figure 2 and Figure 1 the controller 110 shown; Figure 3 and Figures 5A to 5H the electronic device 120 shown; and / or suitable combinations thereof.
[0087] As shown in Figure 1 , the content delivery scenarios include a physical environment 105 and an XR environment 128 displayed on a display 122 of an electronic device 120 (e.g., associated with a user 150). While the user 150 is physically present within the physical environment 105, the electronic device 120 presents the XR environment 128 to the user 150, which physical environment includes a door 115 that is currently located within the FOV 111 of an outward-facing image sensor of the electronic device 120. Thus, in some implementations, the user 150 holds the electronic device 120 in their hand, similar to the operating environment 100 in Figure 5A .
[0088] In other words, in some implementations, the electronic device 120 is configured to present XR content and enable optical pass-through or video pass-through of at least a portion of the physical environment 105 on the display 122 (e.g., the door 115). For example, the electronic device 120 corresponds to a mobile phone, a tablet, a laptop, a near-eye system, a wearable computing device, and the like.
[0089] like Figure 5A As shown, during instance 500 of a content delivery scenario (e.g., associated with time T0), electronic device 120 presents an XR environment 128 including an XR object 502 (e.g., a cube). Figure 4A As shown, the XR environment 128 also includes visualization 508 of the user 150's gaze direction or gaze vector. Depending on various specific implementations, as referenced above... Figure 5A As described, the eye-tracking engine 412 can determine and update the gaze direction or gaze vector 413 over time. Those skilled in the art will understand that the visualization 508 can be removed in various embodiments or replaced with other forms or configurations in various other embodiments. Those skilled in the art will also understand that the user 150 can interact with the XR object 502 within the XR environment 128 based on various inputs (e.g., eye-tracking input, hand / limb tracking input, voice commands, etc.) such as zooming, panning, rotating, annotating, modifying the XR object 502, etc., in various embodiments.
[0090] like Figure 5A As shown, during instance 500, the visualization 508 of the user 150's gaze direction points to XR object 502. Figure 5A An attention accumulator 503 associated with the XR object 502 is also shown, which includes a current value 511 at time T0, which corresponds to the length of time since T0 that the user 150's gaze has been directed toward the XR object 502 relative to a reference time window (e.g., the past X seconds, Y frames, Z cycles, etc.). Figure 5A As shown, the attention accumulator 503 associated with the XR object 502 also includes an increment indicator 509, which indicates that the current value 511 of the attention accumulator 503 associated with the XR object 502 increases within time T0. Furthermore, in Figure 5C In this embodiment, the attention accumulator 503 associated with the XR object 502 also includes a first (selection) threshold 501. Those skilled in the art will understand that the visual representation of the attention accumulator 503 can be removed in various implementations or replaced with other forms or configurations in various other implementations.
[0091] In some specific implementations, based on the determination that the current value of the attention accumulator 503 associated with the XR object 502 has breached or exceeded a first (selection) threshold 501, the electronic device 120 selects the XR object 502, such as... Figure 5C As shown, and by altering the appearance of the XR object 502 to indicate that it has been selected, such as by presenting a border or frame 522 surrounding the XR object 502. Figure 5Ais shown. In some implementations, in accordance with a determination that the current value of the attention accumulator 503 associated with the XR object 502 does not breach or exceed the first (selection) threshold 501, the electronic device 120 forgoes selecting the XR object 502, as Figure 5B and Figure 5B are shown.
[0092] According to some implementations, the first (selection) threshold 501 corresponds to a predefined value or deterministic value, such as a predefined dwell time timer. According to some implementations, the first (selection) threshold 501 corresponds to a non-deterministic value that is dynamically determined or selected based on eye tracking accuracy, current foreground application, XR object classification / type, UI element classification / type, user history, user preferences, and the like.
[0093] As Figure 5B shown, during an instance 510 of the content delivery scenario (e.g., associated with time T1), the visualization 508 of the gaze direction of the user 150 remains pointed at the XR object 502. Figure 5B Also shown is the attention accumulator 503 associated with the XR object 502, which includes a current value 521 at time T1 that corresponds to a length of time that the gaze direction of the user 150 has been pointed at the XR object 502 since T1 relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, and the like). As Figure 5B shown, the attention accumulator 503 associated with the XR object 502 also includes an increase indicator 509 that indicates that the current value 521 increased within time T1 relative to time TO. In Figure 5C , the current value 521 does not breach or exceed the first (selection) threshold 501.
[0094] As Figure 5C shown, during an instance 520 of the content delivery scenario (e.g., associated with time T2), the visualization 508 of the gaze direction of the user 150 remains pointed at the XR object 502. Figure 5C Also shown is the attention accumulator 503 associated with the XR object 502, which includes a current value 531 at time T2 that corresponds to a length of time that the gaze direction of the user 150 has been pointed at the XR object 502 since T2 relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, and the like). As Figure 5C shown, the attention accumulator 503 associated with the XR object 502 also includes an increase indicator 509 that indicates that the current value 531 increased within time T2 relative to time T1. In Figure 5C , the current value 531 breaches or exceeds the first (selection) threshold 501.
[0095] As Figure 5DAs shown, electronic device 120 selects XR object 502 by determining that the current value 531 of attention accumulator 503 associated with XR object 502 has exceeded or surpassed a first (selection) threshold 501, and changes the appearance of XR object 502 (e.g., a cube) by presenting a border or frame 522 surrounding XR object 502. Those skilled in the art will understand that the manner in which the appearance of XR object 502 is changed to indicate its selection may vary in various other embodiments, such as changes in the color, texture, brightness, etc. of XR object 502. Those skilled in the art will also understand that electronic device 120 may provide other feedback to indicate that XR object 502 has been selected in various other embodiments, such as haptic feedback, auditory feedback, etc. In some embodiments, electronic device 120 may also perform functions associated with XR object 502 by determining that the current value 531 of attention accumulator 503 associated with XR object 502 has exceeded or surpassed the first (selection) threshold 501.
[0096] like Figure 5D As shown, during instance 530 of the content delivery scenario (e.g., associated with time T3), the visualization 508 of the user 150's gaze direction no longer points to the XR object 502. Figure 5D An attention accumulator 503 associated with the XR object 502 is also shown, which includes a current value 541 at time T3, which corresponds to the length of time since T3 that the user 150's gaze has been directed toward the XR object 502 relative to a reference time window (e.g., the past X seconds, Y frames, Z cycles, etc.). Figure 5D As shown, the attention accumulator 503 associated with the XR object 502 also includes a decrease indicator 539, which indicates that the current value 541 decreases relative to time T2 within time T3.
[0097] In addition, Figure 5H In this embodiment, the attention accumulator 503 associated with the XR object 502 also includes a second (deselection) threshold 532. In some implementations, based on the determination that the current value of the attention accumulator 503 associated with the XR object 502 exceeds or falls below the second (deselection) threshold 532, the electronic device 120 deselects the XR object 502 and alters the appearance of the XR object 502 to indicate its deselection by, for example, removing the border or frame 522 surrounding the XR object 502. Figure 5F As shown. In some specific implementations, based on the determination that the current value of the attention accumulator 503 associated with the XR object 502 has not exceeded or fallen below a second (deselection) threshold 532, the electronic device 120 maintains the selection of the XR object 502 and the rendering of the border or frame 522 surrounding the XR object 502 to indicate continued selection, such as... Figure 5G and Figure 5DAs shown. Figure 5D As shown, the current value 541 of the attention accumulator 503 associated with the XR object 502 at time T3 has not exceeded or decreased below the second (deselect) threshold 532. Therefore, in Figure 5E In this process, the electronic device 120 maintains the selection of the XR object 502.
[0098] In some implementations, the second (deselection) threshold 532 corresponds to a predefined or deterministic value, such as a predefined dwell time timer. In some implementations, the second (deselection) threshold 532 corresponds to a non-deterministic value that is dynamically determined or selected based on eye tracking accuracy, the current foreground application, XR object classification / type, UI element classification / type, user history, user preferences, etc.
[0099] like Figure 5E As shown, during instance 540 of the content delivery scenario (e.g., associated with time T4), the visualization 508 of the user 150's gaze direction is again pointed to the XR object 502. Figure 5E An attention accumulator 503 associated with the XR object 502 is also shown, which includes a current value 551 at time T4, which corresponds to the length of time since T4 that the user 150's gaze has been directed toward the XR object 502 relative to a reference time window (e.g., the past X seconds, Y frames, Z cycles, etc.). Figure 5E As shown, the attention accumulator 503 associated with the XR object 502 also includes an increment indicator 509, which indicates that the current value 551 increases relative to time T3 within time T4. Figure 5E As shown, the current value 551 of the attention accumulator 503 associated with XR object 502 at time T4 has not exceeded or decreased below the second (deselect) threshold 532. Therefore, in Figure 5F In this process, the electronic device 120 maintains the selection of the XR object 502.
[0100] like Figure 5F As shown, during instance 550 of the content delivery scenario (e.g., associated with time T5), the visualization 508 of the user 150's gaze direction no longer points to the XR object 502. Figure 5F An attention accumulator 503 associated with the XR object 502 is also shown, which includes a current value 561 at time T5, which corresponds to the length of time since T5 that the user 150's gaze has been directed toward the XR object 502 relative to a reference time window (e.g., the past X seconds, Y frames, Z cycles, etc.). Figure 5F As shown, the attention accumulator 503 associated with the XR object 502 also includes a decrease indicator 539, which indicates that the current value 561 decreases relative to time T4 within time T5. Figure 5FAs shown, the current value 561 of the attention accumulator 503 associated with XR object 502 at time T5 has not exceeded or decreased below the second (deselect) threshold 532. Therefore, in Figure 5G In this process, the electronic device 120 maintains the selection of the XR object 502.
[0101] like Figure 5G As shown, during instance 560 of the content delivery scenario (e.g., associated with time T6), the visualization 508 of the user 150's gaze direction does not point to the XR object 502. Figure 5G An attention accumulator 503 associated with the XR object 502 is also shown, which includes a current value 571 at time T6, which corresponds to the length of time since T6 that the user 150's gaze has been directed toward the XR object 502 relative to a reference time window (e.g., the past X seconds, Y frames, Z cycles, etc.). Figure 5G As shown, the attention accumulator 503 associated with the XR object 502 also includes a decrease indicator 539, which indicates that the current value 571 decreases relative to time T5 within time T6. Figure 5G As shown, the current value 571 of the attention accumulator 503 associated with XR object 502 at time T6 has not exceeded or decreased below the second (deselect) threshold 532. Therefore, in Figure 5H In this process, the electronic device 120 maintains the selection of the XR object 502.
[0102] like Figure 5H As shown, during instance 570 of the content delivery scenario (e.g., associated with time T7), the visualization 508 of the user 150's gaze direction does not point to the XR object 502. Figure 5H An attention accumulator 503 associated with the XR object 502 is also shown, which includes a current value 581 at time T7, corresponding to the length of time since T7 that the user 150's gaze has been directed toward the XR object 502 relative to a reference time window (e.g., the past X seconds, Y frames, Z cycles, etc.). Figure 5H As shown, the attention accumulator 503 associated with the XR object 502 also includes a decrease indicator 539, which indicates that the current value 581 decreases relative to time T6 within time T7.
[0103] like Figure 5H As shown, the current value 581 of the attention accumulator 503 associated with XR object 502 at time T7 exceeds or decreases below the second (deselect) threshold 532. Figures 6A to 6HAs shown, the electronic device 120 deselects the XR object 502 and changes the appearance of the XR object 502 (e.g., a stereoscopic cube) by removing the border or frame 522 that surrounds the XR object 502 in accordance with determining that the current value 581 of the attention accumulator 503 associated with the XR object 502 breaks or falls below the second (deselection) threshold 532.
[0104] Figures 6A to 6H A sequence of examples 600-670 of content delivery scenarios is shown in accordance with some implementations. While some specific features are shown, one of skill in the art will recognize from the present disclosure that various other features have not been shown for brevity and to not obscure more relevant aspects of the implementations disclosed herein. Figures 5A to 5H Similar to and adapted from Figures 5A to 5H Thus, Figures 6A to 6H and Figure 1 Common reference numbers are used in Figure 2 and Figure 1 Controller 110 shown in Figure 3 and Figures 6A to 6H Electronic device 120 shown in; and / or suitable combinations thereof.
[0105] As shown in Figure 1 The content delivery scenario includes a physical environment 105 and an XR environment 128 displayed on a display 122 of an electronic device 120 (e.g., associated with a user 150). While the user 150 is physically present within the physical environment 105, which includes a door 115 that is currently located within a FOV 111 of an outward-facing image sensor of the electronic device 120, the electronic device 120 presents the XR environment 128 to the user 150. Thus, in some implementations, the user 150 holds the electronic device 120 in their hand, similar to the operating environment 100 in Figure 6A
[0106] In other words, in some implementations, the electronic device 120 is configured to present XR content and enable optical pass-through or video pass-through of at least a portion of the physical environment 105 on the display 122 (e.g., the door 115). For example, the electronic device 120 corresponds to a mobile phone, a tablet, a laptop, a near-eye system, a wearable computing device, etc.
[0107] As shown in Figure 6A As shown, during an instance 600 of a content delivery scenario (e.g., associated with time TO), the electronic device 120 presents an XR environment 128 that includes an XR object 502 (e.g., a stereoscopic cube), an XR object 604 (e.g., a stereoscopic cylinder), and a virtual agent (VA) 606. As Figure 4A As shown, the XR environment 128 also includes a visualization 508 of the gaze direction or gaze vector of the user 150. According to various implementations, the gaze direction or gaze vector 413 can be determined and updated over time by the eye tracking engine 412, as described above with reference to FIG. 4. As will be appreciated by one of ordinary skill in the art, the visualization 508 can be removed in various implementations or replaced with other forms or configurations in various other implementations. As will be appreciated by one of ordinary skill in the art, the user 150 can interact with the XR object 502, the XR object 604, or the VA 606 within the XR environment 128 based on various inputs (e.g., eye tracking inputs, hand / limb tracking inputs, voice commands, etc.) such as zooming, panning, rotating, annotating, modifying the XR object 502, the XR object 604, or the VA 606, etc. in various implementations. Figure 6A As described, the eye tracking engine 412 can determine and update the gaze direction or gaze vector 413 over time. As will be appreciated by one of ordinary skill in the art, the visualization 508 can be removed in various implementations or replaced with other forms or configurations in various other implementations. As will be appreciated by one of ordinary skill in the art, the user 150 can interact with the XR object 502, the XR object 604, or the VA 606 within the XR environment 128 based on various inputs (e.g., eye tracking inputs, hand / limb tracking inputs, voice commands, etc.) such as zooming, panning, rotating, annotating, modifying the XR object 502, the XR object 604, or the VA 606, etc. in various implementations.
[0108] As shown, during the instance 600, the visualization 508 of the gaze direction of the user 150 is directed at the XR object 502. Figure 6A As shown, during the instance 600, the visualization 508 of the gaze direction of the user 150 is directed at the XR object 502. Figure 6A Also shown is an attention accumulator 503 associated with the XR object 502 that includes a current value 611A at time TO, which corresponds to a length of time that the gaze direction of the user 150 has been directed at the XR object 502 since TO relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, etc.).
[0109] Figure 6A Also shown are an attention accumulator 605 associated with the XR object 604 that includes a current value 611B (e.g., a null value) and an attention accumulator 607 associated with the VA 606 that includes a current value 611C (e.g., a null value). Figure 6A Also shown is a ranked list 601 within time TO that includes the current value 611A of the attention accumulator 503 associated with the XR object 604 ranked higher than the current value 611B of the attention accumulator 605 associated with the XR object 502 and the current value 611C of the attention accumulator 607 associated with the VA 606. As will be appreciated by one of ordinary skill in the art, the visual representations of the attention accumulators 503, 605, and 607 along with the ranked list 601 can be removed in various implementations or replaced with other forms or configurations in various other implementations.
[0110] As shown, during the instance 600, the visualization 508 of the gaze direction of the user 150 is directed at the XR object 502. Figure 6AAs shown, the attention accumulator 503 associated with the XR object 502 includes an increment indicator 509, which indicates that the current value 611A of the attention accumulator 503 associated with the XR object 502 increases within time T0. Furthermore, in Figure 6C In this context, attention accumulators 503, 605, and 607 also include a first (selection) threshold 501. In some specific implementations, based on the determination that the current value of the attention accumulator 503 associated with the XR object 502 exceeds the first (selection) threshold 501, the electronic device 120 selects the XR object 502, such as... Figure 6C As shown, and by altering the appearance of the XR object 502 to indicate that it has been selected, such as by presenting a border or frame 522 surrounding the XR object 502. Figure 6A As shown. In some specific implementations, based on the determination that the current value of the attention accumulator 503 associated with the XR object 502 has not exceeded or surpassed a first (selection) threshold 501, the electronic device 120 abandons the selection of the XR object 502, such as... Figure 6B and Figure 6A As shown. In Figure 6B In the current value 611A, the first (selection) threshold 501 has not been exceeded or broken.
[0111] like Figure 6B As shown, during instance 610 of the content delivery scenario (e.g., associated with time T1), the visualization 508 of the user 150's gaze direction remains pointed to the XR object 502. Figure 6B An attention accumulator 503 associated with the XR object 502 is also shown, which includes a current value 621A at time T1, which corresponds to the length of time since T1 that the user 150's gaze has been directed toward the XR object 502 relative to a reference time window (e.g., the past X seconds, Y frames, Z cycles, etc.). Figure 6B As shown, the attention accumulator 503 associated with the XR object 502 also includes an increment indicator 509, which indicates that the current value 621A increases relative to time T0 within time T1. Figure 6B In the current value 621A, the first (selection) threshold 501 has not been exceeded or broken.
[0112] Figure 6B Attention accumulator 605 associated with XR object 604, including current value 621B (e.g., null value), and attention accumulator 607 associated with VA 606, including current value 621C (e.g., null value), are also shown. Figure 6CAlso shown is a list 601 sorted by rank within time T1, which includes the current value 621B of the attention accumulator 605 associated with XR object 502 and the current value 621C of the attention accumulator 607 associated with VA 606, as well as the current value 621A of the attention accumulator 503 associated with XR object 604.
[0113] like Figure 6C As shown, during instance 620 of the content delivery scenario (e.g., associated with time T2), the visualization 508 of the user 150's gaze direction remains pointed to the XR object 502. Figure 6C An attention accumulator 503 associated with the XR object 502 is shown, which includes a current value 631A at time T2, which corresponds to the length of time since T2 that the user 150's gaze has been directed toward the XR object 502 relative to a reference time window (e.g., the past X seconds, Y frames, Z cycles, etc.). Figure 6C As shown, the attention accumulator 503 associated with the XR object 502 also includes an increment indicator 509, which indicates that the current value 631A increases relative to time T1 within time T2. Figure 6C In the current value 631A, the threshold value has exceeded the first (selection) threshold of 501. For example... Figure 6C As shown, the electronic device 120 selects the XR object 502 based on the determination that the current value 631A of the attention accumulator 503 associated with the XR object 502 has exceeded or surpassed a first (selection) threshold 501, and changes the appearance of the XR object 502 (e.g., a 3D cube) by presenting a border or frame 522 surrounding the XR object 502.
[0114] Figure 6C Attention accumulator 605 associated with XR object 604, including current value 631B (e.g., null value), and attention accumulator 607 associated with VA 606, including current value 631C (e.g., null value), are also shown. Figure 6D Also shown is a list 601 sorted by rank within time T2, which includes the current value 631A of the attention accumulator 503 associated with XR object 502, which is ranked higher than the current value 631B of the attention accumulator 605 associated with XR object 604 and the current value 631C of the attention accumulator 607 associated with VA 606.
[0115] like Figure 6D As shown, during instance 630 of the content delivery scenario (e.g., associated with time T3), the visualization 508 of the user 150's gaze direction no longer points to XR object 502, but instead points to XR object 604. Therefore, Figure 6DA gaze accumulator 503 associated with the XR object 502 is shown, including a current value 641A at time T3, which corresponds to a length of time that the gaze direction of the user 150 has been directed at the XR object 502 since T3, relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, etc.). As shown, Figure 6D The gaze accumulator 503 associated with the XR object 502 also includes a decrease indicator 539, indicating that the current value 641A decreased relative to time T2 within time T3.
[0116] Figure 6D A gaze accumulator 605 associated with the XR object 604 is also shown, including a current value 641B at time T3, which corresponds to a length of time that the gaze direction of the user 150 has been directed at the XR object 604 since T3, relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, etc.). As shown, Figure 6D The gaze accumulator 605 associated with the XR object 604 also includes an increase indicator 509, indicating that the current value 641B increased relative to time T2 within time T3. As shown, Figure 6D The current value 641B of the gaze accumulator 605 associated with the XR object 604 at time T3 does not breach or exceed the first (selection) threshold 501, as shown. Figure 6D A gaze accumulator 607 associated with the VA 606 is also shown, including a current value 641C (e.g., a null value).
[0117] Figure 6D A ranked list 601 within time T3 is also shown, including the current value 641A of the gaze accumulator 503 associated with the XR object 502, which is ranked higher than the current value 641B of the gaze accumulator 605 associated with the XR object 604, which is ranked higher than the current value 641C of the gaze accumulator 607 associated with the VA 606. Thus, in Figure 6E the electronic device 120 maintains selection of the XR object 502. Notably, since the attention (e.g., gaze direction) of the user 150 was directed at the XR object 502 for a length of time sufficient to satisfy the first (selection) threshold 501, the electronic device 120 can maintain selection of the XR object 502, even though the gaze direction of the user 150 was temporarily directed at the XR object 604. This advantageously prevents de-selection of a virtual object that the user expressed strong, recent interest in (e.g., represented by a relatively high current value 641A) due to, e.g., eye tracking noise, unintentional eye movement, etc.
[0118] As Figure 6EAs shown, during an instance 640 of the content delivery scenario (e.g., associated with time T4), the visualization 508 of the gaze direction of the user 150 remains pointed towards the XR object 604. To this end, Figure 6E The attention accumulator 503 associated with the XR object 502 is shown to include a current value 651A at time T4, which corresponds to a length of time that the gaze direction of the user 150 has been pointed towards the XR object 502 since T4 relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, etc.). As Figure 6E The attention accumulator 503 associated with the XR object 502 is shown to further include a decrease indicator 539, which indicates that the current value 651A decreased relative to time T3 within time T4.
[0119] Figure 6E The attention accumulator 605 associated with the XR object 604 is also shown to include a current value 651B at time T4, which corresponds to a length of time that the gaze direction of the user 150 has been pointed towards the XR object 604 since T4 relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, etc.). As Figure 6E The attention accumulator 605 associated with the XR object 604 is shown to further include an increase indicator 509, which indicates that the current value 651B increased relative to time T3 within time T4. Figure 6E The attention accumulator 607 associated with the VA 606 is also shown to include a current value 651C (e.g., a null value).
[0120] As Figure 6E The current value 651B of the attention accumulator 605 associated with the XR object 604 at time T4 is shown to be the largest in the rank-ordered list 601 and is also greater than the current value 651A of the attention accumulator 503 associated with the XR object 604 at time T4, and the current value 651B of the attention accumulator 605 associated with the XR object 502 at time T4 breaks or exceeds the first (selection) threshold 501. Figure 6E The rank-ordered list 601 within time T4 is also shown to include the current value 651B of the attention accumulator 605 associated with the XR object 604 ranked higher than the current value 651A of the attention accumulator 503 associated with the XR object 502, which is ranked higher than the current value 641C of the attention accumulator 607 associated with the VA 606. Accordingly, as Figure 6F The electronic device 120 is shown to deselect the XR object 502 and select the XR object 604. In addition, the electronic device 120 changes the appearance (e.g., a stereoscopic cube) of the XR object 604 by presenting a border or frame 522 that surrounds the XR object 604.
[0121] As Figure 6F shown, during an instance 650 of the content delivery scenario (e.g., associated with time T5), the visualization 508 of the gaze direction of the user 150 no longer points to the XR object 604, and instead points to the VA 606. To this end, Figure 6F a focus accumulator 503 associated with the XR object 502 is shown, which includes a current value 661A at time T5 that corresponds to a length of time that the gaze direction of the user 150 has been pointed at the XR object 502 since T5 relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, etc.). As Figure 6F shown, the focus accumulator 503 associated with the XR object 502 also includes a decrease indicator 539 that indicates that the current value 661A decreased relative to time T4 within time T5.
[0122] Figure 6F A focus accumulator 605 associated with the XR object 604 is also shown, which includes a current value 661B at time T5 that corresponds to a length of time that the gaze direction of the user 150 has been pointed at the XR object 604 since T5 relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, etc.). As Figure 6F shown, the focus accumulator 605 associated with the XR object 604 also includes a decrease indicator 539 that indicates that the current value 661B decreased relative to time T4 within time T5.
[0123] Figure 6F A focus accumulator 607 associated with the VA 606 is also shown, which includes a current value 661C at time T5 that corresponds to a length of time that the gaze direction of the user 150 has been pointed at the VA 606 since T5 relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, etc.). As Figure 6F shown, the focus accumulator 607 associated with the VA 606 also includes an increase indicator 509 that indicates that the current value 661C increased relative to time T4 within time T5. As Figure 6F shown, the current value 661C of the focus accumulator 607 associated with the VA 606 at time T5 does not breach or exceed the first (selection) threshold 501.
[0124] Figure 6F A ranked list 601 within time T5 is also shown, which includes the current value 661B of the focus accumulator 605 associated with the XR object 604 ranked higher than the current value 661C of the focus accumulator 607 associated with the VA 606, which is ranked higher than the current value 661A of the focus accumulator 503 associated with the XR object 502. Thus, inFigure 6G In particular, the electronic device 120 maintains the selection of the XR object 604.
[0125] As Figure 6G shown, during an instance 660 of the content delivery scenario (e.g., associated with time T6), the visualization 508 of the gaze direction of the user 150 no longer points to the VA 606, and instead points to the XR object 604. To this end, Figure 6G shown, the attention accumulator 503 associated with the XR object 502 includes a current value 671A at time T6 (e.g., a null value) that corresponds to a length of time that the gaze direction of the user 150 has been pointed at the XR object 502 since T6 relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, etc.). As Figure 6G shown, the attention accumulator 503 associated with the XR object 502 also includes a decrease indicator 539 that indicates that the current value 671A decreased relative to time T5 within time T6.
[0126] Figure 6G Also shown is the attention accumulator 605 associated with the XR object 604 that includes a current value 671B at time T6 that corresponds to a length of time that the gaze direction of the user 150 has been pointed at the XR object 604 since T6 relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, etc.). As Figure 6G shown, the attention accumulator 605 associated with the XR object 604 also includes an increase indicator 509 that indicates that the current value 671B increased relative to time T5 within time T6.
[0127] Figure 6G Also shown is the attention accumulator 607 associated with the VA 606 that includes a current value 671C at time T6 that corresponds to a length of time that the gaze direction of the user 150 has been pointed at the VA 606 since T6 relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, etc.). As Figure 6G shown, the attention accumulator 607 associated with the VA 606 also includes a decrease indicator 539 that indicates that the current value 671C decreased relative to time T5 within time T6.
[0128] Figure 6G Also shown is a ranked list 601 within time T6 that includes the current value 671B of the attention accumulator 605 associated with the XR object 604 that is ranked higher than the current value 671C of the attention accumulator 607 associated with the VA 606 that is ranked higher than the current value 671A of the attention accumulator 503 associated with the XR 502. Thus, in Figure 6HIn particular, the electronic device 120 maintains the selection of the XR object 604.
[0129] As Figure 6H shown, during an instance 670 of the content delivery scenario (e.g., associated with time T7), the visualization 508 of the gaze direction of the user 150 no longer points to the XR object 604. To this end, Figure 6H a focus of attention accumulator 503 associated with the XR object 502 is shown, which includes a current value 681A at time T7 (e.g., a null value) that corresponds to a length of time that the gaze direction of the user 150 has been pointed at the XR object 502 since T7 relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, etc.).
[0130] Figure 6H A focus of attention accumulator 605 associated with the XR object 604 is also shown, which includes a current value 681B at time T7 that corresponds to a length of time that the gaze direction of the user 150 has been pointed at the XR object 604 since T7 relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, etc.). As Figure 6H shown, the focus of attention accumulator 605 associated with the XR object 604 also includes a decrease indicator 539 that indicates that the current value 681B decreased relative to time T6 within time T7.
[0131] Figure 6H A focus of attention accumulator 607 associated with the VA 606 is also shown, which includes a current value 681C at time T7 (e.g., a null value) that corresponds to a length of time that the gaze direction of the user 150 has been pointed at the VA 606 since T7 relative to a reference time window (e.g., past X seconds, Y frames, Z cycles, etc.). As Figure 6H shown, the focus of attention accumulator 607 associated with the VA 606 also includes a decrease indicator 539 that indicates that the current value 681C decreased relative to time T6 within time T7.
[0132] Figure 7A to 7C A ranked list 601 within time T7 is also shown, which includes the current value 681B of the focus of attention accumulator 605 associated with the XR object 604 that is ranked higher than the current value 681A of the focus of attention accumulator 503 associated with the XR object 502 and the current value 681C of the focus of attention accumulator 607 associated with the VA 606. Accordingly, in Figure 1 particular, the electronic device 120 maintains the selection of the XR object 604.
[0133] While the above examples show attention accumulators with the same first (selection) threshold 501 and the same second (deselection) threshold 532, one of ordinary skill in the art will appreciate that, in other implementations, different attention accumulators can have the same or different first and second thresholds. These values can correspond to predetermined values or non-deterministic values that are dynamically selected or determined based on eye tracking accuracy, current foreground application, UI element type, user history, user preferences, etc. Additionally, in some implementations, the rate at which the value of an accumulator increases can be the same or different from the rate at which it decreases. Furthermore, the rate at which the value of an accumulator increases or decreases can be the same or different from those rates of other accumulators. These values can correspond to predetermined values or non-deterministic values that are dynamically selected or determined based on eye tracking accuracy, current foreground application, UI element type, user history, user preferences, etc.
[0134] Figure 3 A flowchart representation of a method 700 of selecting UI elements with eye tracking based attention accumulators is shown in accordance with some implementations. In various implementations, the method 700 is performed at a computing system that includes non-transitory memory and one or more processors, where the computing system is communicatively coupled to a display device and one or more input devices (e.g., Figure 1 and Figure 2 the electronic device 120 shown; Figure 4A and Figure 4B the controller 110 shown; or suitable combinations thereof). In some implementations, the method 700 is performed by processing logic, including hardware, firmware, software, or a combination thereof. In some implementations, the method 700 is performed by a processor executing code stored in a non-transitory computer-readable medium (e.g., memory). In some implementations, the electronic system corresponds to one of a tablet computer, a laptop computer, a mobile phone, a near-eye system, a wearable computing device, etc. In some implementations, the one or more input devices correspond to a computer vision (CV) engine using image streams from one or more outward-facing image sensors, a finger / hand / limb tracking engine, an eye tracking engine, a touch-sensitive surface, one or more microphones, etc.
[0135] As discussed above, eye tracking can be noisy and / or inaccurate. A dwell time timer can be used to select UI elements with eye tracking input. However, using a dwell time timer to switch between UI elements can result in additional user experience (UX) issues, such as discontinuities or jank. In turn, the methods described herein reduce eye tracking noise and improve overall UX by employing a rank-ordered attention accumulator scheme when selecting UI elements based on eye tracking input.
[0136] As represented by block 710, while currently selecting (e.g., focusing on) a first user interface (UI) element, the method 700 includes detecting a first gaze direction directed at a second UI element different from the first UI element. For example, with reference to Figure 4C and Figure 6D The computing device or a portion thereof (e.g., the eye tracking engine 412) obtains (e.g., receives, retrieves, or detects / determines / generates) an eye tracking vector 413 (e.g., the first gaze direction) for a current time period and updates the eye tracking vector 413 over time. For example, with reference to Figures 6A to 6H The computing device or a portion thereof (e.g., the rendering engine 450) renders a user interface (UI) or an extended reality (XR) environment that includes the first UI element and the second UI element. As one example, with reference to Figure 6D While currently selecting (e.g., as indicated by the border or frame 522 that surrounds the XR object 604) the XR object 502 (e.g., the first UI element), the electronic device 120 detects a gaze direction directed at the XR object 502 (e.g., the second UI element).
[0137] In some implementations, the display device corresponds to a transparent lens assembly, and wherein presenting the UI or the XR environment includes projecting the UI or the XR environment onto the transparent lens assembly. In some implementations, the display device corresponds to a near-eye system, and wherein presenting the UI or the XR environment includes compositing the UI or the XR environment with one or more images of the physical environment captured by an outward-facing image sensor.
[0138] In some implementations, the appearance of the currently selected UI element is different from the unselected UI elements, such as a different color, texture, brightness, highlighted, etc. In some implementations, a border, frame, etc. surrounds the currently selected UI element. In some implementations, a spotlight, shadow, etc. is cast on or around the currently selected UI element.
[0139] In some implementations, as represented by block 712, the first UI element and the second UI element correspond to one of a selectable affordance, a selectable button, an interactive (non-binary) UI element (e.g., a dial, a slider, etc.), a notification, an extended reality (XR) object, etc. As one example, with reference to Figure 6D The electronic device 120 presents an XR environment 128 that includes the XR object 502 (e.g., the first UI element), the XR object 604 (e.g., the second UI element), and the VA 606 (e.g., the third UI element).
[0140] As represented by block 720, in response to detecting the first gaze direction pointing to the second UI element, the method 700 includes decreasing the first attention accumulator value associated with the first UI element and increasing the second attention accumulator value associated with the second UI element based on a length of time that the first gaze direction pointed to the second UI element. As one example, refer to FIG. 6B, which illustrates a decrease in the current value 641 A of the attention accumulator 503 associated with the XR object 604 and an increase in the current value 641B of the attention accumulator 605 associated with the XR object 502. Figure 6D The attention accumulator 503 associated with the XR object 604 includes a decrease indicator 539 indicating that the current value 641A decreased relative to time T2 over time T3, and the attention accumulator 605 associated with the XR object 502 includes an increase indicator 509 indicating that the current value 641B increased relative to time T2 over time T3, because Figure 6C the gaze direction in 604 is no longer pointing to the XR object 604, and instead is pointing to the XR object 502.
[0141] In some implementations, as represented by block 722, the first attention accumulator value and the second attention accumulator value are stored in a ranked list of attention accumulator values. As one example, refer to FIG. 6C, which illustrates a ranked list 601 of attention accumulator values over time T3. Figure 6D In some implementations, as represented by block 722, the first attention accumulator value and the second attention accumulator value are stored in a ranked list of attention accumulator values. As one example, refer to FIG. 6C, which illustrates a ranked list 601 of attention accumulator values over time T3.
[0142] In some implementations, the computing system selects the UI element associated with the highest attention accumulator value in the ranked list. In some implementations, the computing system selects the UI element associated with the highest attention accumulator value in the ranked list so long as the associated attention accumulator value exceeds a first (selection) threshold for the last N time periods. In some implementations, when multiple attention accumulator values are included in the ranked list, the attention accumulator values decrease / increase more quickly than when a single attention accumulator value is included in the ranked list. In some implementations, when multiple non-zero attention accumulator values are included in the ranked list, the attention accumulator values decrease / increase more quickly than when a single attention accumulator value is included in the ranked list.
[0143] In some implementations, in contrast to when a single attention accumulator value is included in the rank-ordered list, when multiple attention accumulator values are included in the rank-ordered list, the computing system determines or modifies the first (selection) threshold and the second (deselection) threshold. In other words, in contrast to when a single attention accumulator value is included in the rank-ordered list, when multiple attention accumulator values are included in the rank-ordered list, UI elements can lose selection more quickly. In one example, when the computing system presents two UI elements within the UI, whose attention accumulator values are thus between 0 and 1.0, the first (selection) threshold corresponds to 0.75, and the second (deselection) threshold corresponds to 0.25. Continuing the example, when the first UI element is currently selected / focused and the attention accumulator value of the second UI element exceeds 0.75, if the attention accumulator value of the second UI element exceeds the attention accumulator value of the first UI element (even if the attention accumulator value of the first UI element exceeds 0.25), the computing system deselects / defocuses the first UI element and selects / focuses the second UI element.
[0144] As represented by block 730, in accordance with a determination that the second attention accumulator value associated with the second UI element does not exceed the first attention accumulator value associated with the first UI element, the method 700 includes maintaining selection of the first UI element. As one example, the attention accumulator 503 associated with the XR object 604 decreases from a value 631A in Figure 6C to a value 641A in Figure 6D and the attention accumulator 605 associated with the XR object 502 increases from a value 631B in Figure 6D to a value 641B in Figure 6D because the gaze direction in Figure 6D is no longer directed at the XR object 502 and instead is directed at the XR object 604. Continuing the example, as shown by the rank-ordered list 601 over time T3 in Figure 6D the current value 641A of the attention accumulator 503 associated with the XR object 502 is the largest in the rank-ordered list 601 and is also greater than the current value 641B of the attention accumulator 605 associated with the XR object 604. Accordingly, in Figure 6C the electronic device 120 maintains selection of the XR object 502.
[0145] In some implementations, as represented by block 732, the method 700 also maintains selection of the first UI element in accordance with a determination that the first attention accumulator value associated with the first UI element does not break or fall below the second (deselection) threshold. For example, with reference to Figure 6D the electronic device 120 maintains selection of the XR object 502 when the attention accumulator 503 associated with the XR object 502 decreases from Figure 6Dthe value 631A in the ranked list 601 at time T2 is reduced to Figure 6D the selection of the XR object 502 is maintained after the value 641A in the ranked list 601 at time T3 because Figure 6D the value 641A in the ranked list 601 at time T3 is greater than the second (deselection) threshold 532 and Figure 6E the value 641A in the ranked list 601 at time T3 is the greatest.
[0146] In some implementations, the method 700 deselects the first UI element according to a determination that the first attention accumulator value associated with the first UI element breaks or falls below a second (deselection) threshold. In some implementations, the second (deselection) threshold corresponds to a predetermined, predefined, or deterministic value. In some implementations, the second (deselection) threshold corresponds to a non-deterministic value that is dynamically selected or determined based on eye tracking accuracy, current foreground application, UI element type, user history, user preferences, and the like.
[0147] As represented by block 740, according to a determination that the second attention accumulator value associated with the second UI element exceeds the first attention accumulator value associated with the first UI element, the method 700 includes deselecting the first UI element and selecting the second UI element. In some implementations, the method 700 also deselects the first UI element and selects the second UI element according to a determination that the second attention accumulator value associated with the second UI element breaks or exceeds a first (selection) threshold. Further, in some implementations, the method 700 also deselects the first UI element and selects the second UI element according to a determination that the second attention accumulator value associated with the second UI element is the greatest in the ranked list.
[0148] As one example, the attention accumulator 503 associated with the XR object 604 is reduced from Figure 6D the value 631A in the ranked list 601 at time T2 to Figure 6E the value 651A in the ranked list 601 at time T3, and the attention accumulator 605 associated with the XR object 502 is increased from Figure 6E the value 641B in the ranked list 601 at time T2 to Figure 6E the value 651B in the ranked list 601 at time T3 because Figure 6E the gaze direction in the ranked list 601 at time T3 remains directed at the XR object 604. Continuing with this example, as shown by the ranked list 601 at time T3 in Figure 6E the current value 651B of the attention accumulator 605 associated with the XR object 604 is the greatest in the ranked list 601 and also greater than the current value 651A of the attention accumulator 503 associated with the XR object 502. Further, the current value 651B of the attention accumulator 605 associated with the XR object 604 breaks or exceeds the first (selection) threshold 501 in the ranked list 601 at time T3. Accordingly, as shown by the ranked list 601 at time T3 in Figure 6E the XR object 604 is selected and the XR object 502 is deselected. Figure 6EAs shown, the electronic device 120 deselects the XR object 502 and selects the XR object 604. In addition, the electronic device 120 changes the appearance of the XR object 604 (e.g., a stereoscopic cube) by presenting a border or frame 522 that surrounds the XR object 604.
[0149] In some implementations, as represented by block 742, the method 700 also deselects the first UI element and selects the second UI element according to a determination that the second attention accumulator value associated with the second UI element breaches or exceeds the first (selection) threshold. In this example, the method 700 deselects the first UI element even though the first attention accumulator value associated with the first UI element does not breach or fall below the second (selection) threshold. For example, with reference to Figure 6D , the electronic device 120 deselects the XR object 502 (e.g., the first UI element) and selects the XR object 604 (e.g., the second UI element) because the value 651B exceeds or breaches the first (selection) threshold 501 and the value 651B is greater than the value 651A.
[0150] In some implementations, as represented by block 744, the method 700 also deselects the first UI element and selects the second UI element according to a determination that the second attention accumulator value associated with the second UI element breaches or exceeds the first (selection) threshold and additionally according to a determination that the second attention accumulator value associated with the second UI element is the largest in the rank-ordered list. Continuing the above, the method 700 deselects the first UI element even though the first attention accumulator value associated with the first UI element does not breach or fall below the second (selection) threshold. For example, with reference to Figure 6E , the electronic device 120 deselects the XR object 502 (e.g., the first UI element) and selects the XR object 604 (e.g., the second UI element) because the value 651B exceeds or breaches the first (selection) threshold 501 and the value 651B is the largest in the rank-ordered list 601.
[0151] In some implementations, the first (selection) threshold corresponds to a predetermined, predefined, or deterministic value. In some implementations, the first (selection) threshold corresponds to a non-deterministic value that is dynamically selected or determined based on eye tracking accuracy, a current foreground application, a UI element type, a user history, a user preference, and the like.
[0152] In some implementations, the first UI element is deselected after the first attention accumulator value associated with the first UI element decreases for at least two consecutive time periods. In some implementations, the respective time periods correspond to a CPU cycle, a frame refresh rate, a deterministic amount of time (e.g., 1 ms), a non-deterministic amount of time, and the like. As one example, the attention accumulator 503 associated with the XR object 502 decreases for two consecutive time periods 505A and 505B.Figure 6D and Figure 6D decreases because the gaze direction is not directed at the XR object 502 for two consecutive time periods in Figure 6E and 6E .
[0153] In some implementations, the second UI element is selected after the second attention accumulator value associated with the second UI element increases for at least two consecutive time periods. In some implementations, the respective time periods correspond to CPU cycles, frame refresh rates, deterministic amounts of time (e.g., 1 ms), non-deterministic amounts of time, etc. As one example, the attention accumulator 605 associated with the XR object 604 increases in Figure 6D and Figure 6D because the gaze direction is directed at the XR object 604 for two consecutive time periods in Figure 6E and 6E .
[0154] In some implementations, as represented by block 750, in response to deselecting the first UI element, the method 700 includes changing an appearance of the first UI element. In some implementations, changing the appearance of the first UI element corresponds to changing a color, texture, shape, brightness, etc. of the first UI element. In some implementations, changing the appearance of the first UI element corresponds to removing a border or frame that surrounds the first UI element, removing a spotlight on or around the first UI element, removing a shadow on or around the first UI element, presenting a deselection animation, etc. As one example, the computing system changes the appearance of the XR object 502 (e.g., the first UI element) by removing the frame or border 522 that surrounds the XR object 502 to indicate that the XR object 502 has been deselected between Figure 6D and Figure 6E .
[0155] In some implementations, as represented by block 760, in response to selecting the second UI element, the method 700 includes changing an appearance of the second UI element. In some implementations, changing the appearance of the second UI element corresponds to changing a color, texture, shape, brightness, etc. of the second UI element. In some implementations, changing the appearance of the second UI element corresponds to presenting a border or frame that surrounds the second UI element, presenting a spotlight on or around the second UI element, presenting a shadow on or around the second UI element, presenting a selection animation, etc. As one example, the computing system changes the appearance of the XR object 604 (e.g., the second UI element) by presenting the frame or border 522 that surrounds the XR object 604 to indicate that the XR object 604 has been selected between Figure 6E and Figure 6E .
[0156] In some implementations, as represented by block 770, in response to selecting the second UI element, the method 700 includes performing an operation associated with the second UI element. As one example, the computing system can perform an operation (e.g., zoom in, rotate, pan, scale, etc.) associated with the XR object 604 (e.g., the second UI element) in response to selecting the XR object 604 in Figure 6F
[0157] In some implementations, as represented by block 780, the method includes: while the second UI element is currently selected (e.g., focused), detecting a second gaze direction that is directed at the first UI element; in response to detecting the second gaze direction that is directed at the first UI element: decreasing a second attention accumulator value associated with the second UI element and increasing a first attention accumulator value associated with the first UI element based on a length of time that the second gaze direction is directed at the first UI element; in accordance with a determination that the first attention accumulator value associated with the first UI element exceeds the second attention accumulator value associated with the second UI element, deselecting the second UI element and selecting the first UI element; and in accordance with a determination that the first attention accumulator value associated with the first UI element does not exceed the second attention accumulator value associated with the second UI element, maintaining selection of the second UI element.
[0158] As one example, the attention accumulator 605 associated with the XR object 604 is decreased from a value 651B in Figure 6E to a value 661B in Figure 6F and the attention accumulator 607 associated with the VA 606 is increased from a value 651C in Figure 6F to a value 661C in Figure 6F because the gaze direction in Figure 6F is no longer directed at the XR object 604 and is instead directed at the VA 606. Continuing with this example, as shown by the rank ordered list 601 over time T5 in Figure 5A the current value 661B of the attention accumulator 605 associated with the XR object 604 is the largest in the rank ordered list 601 and is also greater than the current value 661C of the attention accumulator 607 associated with the VA 606. Accordingly, in Figure 5B the electronic device 120 maintains selection of the XR object 604.
[0159] In some implementations, the method 700 includes: while the first UI element is not currently selected, detecting a second gaze direction directed at the first UI element; in response to detecting the second gaze direction directed at the first UI element, increasing a first attention accumulator value associated with the first UI element based on a length of time that the second gaze direction is directed at the first UI element; in accordance with a determination that the first attention accumulator value associated with the first UI element exceeds a first (selection) threshold, selecting the first UI element; and in accordance with a determination that the first attention accumulator value associated with the first UI element does not exceed the first (selection) threshold, forgoing selecting the first UI element.
[0160] As one example, with reference to Figure 5A and Figure 5B , in accordance with a determination that the current value of the attention accumulator 503 associated with the XR object 502 does not breach or exceed the first (selection) threshold 501, the electronic device 120 forgoes selecting the XR object 502, as shown in Figure 5C and Figure 5C As another example, with reference to Figure 5C , in accordance with a determination that the current value of the attention accumulator 503 associated with the XR object 502 breaches or exceeds the first (selection) threshold 501, the electronic device 120 selects the XR object 502, as shown in Figure 5F , and changes the appearance of the XR object 502, such as by presenting a border or frame 522 that surrounds the XR object 502 to indicate that it has been selected, as shown in Figure 5G .
[0161] In some implementations, the method 700 includes: while the first UI element is currently selected, detecting a third gaze direction that is not directed at the first UI element; in response to detecting the third gaze direction that is not directed at the first UI element, decreasing the first attention accumulator value associated with the first UI element based on a length of time that the third gaze direction is not directed at the first UI element; in accordance with a determination that the first attention accumulator value associated with the first UI element falls below a second (deselection) threshold, deselecting the first UI element; and in accordance with a determination that the first attention accumulator value associated with the first UI element does not fall below the second (deselection) threshold, maintaining selection of the first UI element.
[0162] In some implementations, the first threshold and the second threshold correspond to the same value. In some implementations, the first threshold and the second threshold correspond to different values. In some implementations, at least one of the first threshold and the second threshold is a predetermined, predefined, or deterministic value. In some implementations, at least one of the first threshold and the second threshold is a non-deterministic value that is dynamically selected or determined based on eye tracking accuracy, current foreground application, UI element type, user history, user preferences, etc. In some implementations, the first threshold and / or the second threshold can dynamically change based on the number of attention accumulator values in the ranked list, etc.
[0163] As one example, referring to Figure 5F and Figure 5G , in accordance with a determination that the current value of the attention accumulator 503 associated with the XR object 502 does not break or fall below the second (deselection) threshold 532, the electronic device 120 maintains selection of the XR object 502 and presentation of the border or frame 522 surrounding the XR object 502 to indicate its continued selection, as shown in Figure 5H and Figure 5H As another example, referring to , in accordance with a determination that the current value of the attention accumulator 503 associated with the XR object 502 breaks or falls below the second (deselection) threshold 532, the electronic device 120 deselects the XR object 502 and changes the appearance of the XR object 502 to indicate its deselection by, for example, removing the border or frame 522 surrounding the XR object 502, as shown in .
[0164] Thus, in some examples, the method 700 can be used to select a single UI element from one or more UI elements in accordance with a determination that its associated attention accumulator value exceeds or breaks the first (selection) threshold and is the highest attention accumulator value of the one or more UI elements. The method 700 can also be used to deselect the selected UI element in accordance with a determination that its associated attention accumulator value falls below or breaks the second (deselection) threshold or that the attention accumulator value of another UI element both exceeds or breaks the first (selection) threshold and is the highest attention accumulator value of the one or more UI elements.
[0165] While various aspects of implementations within the scope of the appended claims are described above, it should be apparent that the various features of implementations described can be embodied in a wide variety of forms and that any specific structure and / or function described above is merely illustrative. Based on the teachings herein one skilled in the art should appreciate that an aspect described herein can be implemented independently of any other aspects and that two or more of these aspects can be combined in various ways. For example, an apparatus can be implemented and / or a method practiced using any number of the aspects described herein. In addition, an apparatus can be implemented and / or a method can be practiced using other structure and / or functionality in addition to or other than one or more of the aspects described herein.
[0166] It will also be understood that, although the terms“first,”“second,” etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first media item could be termed a second media item, and, similarly, a second media item could be termed a first media item, this changing the meanings of the descriptions only so long as the context allows for such designation. The first media item and the second media item are both media items, but they are not the same media item.
[0167] The terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting of the claims. As used in the description of the implementations and the appended claims, the singular forms“a,”“an,” and“the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term“and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms“comprises” and / or“comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0168] As used herein, the term "if' can be construed to mean "when" or "upon" the condition is true or "in response to the determination" or "in accordance with the determination" or "in response to detecting" that a condition is true, depending on the context. Similarly, the phrase "if it is determined [that a condition is true]" or "if [a condition is true]" or "when [a condition is true]" is construed to mean "upon determining" or "in response to determining" or "in accordance with determining" that the condition is true or "when [a condition is true]" or "in response to detecting" that the condition is true, depending on the context.
Claims
1. A method comprising: at a computing system that includes non-transitory memory and one or more processors, wherein the computing system is communicatively coupled to a display device and one or more input devices via a communication interface: while a first user interface (UI) element is currently selected, detecting a first gaze direction that is directed at a second UI element that is different from the first UI element; in response to detecting the first gaze direction that is directed at the second UI element, decreasing a first attention accumulator value associated with the first UI element and increasing a second attention accumulator value associated with the second UI element based on a length of time that the first gaze direction is directed at the second UI element; in accordance with a determination that the second attention accumulator value associated with the second UI element exceeds the first attention accumulator value associated with the first UI element, deselecting the first UI element and selecting the second UI element; and in accordance with a determination that the second attention accumulator value associated with the second UI element does not exceed the first attention accumulator value associated with the first UI element, maintaining selection of the first UI element.
2. The method of claim 1, wherein maintaining selection of the first UI element is in accordance with a determination that the second attention accumulator value associated with the second UI element does not exceed the first attention accumulator value associated with the first UI element and further in accordance with a determination that the first attention accumulator value associated with the first UI element does not breach or fall below a second threshold value.
3. The method of any of claims 1 and 2, wherein deselecting the first UI element and selecting the second UI element is in accordance with a determination that the second attention accumulator value associated with the second UI element exceeds the first attention accumulator value associated with the first UI element and further in accordance with a determination that the second attention accumulator value associated with the second UI element breaches or exceeds a first threshold value.
4. The method of any of claims 1 and 2, wherein deselecting the first UI element and selecting the second UI element is in accordance with a determination that the second attention accumulator value associated with the second UI element exceeds the first attention accumulator value associated with the first UI element and further in accordance with a determination that the second attention accumulator value associated with the second UI element breaches or exceeds a first threshold value and additionally in accordance with a determination that the second attention accumulator value associated with the second UI element is the largest in a ranked list of attention accumulator values that includes at least the first attention accumulator value and the second attention accumulator value.
5. The method of claim 3, wherein the first UI element is deselected even if the first attention accumulator value associated with the first UI element does not breach or fall below a second threshold value that is different from the first threshold value. 6. The method of any one of claims 1, 2, and 5, wherein the first UI element is deselected after the first attention accumulator value associated with the first UI element decreases for at least two consecutive time periods.
7. The method of any one of claims 1, 2, and 5, wherein the second UI element is selected after the second attention accumulator value associated with the second UI element increases for at least two consecutive time periods.
8. The method of any one of claims 1, 2, and 5, wherein the first attention accumulator value and the second attention accumulator value are stored in a ranked list of attention accumulator values.
9. The method of any one of claims 1, 2, and 5, wherein the first UI element and the second UI element correspond to one of a selectable affordance, a selectable button, an interactive UI element, a notification, or an extended reality (XR) object.
10. The method of any one of claims 1, 2, and 5, further comprising: in response to deselecting the first UI element, changing an appearance of the first UI element.
11. The method of any one of claims 1, 2, and 5, further comprising: in response to selecting the second UI element, changing an appearance of the second UI element.
12. The method of any one of claims 1, 2, and 5, further comprising: in response to selecting the second UI element, performing an operation associated with the second UI element.
13. The method of any one of claims 1, 2, and 5, further comprising: while the second UI element is currently selected, detecting a second gaze direction directed at the first UI element; in response to detecting the second gaze direction directed at the first UI element, decreasing the second attention accumulator value associated with the second UI element and increasing the first attention accumulator value associated with the first UI element based on a length of time that the second gaze direction is directed at the first UI element; in accordance with a determination that the first attention accumulator value associated with the first UI element exceeds the second attention accumulator value associated with the second UI element, deselecting the second UI element and selecting the first UI element; and in accordance with a determination that the first attention accumulator value associated with the first UI element does not exceed the second attention accumulator value associated with the second UI element, maintaining selection of the second UI element.
14. The method of any one of claims 1, 2, and 5, further comprising: while no UI element is currently selected, detecting a second gaze direction directed at the first UI element; in response to detecting the second gaze direction directed at the first UI element, increasing the first attention accumulator value associated with the first UI element based on a length of time that the second gaze direction is directed at the first UI element; in accordance with a determination that the first attention accumulator value associated with the first UI element exceeds a first threshold value, selecting the first UI element; and in accordance with a determination that the first attention accumulator value associated with the first UI element does not exceed the first threshold value, maintaining selection of the second UI element. in accordance with a determination that the first attention accumulator value associated with the first UI element does not exceed the first threshold, forgoing selection of the first UI element.
15. The method of claim 14, further comprising: while currently selecting the first UI element, detecting a third gaze direction that is not directed at the first UI element; in response to detecting the third gaze direction that is not directed at the first UI element, decreasing the first attention accumulator value associated with the first UI element based on a length of time that the third gaze direction was not directed at the first UI element; in accordance with a determination that the first attention accumulator value associated with the first UI element falls below a second threshold, deselecting the first UI element; and in accordance with a determination that the first attention accumulator value associated with the first UI element does not fall below the second threshold, maintaining selection of the first UI element.
16. A device, the device comprising: one or more processors; non-transitory memory; an interface to communicate with a display device and one or more input devices; and one or more programs stored in the non-transitory memory that, when executed by the one or more processors, cause the device to perform any of the methods of claims 1-15.
17. Non-transitory memory storing one or more programs, which when executed by one or more processors of a device with an interface to communicate with a display device and one or more input devices, cause the device to perform any of the methods of claims 1-15.
Citation Information
Patent Citations
Enhanced Electronic Gaming Machine with Gaze=Based Features
AU2016273828A1
Signal processing method, device, electronic equipment and storage medium
CN112716506A