A method of interacting with virtual controls and / or affordances for moving virtual objects in a virtual environment

The computer system addresses the inefficiencies in augmented reality interactions by implementing intuitive and feedback-rich interfaces, reducing user input requirements and enhancing the overall experience while conserving energy.

JP7692474B2Active Publication Date: 2025-06-13APPLE INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023519045
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-23
Filing Date
2021-09-25
Publication Date
2025-06-13
Estimated Expiration
2041-09-25

AI Technical Summary

Technical Problem

Existing computer systems for augmented reality have cumbersome, inefficient, and limited methods and interfaces for interacting with environments that include virtual elements, leading to a significant cognitive burden on users and inefficient use of energy, particularly in battery-operated devices.

Method used

The development of a computer system with improved methods and interfaces that reduce the number, degree, and type of inputs required from users by providing enhanced feedback and intuitive interaction with selectable user interface elements, slider elements, and virtual objects in three-dimensional environments.

Benefits of technology

The system enhances user interaction efficiency and intuitiveness, reducing errors and energy consumption, while improving the overall computer-generated reality experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007692474000001
    Figure 0007692474000001
  • Figure 0007692474000002
    Figure 0007692474000002
  • Figure 0007692474000003
    Figure 0007692474000003
Patent Text Reader

Abstract

In some embodiments, the electronic device enhances interaction with virtual objects in a three-dimensional environment. In some embodiments, the electronic device enhances interaction with selectable user interface elements. In some embodiments, the electronic device enhances interaction with slider user interface elements. In some embodiments, the electronic device moves virtual objects in the three-dimensional environment and facilitates access to actions associated with the virtual objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims priority to U.S. Provisional Patent Application No. 63 / 083,802, filed on September 25, 2020, and U.S. Provisional Patent Application No. 63 / 261,555, filed on September 23, 2021, the entire contents of which are incorporated herein by reference for all purposes.

[0002] (Technical Field) This relates generally to a computer system having a display generation component and one or more input devices that present a graphical user interface including, but not limited to, an electronic device that presents a three - dimensional environment including virtual objects via the display generation component.

Background Art

[0003] The development of computer systems for augmented reality has advanced significantly in recent years. Exemplary augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices such as cameras, controllers, joysticks, touch - sensitive surfaces, and touch - screen displays for computer systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Exemplary virtual elements include virtual objects including digital images, videos, text, icons, and control elements such as buttons and other graphics.

[0004] However, methods and interfaces for interacting with environments (such as applications, augmented reality environments, mixed reality environments, and virtual reality environments) that include at least some virtual elements are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback to perform actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems where the manipulation of virtual objects is complex and error-prone impose a significant cognitive burden on the user and degrade the experience in the virtual / augmented reality environment. Additionally, those methods are time-consuming more than necessary, thereby wasting energy. This latter consideration is particularly important in battery-operated devices. SUMMARY OF THE INVENTION

[0005] Accordingly, there is a need for a computer system having improved methods and interfaces for providing a computer-generated experience that makes the interaction with the computer system more efficient and intuitive for the user. Such methods and interfaces can complement or replace conventional methods of providing a computer-generated reality experience to the user. Such methods and interfaces reduce the number, degree, and / or type of inputs from the user by assisting the user in understanding the connection between the input provided and the device response to that input, thereby creating a more efficient human-machine interface.

[0006] The above-described deficiencies and other problems associated with a user interface for a computer system having a display generation component and one or more input devices are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer having an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a "touch screen" or "touch screen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to the display generation component, and the output devices include one or more haptic output generators and one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or sets of instructions stored in the memory for performing a plurality of functions. In some embodiments, the user interacts with the GUI through a stylus and / or finger contact and gestures on a touch-sensitive surface, the movement of the user's eyes and hands in space relative to the user's body when captured by a camera and other motion sensors, and voice input when captured by one or more audio input devices.In some embodiments, the functions executed through the interaction optionally include image editing, drawing, presenting, word processing, spreadsheet creation, game play, making a phone call, video conferencing, sending an email, instant messaging, training support, digital photography, digital video shooting, web browsing, playing digital music, taking notes, and / or playing digital video. The executable instructions for executing those functions are optionally included in a non-transitory computer-readable storage medium or other computer program product configured to be executed by one or more processors.

[0007] There is a need for an improved method and interface for navigating and interacting with a user interface in an electronic device. Such a method and interface may complement or replace conventional methods for interacting with objects in a three-dimensional environment. Such a method and interface reduce the number, degree, and / or type of inputs from the user and generate a more efficient human-machine interface.

[0008] In some embodiments, the electronic device enhances the interaction with selectable user interface elements. In some embodiments, the electronic device enhances the interaction with slider user interface elements. In some embodiments, the electronic device moves virtual objects in a three-dimensional environment and facilitates access to actions associated with the virtual objects.

[0009] Note that the various embodiments described above can be combined with any other embodiments described herein. The functions and advantages described herein are not exhaustive, and in particular, many additional functions and advantages will be apparent to those skilled in the art in view of the drawings, specification, and claims. Further, note that the language used herein has been selected solely for readability and for the purpose of explanation and not for the purpose of defining or limiting the subject matter of the invention.

Brief Description of the Drawings

[0010] To better understand the various embodiments described, the following "Modes for Carrying Out the Invention" should be referred to in conjunction with the following drawings, and like reference numerals refer to corresponding parts throughout the following figures.

[0011]

Figure 1

[0012]

Figure 2

[0013]

Figure 3

[0014]

Figure 4

[0015]

Figure 5

[0016]

Figure 6

[0017]

Figure 7A

Figure 7B

Figure 7C

Figure 7D

[0018]

Figure 8A

Figure 8B

Figure 8C

Figure 8D

Figure 8E

Figure 8F

Figure 8G

Figure 8H

Figure 8I

Figure 8J

Figure 8K

Figure 8L

Figure 8M

[0019]

Figure 9A

Figure 9B

Figure 9C

Figure 9D

Figure 9E

[0020]

Figure 10A

Figure 10B

Figure 10C

Figure 10D

Figure 10E

Figure 10F

Figure 10G

Figure 10H

Figure 10I

Figure 10J

[0021]

Figure 11A

Figure 11B

Figure 11C

Figure 11D

[0022]

Figure 12A

Figure 12B

Figure 12C

Figure 12D

Figure 12E

Figure 12F

Figure 12G

Figure 12H

Figure 12I

Figure 12J

Figure 12K

Figure 12L

Figure 12M

Figure 12N

Figure 12O

[0023]

Figure 13A

Figure 13B

Figure 13C

Figure 13D

Figure 13E

Figure 13F

[0024]

Figure 14A

Figure 14B

Figure 14C

Figure 14D

Figure 14E

Figure 14F

Figure 14G

Figure 14H

Figure 14I

Figure 14J

Figure 14K

Figure 14L

DETAILED DESCRIPTION OF THE INVENTION

[0025] The present disclosure relates to a user interface for providing a computer-generated reality (CGR) experience to a user according to some embodiments.

[0026] The systems, methods, and GUIs described herein provide an improved way for an electronic device to interact with and manipulate objects within a three-dimensional environment. The three-dimensional environment optionally includes one or more virtual objects, one or more representations of real objects (e.g., displayed as a realistic (e.g., "pass-through") representation of the real object or visible to the user through a transparent portion of a display generation component) within the physical environment of the electronic device, and / or a representation of the user within the three-dimensional environment.

[0027] In some embodiments, the electronic device facilitates interaction with selectable user interface elements. In some embodiments, the electronic device presents one or more selectable user interface elements in the three-dimensional environment. In some embodiments, in response to detecting the user's line of sight directed at an individual selectable user interface element, the electronic device updates the appearance of the selectable user interface element, e.g., increasing the z-separation of the selectable user interface element from another portion of the user interface. In some embodiments, the electronic device performs a related action in response to user input including selecting a user interface element, detecting the user's line of sight, and / or the user performing a predetermined gesture with a hand. Enhancing interaction with such selectable user interface elements provides an efficient and intuitive way for the electronic device to make selections and perform actions.

[0028] In some embodiments, the electronic device enhances the interaction with the slider user interface element. In some embodiments, the slider user interface element includes an indication of the current input state of the slider user interface. In some embodiments, in response to detecting the user's line of sight on the slider user interface element, the electronic device updates the slider user interface element to include an indication of a plurality of available input states of the slider user interface element. Optionally, the electronic device changes the current input state of the slider user interface element in response to an input including detecting the user's line of sight and / or detecting a user performing a predetermined hand gesture. Enhancing the interaction with the slider user interface element provides an efficient way to adjust the input state of the slider user interface element and perform an action on the electronic device associated with the slider.

[0029] In some embodiments, the electronic device moves a virtual object in a three-dimensional environment and facilitates access to actions associated with the virtual object. In some embodiments, the electronic device displays a user interface element associated with a virtual object within the virtual environment. In some embodiments, in response to detecting a first input directed to the user interface element, the electronic device initiates a process for moving the associated virtual object within the virtual environment. In some embodiments, in response to detecting a second input directed to the user interface element, the electronic device updates the user interface element to include a plurality of selectable options that, when selected, cause the electronic device to perform an individual action directed to the virtual object. Moving the virtual object and enhancing additional actions directed to the virtual object using the user interface element provides an efficient way to interact with the virtual object.

[0030] In some embodiments, the electronic device facilitates interaction with selectable user interface elements and provides enhanced visual feedback in response to detecting at least a portion of a selection input directed to a selectable user interface element. In some embodiments, the electronic device presents selectable user interface elements within a first container user interface element that is within a second container user interface element. In response to detecting a user's line of sight directed to a selectable user interface element, the electronic device, in some embodiments, updates the appearance of the selectable user interface element and the first container user interface element, such as by increasing the z-separation of the selectable user interface element from the first container user interface element and increasing the z-separation of the first container user interface element from the second container user interface element. In some embodiments, in response to the initiation of a selection input, the electronic device decreases the visual separation between the selectable user interface element and the first container user interface element. In some embodiments, in response to the continuation of a selection input corresponding to decreasing the z-height of the selectable user interface element by an amount that exceeds the visual separation between the selectable user interface element and the first container user interface element, the electronic device continues the visual feedback by decreasing the z-height of the selectable user interface element and the first container user interface element and decreasing the visual separation between the first container user interface element and the second container user interface element as the input continues. In some embodiments, in response to the continuation of a selection input corresponding to decreasing the z-heights of the selectable user interface element and the first container user interface element by an amount that exceeds the amount of z-separation between the first container user interface element and the second container user interface element, the electronic device decreases the z-heights of the selectable user interface element, the first container user interface element, and the second container user interface element as the input continues.Enhancing the interaction with selectable user interface elements in this way provides an efficient and intuitive way for an electronic device to make selections and perform actions.

[0031] Figures 1 - 6 provide an illustration of an exemplary computer system for providing a CGR experience (described below with respect to, for example, methods 800, 1000, 1200, and 1400) to a user. In some embodiments, as shown in FIG. 1, the CGR experience is provided to the user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., speakers 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input devices 125, output devices 155, sensors 190, and peripheral devices 195 are integrated with the display generation component 120 (e.g., within a head-mounted device or a handheld device).

[0032] When describing the CGR experience, various related but distinct environments that a user can perceive and / or interact with (e.g., using inputs detected by computer system 101 which generates audio, visual, and / or tactile feedback corresponding to various inputs provided to computer system 101 for generating the CGR experience) are individually referred to by various terms for the sake of separately discussing them. The following is a subset of these terms.

[0033] Physical environment: The physical environment refers to the physical world in which people can perceive and / or interact without the aid of an electronic system. Physical environments such as a physical park include physical objects such as physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through, for example, vision, touch, hearing, taste, and smell.

[0034] Computer-Generated Reality: In contrast, a computer-generated reality (CGR) environment refers to an environment that is wholly or partially simulated and through which people can perceive and / or interact via an electronic system. In CGR, a subset of a person's body movements or their representations are tracked, and in response, one or more characteristics of one or more virtual objects simulated within the CGR environment are adjusted to behave in accordance with at least one law of physics. For example, a CGR system can detect the rotation of a person's head and, in response, adjust the graphic content and sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. Depending on the situation (e.g., for accessibility reasons), the adjustment of the characteristics of the virtual object(s) in the CGR environment may be made in response to a representation of a body movement (e.g., a voice command). A person may use any one of these senses, including vision, hearing, touch, taste, and smell, to perceive and / or interact with the CGR object. For example, a person can perceive and / or interact with a sound object that creates an audio environment with 3D or spatial spread that provides the perception of a point sound source in 3D space. In another example, a sound object can enable audio transparency that selectively incorporates ambient sound from the physical environment, with or without including computer-generated audio. In some CGR environments, a person may only perceive and / or interact with sound objects.

[0035] Examples of CGR include virtual reality and mixed reality.

[0036] Virtual Reality: A virtual reality (VR) environment refers to an imitation environment designed to be based entirely on computer-generated sensory inputs for one or more senses. A VR environment includes multiple virtual objects that a person can perceive and / or interact with. For example, computer-generated images representing trees, buildings, and avatars of people are examples of virtual objects. A person can perceive and / or interact with the virtual objects in a VR environment through a simulation of the presence of the person within the computer-generated environment and / or through a simulation of a subset of the person's body movements within the computer-generated environment.

[0037] Mixed Reality: In contrast to a VR environment designed to be based entirely on computer-generated sensory inputs, a mixed reality (MR) environment refers to an imitation environment designed to incorporate sensory inputs or representations thereof from the physical environment in addition to including computer-generated sensory inputs (e.g., virtual objects). On the virtual continuum, an MR environment is anywhere between, but not including, the complete physical environment at one end and the virtual reality environment at the other end. In some MR environments, the computer-generated sensory inputs can respond to changes in the sensory inputs from the physical environment. Also, some electronic systems for presenting an MR environment may track the location and / or orientation with respect to the physical environment to enable virtual objects to interact with real objects (i.e., physical articles or representations thereof from the physical environment). For example, the system can take movement into account so that a virtual tree appears stationary with respect to the physical ground.

[0038] Examples of mixed reality include augmented reality and augmented virtuality.

[0039] Augmented Reality: An augmented reality (AR) environment refers to an emulated environment in which one or more virtual objects are superimposed on a physical environment or its representation. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on the transparent or translucent display, whereby a person can use the system to perceive virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture an image or video of the physical environment, which is a representation of the physical environment. The system synthesizes the image or video with virtual objects and presents the composite on the opaque display. A person uses this system to indirectly view the physical environment through the image or video of the physical environment and to perceive virtual objects superimposed on the physical environment. As used herein, a video of a physical environment shown on an opaque display is referred to as a "pass-through video," meaning that the system uses one or more image sensors to capture an image of the physical environment and uses those images when presenting the AR environment on the opaque display. Further alternatively, the system may have a projection system that projects virtual objects, for example as holograms, into the physical environment or onto a physical surface, whereby a person can use the system to perceive virtual objects superimposed on the physical environment. An augmented reality environment also refers to a simulation environment in which a representation of the physical environment is transformed by computer-generated sensory information. For example, when providing a pass-through video, the system may transform one or more sensor images to map to a selected perspective (e.g., viewpoint) different from the perspective captured by the image sensor. As another example, a representation of the physical environment may be transformed by graphically modifying (e.g., magnifying) a portion thereof, whereby the modified portion can be made into a modified version that represents the original captured image but is non-photorealistic. As a further example, a representation of the physical environment may be transformed by graphically removing or obscuring a portion thereof.

[0040] Extended Virtual: An extended virtual (AV) environment refers to an imitative environment in which a virtual environment or a computer-generated environment incorporates one or more sensory inputs from the physical environment. The sensory input can be a representation of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, but people with faces are realistically reproduced from images of physical people. As another example, a virtual object may adopt the shape or color of a physical item imaged by one or more imaging sensors. As a further example, a virtual object can adopt a shadow that coincides with the position of the sun in the physical environment.

[0041] Hardware: There are many different types of electronic systems that enable a person to perceive and / or interact with various CGR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display functionality, windows with integrated display functionality, displays formed as lenses designed to be placed over a person's eyes (similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without tactile feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speakers (singular or plural) and an integrated opaque display. Alternatively, a head-mounted system may be configured to receive an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or videos of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system may have a transparent or translucent display instead of an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed towards a person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scan light sources, or any combination of these technologies. The medium may be an optical waveguide, hologram medium, optical coupler, optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to selectively become opaque. A projection-based system can employ retinal projection technology that projects a graphical image onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or as physical surfaces.In some embodiments, the controller 110 is configured to manage and adjust the user's CGR experience. In some embodiments, the controller 110 includes a suitable combination of software, firmware, and / or hardware. The controller 110 will be described in more detail below with reference to FIG. 2. In some embodiments, the controller 110 is a computing device that is local or remote to the scene 105 (e.g., the physical environment). For example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside the scene 105. In some embodiments, the controller 110 is communicatively coupled to the display generation component 120 (e.g., an HMD, a display, a projector, a touch screen, etc.) via one or more wired or wireless communication channels 144 (e.g., BLUETOOTH, IEEE802.11x, IEEE802.16x, IEEE802.3x, etc.). In another example, the controller 110 is included within the housing (e.g., the physical housing) of one or more of the display generation component 120 (e.g., an HMD, or a portable electronic device including a display and one or more processors, etc.), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or one or more of the peripheral devices 195, or shares the same physical housing or support structure as one or more of the above.

[0042] In some embodiments, the display generation component 120 is configured to provide the user with a CGR experience (e.g., at least the visual component of the CGR experience). In some embodiments, the display generation component 120 includes a suitable combination of software, firmware, and / or hardware. The display generation component 120 will be described in more detail below with reference to FIG. 3. In some embodiments, the functions of the controller 110 are provided by and / or combined with the display generation component 120.

[0043] According to some embodiments, the display generation component 120 provides the user with a CGR experience while the user is virtually and / or physically present within the scene 105.

[0044] In some embodiments, the display generation component is worn on a part of the user's body (e.g., the user's own head or hand). Accordingly, the display generation component 120 includes one or more CGR displays provided for displaying CGR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or tablet) configured to present CGR content, and the user holds a device having a display directed towards the user's field of view and a camera directed towards scene 105. In some embodiments, the handheld device is optionally disposed within a housing worn on the user's head. In some embodiments, the handheld device is optionally disposed on a support in front of the user (e.g., a tripod). In some embodiments, the display generation component 120 is a CGR chamber, housing, or room configured to present CGR content with the user not wearing or holding the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying CGR content (e.g., a device on a handheld or tripod) may be implemented on another type of hardware for displaying CGR content (e.g., an HMD or other wearable computing device). For example, a user interface showing an interaction with CGR content triggered based on an interaction occurring within the space in front of a handheld or tripod-mounted device may be implemented in the same manner as an HMD where the interaction occurs within the space in front of the HMD and the response of the CGR content is displayed via the HMD. Similarly, a user interface showing an interaction with CGR content triggered based on movement of a handheld or tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hand)) may be implemented in the same manner as an HMD caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hand)).

[0045] Although the relevant features of the operating environment 100 are shown in FIG. 1, those skilled in the art will understand from this disclosure that various other features for the sake of simplicity are not shown so as not to obscure more appropriate aspects of the exemplary embodiments disclosed herein.

[0046] FIG. 2 is a block diagram of an example of the controller 110 according to some embodiments. Although certain features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more appropriate aspects of the embodiments disclosed herein. Thus, by way of non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), BLUETOOTH, ZIGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, a memory 220, and one or more communication buses 204 for interconnecting these and various other components.

[0047] In some embodiments, one or more communication buses 204 include circuitry for interconnecting system components and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, and the like.

[0048] Memory 220 includes high-speed random access memory such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220, or the non-transitory computer-readable storage medium of memory 220, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and a CGR experience module 240.

[0049] The operating system 230 includes instructions for processing various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the CGR experience module 240 is configured to manage and coordinate one or more CGR experiences for one or more users (e.g., a single CGR experience for one or more users, or multiple CGR experiences for each group of one or more users). For that purpose, in various embodiments, the CGR experience module 240 includes a data acquisition unit 242, a tracking unit 244, an adjustment unit 246, and a data transmission unit 248.

[0050] In some embodiments, the data acquisition unit 242 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 of FIG. 1 and optionally from one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For that purpose, in various embodiments, the data acquisition unit 242 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0051] In some embodiments, the tracking unit 244 is configured to map the scene 105 and track the position / location of at least the display generation component 120 of FIG. 1 and optionally one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195 with respect to the scene 105 of FIG. 1. For that purpose, in various embodiments, the tracking unit 244 includes instructions and / or logic therefor, as well as heuristics and metadata therefor. In some embodiments, the tracking unit 244 includes a hand tracking unit 243 and / or an eye tracking unit 245. In some embodiments, the hand tracking unit 243 is configured to track the position / location of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand with respect to the scene 105 of FIG. 1, with respect to the display generation component 120, and / or with respect to a coordinate system defined for the user's hand. The hand tracking unit 243 will be described in more detail below with respect to FIG. 4. In some embodiments, the eye tracking unit 245 is configured to track the position and movement of the user's line of sight (or more generally the user's eyes, face, or head) with respect to the scene 105 (e.g., the physical environment and / or the user (e.g., the user's hand)) or with respect to the CGR content displayed via the display generation component 120. The eye tracking unit 245 will be described in more detail below with respect to FIG. 5.

[0052] In some embodiments, adjustment unit 246 is configured to manage and adjust the CGR experience presented to the user by display generation component 120 and optionally by one or more of output device 155 and / or peripheral device 195. For that purpose, in various embodiments, adjustment unit 246 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0053] In some embodiments, data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least display generation component 120 and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. For that purpose, in various embodiments, data transmission unit 248 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0054] Although data acquisition unit 242, tracking unit 244 (including, e.g., eye tracking unit 243 and hand tracking unit 244), adjustment unit 246, and data transmission unit 248 are shown as being present on a single device (e.g., controller 110), it should be understood that in other embodiments, any combination of data acquisition unit 242, tracking unit 244 (including, e.g., eye tracking unit 243 and hand tracking unit 244), adjustment unit 246, and data transmission unit 248 may be disposed within separate computing devices.

[0055] Furthermore, FIG. 2 is more intended to illustrate the functions of various features that may exist in a particular embodiment, as contrasted with the structural overview of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, some of the functional modules separately shown in FIG. 2 can be implemented within a single module, and the various functions of a single functional block can be executed by one or more functional blocks in various embodiments. The actual number of modules, as well as the specific division of particular functions and how functions are allocated among them, vary depending on the implementation form and, in some embodiments, depend in part on a particular combination of hardware, software, and / or firmware selected for a particular implementation form.

[0056] FIG. 3 is a block diagram of an example of a display generation component 120 according to some embodiments. While certain features are shown, it will be understood from this disclosure by those skilled in the art that various other features are not shown for the sake of brevity so as not to obscure more suitable aspects of the embodiments disclosed herein. For that purpose, by way of non-limiting example, in some embodiments, the HMD 120 includes one or more processing units 302 (e.g., microprocessor, ASIC, FPGA, GPU, CPU, processing core, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, BLUETOOTH, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more CGR displays 312, one or more optional inward and / or outward image sensors 314, a memory 320, and one or more communication buses 304 for interconnecting these and various other components.

[0057] In some embodiments, one or more communication buses 304 include circuitry that interconnects system components and controls communication between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of an inertial measurement unit (IMU), accelerometer, gyroscope, thermometer, one or more physiological sensors (e.g., blood pressure monitor, heart rate monitor, blood oxygen sensor, blood glucose sensor, etc.), one or more microphones, one or more speakers, a tactile engine, one or more depth sensors (e.g., structured light, time of flight, etc.), and the like.

[0058] In some embodiments, one or more CGR displays 312 are configured to provide a CGR experience to a user. In some embodiments, one or more CGR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light emitting field effect transistor (OLET), organic light emitting diode (OLED), surface conduction electron emission device display (SED), field emission display (FED), quantum dot light emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, one or more CGR displays 312 correspond to waveguide displays, such as diffractive, reflective, polarizing, holographic, etc. For example, HMD 120 includes a single CGR display. In another example, HMD 120 includes a CGR display for each eye of the user. In some embodiments, one or more CGR displays 312 can present MR or VR content. In some embodiments, one or more CGR displays 312 can present MR or VR content.

[0059] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of a user's face, including the user's eyes (and may be referred to as an eye tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand(s) and optionally the user's arm(s) (and may be referred to as a hand tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene viewed by the user when the HMD 120 is not present (and may be referred to as a scene camera). The one or more optional image sensors 314 can include one or more RGB cameras, one or more infrared (IR) cameras, one or more event-based cameras, and / or the like (e.g., including a complementary metal oxide semiconductor (CMOS) image sensor or a charge coupled device (CCD) image sensor).

[0060] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320, or the non-transitory computer-readable storage medium of memory 320, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 330 and a CGR presentation module 340.

[0061] The operating system 330 includes instructions for processing various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the CGR presentation module 340 is configured to present CGR content to a user via one or more CGR displays 312. For that purpose, in various embodiments, the CGR presentation module 340 includes a data acquisition unit 342, a CGR presentation unit 344, a CGR map generation unit 346, and a data transmission unit 348.

[0062] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 of FIG. 1. For that purpose, in various embodiments, the data acquisition unit 342 includes instructions and / or logic therefor, and heuristics and metadata therefor.

[0063] In some embodiments, the CGR presentation unit 344 is configured to present CGR content via one or more CGR displays 312. For that purpose, in various embodiments, the CGR presentation unit 344 includes instructions and / or logic therefor, and heuristics and metadata therefor.

[0064] In some embodiments, the CGR map generation unit 346 is configured to generate a CGR map (e.g., a 3D map of a composite reality scene or a map of a physical environment in which computer-generated objects can be placed) based on media content data. For that purpose, in various embodiments, the CGR map generation unit 346 includes instructions and / or logic therefor, and heuristics and metadata therefor.

[0065] In some embodiments, data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least controller 110 and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. For that purpose, in various embodiments, data transmission unit 348 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0066] Data acquisition unit 342, CGR presentation unit 344, CGR map generation unit 346, and data transmission unit 348 are shown as being present on a single device (e.g., display generation component 120 of FIG. 1), but in other embodiments, it should be understood that any combination of data acquisition unit 342, CGR presentation unit 344, CGR map generation unit 346, and data transmission unit 348 may be arranged within separate computing devices.

[0067] Furthermore, FIG. 3 is more intended to illustrate the functions of various features that may exist in a particular implementation as contrasted with the structural overview of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules separately shown in FIG. 3 can be implemented within a single module, and the various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the specific division of a particular function and how functions are allocated therebetween, vary depending on the implementation, and in some embodiments, depend in part on a particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0068] FIG. 4 is a schematic diagram of an exemplary embodiment of a hand tracking device 140. In some embodiments, the hand tracking device 140 (FIG. 1) is controlled by a hand tracking unit 243 (FIG. 2) to track the position / location of one or more portions of a user's hand and / or the movement of one or more portions of the user's hand relative to a coordinate system defined for the user's hand (e.g., relative to a physical environment surrounding the user, relative to a display generation component 120, or relative to a portion of the user (e.g., the user's face, eyes, or head)). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0069] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures hand images at a resolution sufficient to distinguish the fingers and their respective positions. The image sensor 404 typically captures images of other parts of the user's body or all of the body and can have either a zoom function or a dedicated sensor with high magnification to capture hand images at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in combination with other image sensors that capture the physical environment of the scene 105 or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are processed as inputs to the controller 110.

[0070] In some embodiments, the image sensor 404 outputs a sequence of frames including 3D map data (and optionally color image data) to the controller 110, thereby extracting high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API) and drives the display generation component 120 accordingly. For example, the user may interact with software operating on the controller 110 by moving the hand 408 and changing the hand's posture.

[0071] In some embodiments, the image sensor 404 projects a spot pattern onto a scene that includes the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the spots of the pattern. This approach is advantageous in that the user does not need to hold or wear any kind of beacon, sensor, or other marker. This gives the depth coordinates of points in the scene relative to a predetermined reference plane at a particular distance from the image sensor 404. In the present disclosure, it is assumed that the image sensor 404 defines a series of orthogonal x, y, and z axes such that the depth coordinates of points in the scene correspond to the z component measured by the image sensor. Alternatively, the hand tracking device 440 can use other 3D mapping methods such as stereoscopic imaging or time-of-flight measurement based on single or multiple cameras or other types of sensors.

[0072] In some embodiments, the hand tracking device 140 captures and processes a temporal sequence of depth maps of the user's hand while the user is moving the hand (e.g., the entire hand or one or more fingers). Software operating on a processor within the image sensor 404 and / or the controller 110 processes the 3D map data to extract hand patch descriptors within these depth maps. The software compares these descriptors to patch descriptors stored in the database 408 based on a previous learning process to estimate the hand pose in each frame. The pose typically includes the 3D locations of the user's hand joints and fingertips.

[0073] Software can also analyze the trajectories of the hand and / or fingers over multiple frames within a sequence to identify gestures. The pose estimation function described herein may be interleaved with the motion tracking function, such that patch-based pose estimation is only performed once every two (or more) frames, while tracking is used to detect changes in pose occurring over the remaining frames. Pose, motion, and gesture information is provided to an application program running on controller 110 via the API described above. This program can, for example, move and modify an image presented on display generation component 120 or perform other functions in response to pose and / or gesture information.

[0074] In some embodiments, the software may be downloaded in electronic form to the controller 110, for example, over a network, or alternatively, may be provided on a tangible non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in a memory associated with the controller 110. Alternatively or additionally, some or all of the described functions of the computer may be implemented in dedicated hardware such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). The controller 110 is shown in FIG. 4 as a separate unit from the image sensor 440, by way of example, but some or all of the processing functions of the controller can be associated with the image sensor 404 by a suitable microprocessor and software, or by dedicated circuitry within the housing of the handtracking device 402, or in other ways. In some embodiments, at least some of these processing functions are performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, a handheld device, or a head-mounted device), or using any other suitable computerized device such as a game console or a media player. The sensing function of the image sensor 404 can similarly be integrated with a computer or other computerized device controlled by the sensor output.

[0075] FIG. 4 further includes a schematic diagram of a depth map 410 captured by an image sensor 404 according to some embodiments. The depth map includes a matrix of pixels having respective depth values, as described above. The pixel 412 corresponding to the hand 406 is segmented from the background and the wrist in this map. The luminance of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z - distance from the image sensor 404, and the tone becomes darker as the depth increases. The controller 110 processes these depth values to identify and segment components (i.e., groups of adjacent pixels) of an image having characteristics of a human hand. These characteristics can include, for example, the overall size, shape, and movement from frame to frame of a sequence of depth maps.

[0076] FIG. 4 also schematically shows a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406 according to some embodiments. In FIG. 4, the skeleton 414 is superimposed on the background 416 of the hand segmented from the original depth map. In some embodiments, the hand (e.g., knuckles, fingertips, center of the palm, the end of the hand connected to the wrist, etc.), and optionally major feature points on the wrist or arm connected to the hand are identified and located on the hand skeleton 414. In some embodiments, the locations and movements of these major feature points over a plurality of image frames are used by the controller 110 to determine, according to some embodiments, a hand gesture or the current state of the hand performed by the hand.

[0077] FIG. 5 shows an exemplary embodiment of the eye tracking device 130 (FIG. 1). In some embodiments, the eye tracking device 130 is controlled by an eye tracking unit 245 (FIG. 2) to track the position and movement of the user's line of sight with respect to the scene 105 or with respect to the CGR content displayed via the display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, if the display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device disposed on a wearable frame, the head-mounted device includes both a component for generating CGR content for viewing by the user and a component for tracking the user's line of sight with respect to the CGR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, if the display generation component is a handheld device or a CGR chamber, the eye tracking device 130 is optionally a device separate from the handheld device or the CGR chamber. In some embodiments, the eye tracking device 130 is a head-mounted device or a part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used with a display generation component worn on the head or a display generation component not worn on the head. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally a part of a non-head-mounted display generation component.

[0078] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that presents a frame including left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display on which the user can directly view the physical environment and display virtual objects. In some embodiments, the display generation component projects virtual objects onto the physical environment. The virtual objects are projected, for example, onto a physical surface or as a hologram, whereby an individual can use the system to observe virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.

[0079] As shown in FIG. 5, in some embodiments, the gaze tracking device 130 includes at least one eye tracking camera (e.g., an infrared (IR) or near-IR (NIR) camera), and an illumination source that emits light (e.g., IR or NIR light) towards the user's eyes (e.g., an IR or NIR light source such as an array or ring of LEDs). The eye tracking camera may be directed towards the user's eyes to directly receive reflected IR or NIR light from the light source, or alternatively, may be directed towards a "hot" mirror disposed between the user's eyes and a display panel that reflects IR or NIR light from the eyes while allowing visual light to pass through to the eye tracking camera. The gaze tracking device 130 optionally captures an image of the user's eyes (e.g., as a video stream captured at 60 - 120 frames per second (fps)), analyzes the image to generate gaze tracking information, and communicates the gaze tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by respective eye tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by an individual eye tracking camera and illumination source.

[0080] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye tracking device for a particular operating environment 100, such as the 3D geometric relationships and parameters of the LED, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at a factory or another facility prior to delivery of the AR / VR device to the end user. The device-specific calibration process may be an automatic calibration process or a manual calibration process. The user-specific calibration process may include an estimation of the eye parameters of a particular user, such as pupil location, foveal position, optical axis, visual axis, interpupillary distance, etc. According to some embodiments, once the device-specific and user-specific parameters for the eye tracking device 130 are determined, the images captured by the eye tracking camera are processed using a glint assist method to determine the user's current visual axis and viewpoint with respect to the display.

[0081] As shown in FIG. 5, an eye tracking device 130 (e.g., 130A or 130B) includes a line-of-sight tracking system including one or more eyepieces 520, at least one eye tracking camera 540 (e.g., an infrared (IR) or near-infrared (NIR) camera) disposed on a side of a user's face where eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye(s) 592. The eye tracking camera 540 is positioned between the user's eye(s) 592 and a display 510 (e.g., a display panel on the left or right side of a head-mounted display, or a display of a handheld device, a projector, etc.), and may be directed toward a mirror 550 that reflects IR or NIR light from the eye(s) 592 while transmitting visible light (as shown, for example, at the top of FIG. 5), or may be directed toward the user's eye(s) 592 to receive the reflected IR or NIR light from the user's eye(s) 592 (as shown, for example, at the bottom of FIG. 5).

[0082] In some embodiments, the controller 110 renders an AR or VR frame 562 (e.g., the left and right frames of the left and right display panels) and provides the frame 562 to the display 510. The controller 110 uses the eye tracking input 542 from the eye tracking camera 540, for example, when processing the frame 562 for display, for various purposes. The controller 110 optionally estimates the user's viewpoint on the display 510 based on the eye tracking input 542 obtained from the eye tracking camera 540 using a glint assist method or other suitable method. The viewpoint estimated from the eye tracking input 542 is optionally used to determine the direction in which the user is currently looking.

[0083] Examples of several possible use cases of the user's current line of sight direction are described below, but this is not intended to be limiting. As an exemplary use case, the controller 110 can render virtual content differently based on the determined line of sight direction of the user. For example, the controller 110 may generate virtual content at a higher resolution in the central visual region determined from the user's current line of sight direction than in the peripheral region. As another example, the controller may position or move virtual content within the view based at least in part on the user's current line of sight direction. As another example, the controller may display specific virtual content within the view based at least in part on the user's current line of sight direction. As another exemplary use case in an AR application, the controller 110 can capture the physical environment of the CGR experience and direct the external camera to focus in the determined direction. The autofocus mechanism of the external camera can then focus on an object or surface within the environment that the user is currently viewing on the display 510. As another exemplary use case, the eyepiece 520 may be a focusable lens, and the line of sight tracking information is used by the controller to adjust the focus of the eyepiece 520 so that the virtual object the user is currently viewing has appropriate binocular convergence to match the convergence of the user's eyes 592. The controller 110 can utilize the line of sight tracking information to direct and adjust the focus of the eyepiece 520 so that the nearby object the user is viewing appears at the correct distance.

[0084] In some embodiments, the eye-tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eyepieces (e.g., eyepiece(s) 520), an eye-tracking camera (e.g., eye-tracking camera(s) 540), and a light source (e.g., light source 530 (e.g., IR LED or NIR LED)) attached to a wearable housing. The light source emits light (e.g., IR or NIR light) toward the user's eye(s) 592. In some embodiments, the light source may be arranged in a ring or circularly around each lens, as shown in FIG. 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520, by way of example. However, more or fewer light sources 530 may be used, and other arrangements and positions of the light sources 530 may be employed.

[0085] In some embodiments, the display 510 emits light within the visible light range and does not emit light within the IR or NIR range, so as not to introduce noise into the eye-tracking system. Note that the location and angle of the eye-tracking camera(s) 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 can be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.

[0086] Embodiments of the eye-tracking system as shown in FIG. 5 can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide the user with an experience of computer-generated reality, virtual reality, augmented reality, and / or augmented virtuality.

[0087] Figure 6 shows a glint-assisted gaze tracking pipeline according to some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a glint-assisted gaze tracking system (e.g., an eye tracking device 130 as shown in FIGS. 1 and 5). The glint-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in the tracking state, the glint-assisted gaze tracking system uses prior information from the previous frame when analyzing the current frame to track the pupil contour and glint within the current frame. When not in the tracking state, the glint-assisted gaze tracking system attempts to detect the pupil and glint within the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in the tracking state.

[0088] As shown in FIG. 6, the gaze tracking camera can capture left and right images of the user's left and right eyes. The captured images are then input into the gaze tracking pipeline for processing starting at 610. As indicated by the arrow returning to element 600, the gaze tracking system can continue to capture images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images may be input into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.

[0089] At 610, for the currently captured image, if the tracking state is yes, the method proceeds to element 640. At 610, if the tracking state is no, as shown at 620, the image is analyzed to detect the user's pupil and glint within the image. At 630, if the pupil and glint are successfully detected, the method proceeds to element 640. If not successfully detected, the method returns to element 610 to process the next image of the user's eyes.

[0090] At 640, when proceeding from element 410, the current frame is analyzed to track the pupil and the glint, based in part on look-ahead information from the previous frame. At 640, when proceeding from element 630, the tracking state is initialized based on the detected pupil and glint within the current frame. The result of the processing at element 640 is checked to confirm that the tracking or detection result is reliable. For example, the result can be checked to determine whether a sufficient number of glints for performing pupil and gaze estimation are successfully tracked or detected in the current frame. At 650, if the result is not reliable, the tracking state is set to no, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (or remains yes if already yes), and the pupil and glint information is passed to element 680 to estimate the user's viewpoint.

[0091] FIG. 6 is intended to function as an example of an eye tracking technique that can be used in a particular implementation. As will be recognized by those skilled in the art, other eye tracking techniques, whether currently existing or developed in the future, can be used in computer system 101 in place of, or in combination with, the glint-assisted eye tracking technique described herein to provide a CGR experience to a user according to various embodiments.

[0092] Accordingly, the description herein describes several embodiments of a three-dimensional environment (e.g., a CGR environment) that includes representations of real-world objects and representations of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table that exists in a physical environment that is captured and displayed in the three-dimensional environment (e.g., actively via a camera and display of an electronic device, or passively via a transparent or translucent display of the electronic device). As described above, the three-dimensional environment is optionally a mixed reality system based on a physical environment that is captured by one or more sensors of a device and displayed via a display generation component. As a mixed reality system, the device can optionally selectively display portions and / or objects of the physical environment such that each portion and / or object of the physical environment appears to exist in the three-dimensional environment that is displayed by the electronic device. Similarly, the device can optionally display virtual objects within the three-dimensional environment such that the virtual objects appear to exist in the real world (e.g., the physical environment) by placing the virtual objects at respective locations within the three-dimensional environment that have corresponding locations in the real world. For example, the device can optionally display a vase such that it appears as if an actual vase is placed on a table in the physical environment. In some embodiments, each location within the three-dimensional environment has a corresponding location within the physical environment. Thus, when the device is described as displaying a virtual object at an individual location relative to a physical object (e.g., a location near or on the user's hand, or a location near or on a physical table, etc.), the device displays the virtual object at a particular location within the three-dimensional environment such that the virtual object appears to be near or on the physical object of the physical world (e.g., the virtual object is displayed at a location within the three-dimensional environment that corresponds to the location within the physical environment where the virtual object would be displayed if the virtual object were a real object at that particular location).

[0093] In some embodiments, real-world objects present in the physical environment that are displayed in the three-dimensional environment can interact with virtual objects that exist only in the three-dimensional environment. For example, the three-dimensional environment can include a table and a vase placed on the table, where the table is a view (or representation) of the physical table in the physical environment and the vase is a virtual object.

[0094] Similarly, a user can optionally use one or more hands to interact with virtual objects in the three-dimensional environment as if the virtual objects were real objects in the physical environment. For example, as described above, one or more sensors of the device can optionally capture one or more of the user's hands and display a representation of the user's hands in the three-dimensional environment (e.g., in a manner similar to displaying real-world objects in the three-dimensional environment as described above), or in some embodiments, the user's hands can be seen through the display generating components via the ability to see the physical environment through the user interface due to transparency / semi-transparency of a portion of the display generating components displaying the user interface, or the projection of the user interface onto a transparent / semi-transparent surface, the projection of the user interface onto the user's eyes or the field of view of the user's eyes. Thus, in some embodiments, the user's hands are displayed at separate locations in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment that can interact with virtual objects in the three-dimensional environment as if they were actual physical objects in the physical environment. In some embodiments, the user can move their hands to cause a representation of their hands in the three-dimensional environment to move in tandem with the movement of the user's hands.

[0095] In some of the embodiments described below, the device can arbitrarily determine the "effective" distance between a physical object in the physical world and a virtual object in a three-dimensional environment for the purpose of determining, for example, whether a physical object is interacting with a virtual object (e.g., whether a hand is touching, grasping, holding a virtual object, etc., or is within a threshold distance from a virtual object). For example, the device determines the distance between the user's hand and a virtual object when determining whether the user is interacting with the virtual object and / or how the user is interacting with the virtual object. In some embodiments, the device determines the distance between the user's hand and a virtual object by determining the distance between the location of the hand in the three-dimensional environment and the location of the virtual object of interest in the three-dimensional environment. For example, one or more of the user's hands are placed at a particular position in the physical world, and the device optionally captures and displays at a corresponding particular position in the three-dimensional environment (e.g., the position in the three-dimensional environment where the hand would be displayed if the hand were a virtual hand rather than a physical hand). The position of the hand in the three-dimensional environment is optionally compared with the position of the virtual object of interest in the three-dimensional environment to determine the distance between one or more of the user's hands and the virtual object. In some embodiments, the device optionally determines the distance between a physical object and a virtual object by comparing positions in the physical world (e.g., as opposed to comparing positions in the three-dimensional environment). For example, when determining the distance between one or more of the user's hands and a virtual object, the device optionally determines the corresponding location of the virtual object in the physical world (e.g., the position where the virtual object would be placed in the physical world if the virtual object were a physical object rather than a virtual object), and then determines the distance between the corresponding physical position and one or more of the user's hands. In some embodiments, the same technique is optionally used to determine the distance between any physical object and any virtual object.Thus, as described herein, when determining whether a physical object is in contact with a virtual object or whether the physical object is within a threshold distance of the virtual object, the device optionally executes any of the techniques described above to map the location of the physical object into the three-dimensional environment and / or to map the location of the virtual object into the physical world.

[0096] In some embodiments, the same or similar techniques are used to determine where and in what direction the user's line of sight is directed and / or where and in what direction a physical stylus held by the user is directed. For example, if the user's line of sight is directed at a particular position within the physical environment, the device optionally determines the corresponding position within the three-dimensional environment, and if a virtual object is located at the corresponding virtual position, the device optionally determines that the user's line of sight is directed at that virtual object. Similarly, the device can optionally determine where in the physical world the stylus is pointing based on the orientation of the physical stylus. In some embodiments, based on this determination, the device determines the corresponding virtual position within the three-dimensional environment corresponding to the location in the physical world where the stylus is pointing and optionally determines that the stylus is pointing at the corresponding virtual position within the three-dimensional environment.

[0097] Similarly, the embodiments described herein may refer to the location of a user (e.g., a user of a device) and / or the location of a device in a three-dimensional environment. In some embodiments, the user of the device is holding, wearing, or otherwise located at or near the electronic device. Thus, in some embodiments, the location of the device is used as a proxy for the location of the user. In some embodiments, the location of the device and / or the user in the physical environment corresponds to an individual location in the three-dimensional environment. In some embodiments, the individual location is the location where the "camera" or "view" of the three-dimensional environment extends. For example, the location of the device is a location within the physical environment (and its corresponding location within the three-dimensional environment) from where, if the user were standing facing an individual portion of the physical environment being displayed by the display generation component, the user would see the objects within the physical environment in the same position, orientation, and / or size as that displayed by the display generation component of the device (e.g., from an absolute perspective and / or relative to each other). Similarly, if a virtual object displayed in the three-dimensional environment were a physical object within the physical environment (e.g., located at the same location within the same physical environment as the three-dimensional environment and having the same size and orientation within the same physical environment as the three-dimensional environment), the location of the device and / or the user would be the position from where the user would see the virtual object within the physical environment in the same position, orientation, and / or size as that displayed by the display generation component of the device (e.g., from an absolute perspective and / or relative to each other and to the objects in the real world).

[0098] In this disclosure, various input methods are described with respect to interaction with a computer system. If one example is provided using one input device or input method and another example is provided using another input device or input method, it should be understood that each example may be compatible with and optionally utilize the input device or input method described in another example. Similarly, various output methods are described with respect to interaction with a computer system. If one example is provided using one output device or output method and another example is provided using another output device or output method, it should be understood that each example may be compatible with and optionally utilize the output device or output method described in another example. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment via a computer system. If one example is provided using interaction with a virtual environment and another example is provided using a mixed reality environment, it should be understood that each example may be compatible with and optionally utilize the method described in another example. Accordingly, this disclosure discloses embodiments that are combinations of features of multiple examples without comprehensively listing all features of the embodiments in the description of each exemplary embodiment.

[0099] The processes described below enhance the operability of a device and streamline the user interface with the device by various techniques, including, for example, helping the user make appropriate inputs when operating / interacting with the device and reducing user errors, such as providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional controls being displayed, performing an operation without requiring further user input when a set of conditions is met, and / or other techniques. These techniques also reduce power usage and improve the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0100] Furthermore, in the methods described herein, conditioned on one or more conditions being satisfied by one or more steps, it should be understood that the described methods can be repeated in multiple iterations such that, over the course of the repetitions, all of the conditions conditioned on by the steps of the method are satisfied in different repetitions of the method. For example, if a method requires performing a first step when a condition is satisfied and a second step when the condition is not satisfied, one of ordinary skill in the art will understand that the steps recited in the claims will be repeated in a particular order until the condition is satisfied and then becomes not satisfied. Thus, a method described in terms of one or more steps that depend on one or more satisfied conditions can be rewritten as a method that is repeated until each condition described in the method is satisfied. However, this is not required for claims to a system or computer-readable medium that includes instructions for performing conditional operations based on the fulfillment of corresponding one or more conditions, and thus can determine whether an event is satisfied without explicitly repeating the steps of the method until all conditions for the steps of the method being conditional are satisfied. One of ordinary skill in the art will also understand that, similar to a method with conditional steps, a system or computer-readable storage medium can repeat the steps of the method as many times as necessary to ensure that all of the conditional steps are executed. User Interface and Related Processes

[0101] Attention is now directed to embodiments of a user interface (“UI”) and related processes that can be implemented in a computer system such as a portable multifunctional device or a head-mounted device that includes a display generation component, one or more input devices, and (optionally) one or more cameras.

[0102] Figures 7A - 7D illustrate examples of how an electronic device can enhance its interaction with selectable user interface elements, according to some embodiments.

[0103] FIG. 7A shows an electronic device 101 that displays a three-dimensional environment 702 on a user interface via a display generation component 120. In some embodiments, it should be understood that the electronic device 101 may utilize one or more of the techniques described with reference to FIGS. 7A-7D in a two-dimensional environment without departing from the scope of the present disclosure. As described above with reference to FIGS. 1-6, the electronic device 101 optionally includes a display generation component 120 (e.g., a touch screen) and a plurality of image sensors 314. The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor, and the electronic device 101 can be used to capture one or more images of the user or a part of the user while the user is interacting with the electronic device 101. In some embodiments, the display generation component 120 is a touch screen that can detect gestures and movements of the user's hand. In some embodiments, the user interface shown below can also be implemented on a head-mounted display that includes a display generation component that displays the user interface to the user, a sensor that detects the physical environment and / or the movement of the user's hand (e.g., an external sensor facing outward from the user), and / or a sensor that detects the user's line of sight (e.g., an internal sensor facing inward toward the user's face).

[0104] As shown in FIG. 7A, the three-dimensional environment 702 includes a dialog box 706 that includes text 710 and a plurality of selectable options 708a - d. In some embodiments, the electronic device 101 presents the three-dimensional environment 702 from the perspective of the user of the electronic device 101 within the three-dimensional environment 702. Thus, in some embodiments, the electronic device 101 displays one or more objects within the three-dimensional environment 702 at various distances from the perspective of the user within the three-dimensional environment 702 (e.g., at various z-heights). For example, the dialog box 706 is displayed with a shadow indicating the z-height of the dialog box relative to a reference frame in the three-dimensional environment 706 (e.g., the physical environment of the device 101). As another example, in FIG. 7A, the options 708a - 708d are displayed without a shadow, indicating that the options 708a - 708d are displayed at the same z-height as the rest of the dialog box 706 within the three-dimensional environment 702.

[0105] In some embodiments, in response to detecting a selection of one of the plurality of selectable options 708a - d, the electronic device 101 performs an action associated with the selected option. In some embodiments, the selectable options 708a - d are related to the text 710 included in the dialog box 706. For example, the text 710 describes a feature or setting having a plurality of available configurations, and the selectable options 708a - d are selectable to configure the electronic device 101 according to the individual configurations of the feature or setting described by the text 710.

[0106] In some embodiments, as described below with reference to FIGS. 7B - 7D, the electronic device 101 detects the selection of options 708a - d in response to user input including detection of the user's line of sight and detection of positions and / or gestures performed by the user's hand 704. For example, the electronic device 101, without detecting additional input, in response to detecting the user's line of sight directed at an individual option for a predetermined period (e.g., 0.2, 0.5, 1, 2 seconds, etc.), moves the hand with one or more fingers extended to a location corresponding to the 3D environment 702 corresponding to the option (e.g., in a pointing gesture), and the user "presses" an individual option by pushing it by a predetermined amount from the user, and / or in response to detecting that the user performs a predetermined gesture (e.g., touching the thumb to another finger of the same hand (e.g., index finger, middle finger, ring finger, little finger) (e.g., pinch gesture)) with the hand 704 while looking at an individual option, selects options 708a - d. In some embodiments, the electronic device 101 changes the appearance of an individual option while a selection is being detected. Changing the appearance of an individual option while a selection is being detected provides the user with feedback that a selection is being detected and allows the user to correct a selection error before the selection is made.

[0107] FIG. 7A shows a dialog box 706 while the electronic device 101 does not detect an input directed to any of the selectable options 708a - d. As described above, the input directed to the selectable options 708a - d optionally includes detecting the user's line of sight directed to one of the selectable options 708a - d and / or detecting the user's hand 704 at an individual gesture and / or position directed to one of the options 708a - d. In FIG. 7A, the line of sight of the user (not shown) is not directed to one of the selectable options 708a - d. The user's hand 704 in FIG. 7A is optionally not performing one of a predetermined gesture (e.g., a pointing gesture or a pinch gesture) and / or not in a position corresponding to the location of one of the selectable options 708a - d within the three - dimensional environment 702.

[0108] In FIG. 7B, the electronic device 101 detects the start of an input directed to option A 708a. For example, the electronic device 101 detects the user's line of sight 712 directed to option A 708a for a predetermined time threshold (e.g., 0.1, 0.2, 0.5, 1 second, etc.) without detecting an additional input (e.g., an input including detecting the user's hand 704). As another example, the electronic device 101 detects the user's hand 704 extended towards the location of option A 708a within the three-dimensional environment 702 with one or more (or all) fingers extended. In some embodiments, the user's hand 704 makes a pointing gesture (e.g., one or more extended fingers, but not all). In some embodiments, the user's hand 704 does not make a pointing gesture (e.g., all fingers of the hand 704 are extended). In some embodiments, the location of the hand is within a threshold distance (e.g., 1, 2, 10, 20, 30 centimeters, etc.) from the location corresponding to option 708a within the three-dimensional environment 702. In some embodiments, the hand makes the start of a pinch gesture, such as when the thumb and another finger are within a threshold distance (e.g., 0.5, 1, 2 centimeters, etc.) of each other. The electronic device 101 optionally simultaneously detects the user's line of sight 712 on option A 708a and the gesture and / or position of the hand 704 described above.

[0109] In response to detection of the line of sight 712 and / or gesture and / or position of the hand 704, the electronic device 101 gradually increases the z separation between option A 708a and the remainder of the dialog box 706 while the line of sight and / or hand gesture and / or position is being detected. In some embodiments, the electronic device 101 moves option 708a towards the user in the three-dimensional environment 702 and / or moves the remainder of the dialog box 706 away from the user in the three-dimensional environment 702. In some embodiments, the three-dimensional environment 702 includes a hierarchical level of a plurality of possible z heights from the perspective of the user within the three-dimensional environment 702. For example, in FIG. 7B, option A 708a is presented at a first hierarchical level and the remainder of the dialog box 706 is presented at a second (e.g., lower) hierarchical level. In some embodiments, the three-dimensional environment 702 includes additional objects at additional hierarchical levels at additional z heights from the user within the three-dimensional environment 702.

[0110] FIG. 7C shows the electronic device 101 detecting a selection of option A 708a. In some embodiments, in response to detecting a selection of option A 708a while displaying the three-dimensional environment 702 shown in FIG. 7B, the electronic device 101 decreases the z separation between option A 708a and the remainder of the dialog box 706. In some embodiments, when the z height of option A 708a reaches the z height of the remainder of the dialog box 706, the electronic device 101 selects option A 708a and updates the color of option A 708a.

[0111] In some embodiments, the selection of option A 708a is detected by detecting the user's line of sight 712 on option A 708a for a second threshold time (e.g., 0.2, 0.5, 1, 2 seconds, etc.) that is longer than the duration of the user's line of sight 712 in FIG. 7B without detecting an additional input (e.g., via the user's hand 704). In some embodiments, while detecting the line of sight 712 on option A 708a for a time longer than the time corresponding to FIG. 7B, the electronic device 101 gradually reduces the z separation between option A 708a and the rest of the dialog box 706. In some embodiments, when the line of sight 712 is detected on option A 708a for the second threshold time and the electronic device 101 displays option A 708a at the same z height as the dialog box 706, the electronic device 101 updates the color of option A 708a and executes an operation according to option A 708a.

[0112] In some embodiments, the selection of option A 708a is detected in response to detecting that the user performs a pinch gesture with hand 704 while the user's line of sight 712 is directed at option A 708a. In some embodiments, in response to detecting that the user performs a pinch gesture with hand 704 while detecting the line of sight 712 on option A 708a, the electronic device 101 reduces the z-separation between option A 708a and the rest of the dialog box 706 at a rate faster than the rate at which the electronic device 101 reduces the z-separation between option A 708a and the rest of the dialog box 706 in response to line-of-sight only input. In some embodiments, when option A 708a reaches the same z-height as the rest of the dialog box 706 (e.g., in response to the pinch gesture being maintained for a threshold time (e.g., 0.1, 0.2, 0.5 seconds, etc.)), the electronic device 101 updates the color 708a of option A 708a. In some embodiments, while displaying the three-dimensional environment 702 as shown in FIG. 7C, the electronic device 101 performs an action associated with option A 708a in response to detecting the end of the pinch gesture (e.g., the user releases the thumb from the other finger).

[0113] In some embodiments, the selection of option A 708a is detected by the user's hand 704 making a pointing gesture in response to detecting that the user "pressed" option A 708a, and moving the user's hand 704 from a location corresponding to the location of option A 708a in the three-dimensional environment 702 shown in FIG. 7B in a direction towards the remainder of the dialog box 706. In some embodiments, the electronic device 101 gradually decreases the z-separation between option A 708a and the remainder of the dialog box 706 at a predetermined rate while the user's line of sight 712 is directed towards option A 708a without a pinch gesture (or pointing gesture) being detected. In some embodiments, as described above, when the user "presses" option A 708a towards the remainder of the dialog box 706, the electronic device 101 updates the z-height of option A 708a in accordance with the movement of the user's hand 704 towards the remainder of the dialog box 706. For example, the speed and distance at which the electronic device 101 updates the z-position of option A 708a corresponds to the speed and distance of the movement of the user's hand 704 towards the remainder of the dialog box 706. In some embodiments, while maintaining the pointing gesture corresponding to the electronic device 101 displaying option A 708a at the same z-height as the remainder of the dialog box 706, the electronic device updates the color 708a of option A 708a in response to the movement of the user's hand 704. In some embodiments, after the electronic device 101 "pressed" option A 708a to the same z-height as the remainder of the dialog box 706 (e.g., the user stops performing the pointing gesture and the user moves the hand 704 away from the location corresponding to option A 708a), in response to detecting that the user has detected the end of "pressing" option A 708a, the electronic device performs an action associated with option A 708a.

[0114] In some embodiments, in response to an input that initially meets the selection criteria but ultimately does not meet the selection criteria, the electronic device 101 begins to reduce the z-separation of option A 708a shown in FIG. 7B, but does not select option A 708a as described above. For example, the electronic device 101 detects the user's line of sight 712 over a duration that is longer than the duration corresponding to the display of the three-dimensional environment 702 as shown in FIG. 7B but shorter than the duration corresponding to the display of the three-dimensional environment 702 as shown in FIG. 7C. As another example, the electronic device 101 detects a user who performs a pinch gesture with the hand 101 while the user's line of sight 712 is directed towards option 708a for an amount of time less than a predetermined threshold (e.g., 0.1, 0.2, 0.5, 1 second, etc.). As another example, the user "pushes" option A 708a by a distance that is less than the z-separation between option A 708a in FIG. 7B from the z-height shown in FIG. 7B. In some embodiments, in response to detecting an input corresponding to beginning to select option A 708a without actually selecting option A 708a, the electronic device 101 animates the separation between option A 708a and the rest of the dialog box 706 and increases it by inertia to the separation shown in FIG. 7B. For example, the electronic device 101 shows increasing the z-separation between option A 708a and the rest of the dialog box 706 following a deceleration of the z-separation between option A 708a and the rest of the dialog box 706.

[0115] In some embodiments, as shown in FIG. 7C, in response to detecting a selection input following satisfaction of the selection criteria while displaying the three-dimensional environment 702, the electronic device 101, as shown in FIG. 7D, gradually moves the dialog box 706, which includes the text 710 and the options 708a - d, away from the user within the three-dimensional environment 702 while maintaining the updated color of option A 708a. For example, the electronic device 101 detects a user who is maintaining a pinching gesture while the user's line of sight 712 is directed at option A 708a for a threshold time (e.g., 0.1, 0.2, 0.5, 1 second, etc.) longer than the threshold time corresponding to the selection of option A 708a. In some embodiments, in response to detecting that the user is maintaining the selection of the pinch gesture for option 708a, the electronic device 101 gradually moves the dialog box 706 away from the user at a predetermined speed (e.g., as time continues to elapse).

[0116] As another example, the electronic device 101 detects the user's "pushing" option A 708a beyond the z-height of the dialog box 706 shown in FIG. 7C by moving the hand 704 to a location corresponding to the z-height behind the dialog box 706 in FIG. 7C. In some embodiments, in response to detecting the user's "pushing" option A 708a beyond the z-height of the dialog box 706 in FIG. 7C, the electronic device 101 continues to move the dialog box 706 according to the speed and distance of the movement of the user's hand 704 within the three-dimensional environment 702, as shown in FIG. 7D.

[0117] In some embodiments, in response to detecting the end of an input that pushes the dialog box 706 away from the user within the three-dimensional environment 702, the electronic device 101 performs an action corresponding to option A 708a and returns the dialog box 706 to the position in the three-dimensional environment 702 shown in FIG. 7A (e.g., the original position of the dialog box 706 in the three-dimensional environment 702). In some embodiments, the dialog box 706 moves from the position in FIG. 7D to the position in FIG. 7A by inertia in a manner similar to the inertial movement of option A 708a described above with reference to FIG. 7B.

[0118] In some embodiments, the electronic device 101 displays additional user interface elements at the z-height behind the dialog box 706. For example, the dialog box 706 is displayed in a user interface that is displayed behind the dialog box 706 in the three-dimensional environment 702. In response to an input (e.g., a pinch gesture or a push gesture) that pushes the dialog box 706 beyond the user interface object behind it, the electronic device 101 optionally pushes back each of the objects behind the dialog box 706, the dialog box 706, and option 708a (and options 708b - d) according to additional input. In some embodiments, the user interface behind the dialog box 706 is the three-dimensional environment 702 itself, and the device 101 moves the three-dimensional environment 702 away from the user's perspective as if the user were moving in a direction opposite to the direction in which the three-dimensional environment 702 is being pushed.

[0119] In some embodiments, in response to detecting the end of an input that pushes an object behind the dialog box 706 away from the user in the three-dimensional environment 702, the electronic device 101 performs an action corresponding to option A 708a and returns the object dialog box 706 including the text 710 and options 708a - d to a position within the three-dimensional environment 702 before the start of the selection of option A 708a is detected (e.g., the state of the three-dimensional environment 702 shown in FIG. 7A). In some embodiments, the object behind the dialog box 706 moves from the position pushed away from the user to the initial position by inertia in a manner similar to the inertial movement of option A 708a described above with reference to FIG. 7B.

[0120] FIGS. 8A - 8M are flowcharts showing a method for automatically updating the orientation of virtual objects in a three-dimensional environment based on a user's perspective according to some embodiments. In some embodiments, method 800 is executed in a computer system (e.g., the computer system 101 of FIG. 1 such as a tablet, smartphone, wearable computer, or head-mounted device) including a display generation component (e.g., the display generation component 120 of FIGS. 1, 3, and 4) (e.g., a head-up display, a display, a touch screen, a projector, etc.) and one or more cameras (e.g., a camera facing downward with the user's hand (e.g., a color sensor, an infrared sensor, and other depth detection cameras), or a camera facing forward from the user's head). In some embodiments, method 800 is stored in a non-transitory computer-readable storage medium and is executed by instructions executed by one or more processors of the computer system, such as one or more processors 202 (e.g., the control unit 110 of FIG. 1A) of the computer system 101. Some operations of method 800 are optionally combined, and / or the order of some operations is optionally changed.

[0121] In some embodiments, such as in FIG. 7A, method 800 is executed on an electronic device that communicates with a display generation component and one or more input devices (e.g., a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer). In some embodiments, the display generation component is an integrated display (optionally a touch screen display) with the electronic device, an external display such as a monitor, projector, television, or a hardware component (optionally integrated or external) that projects a user interface or makes the user interface visible to one or more users. In some embodiments, the one or more input devices include an electronic device or component that can receive user input (e.g., capture user input, detect user input, etc.) and transmit information related to the user input to the electronic device. Examples of input devices include a touch screen, a mouse (e.g., external), a trackpad (optionally integrated or external), a touchpad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the electronic device), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, an eye tracking device, and / or a motion sensor (e.g., a hand tracking device, a hand motion sensor), etc. In some embodiments, the electronic device communicates with a hand tracking device (e.g., one or more cameras, a depth sensor, a proximity sensor, a touch sensor (e.g., a touch screen, a trackpad)). In some embodiments, the hand tracking device is a wearable device such as a smart glove. In some embodiments, the hand tracking device is a handheld input device such as a remote control or a stylus.

[0122] In some embodiments, such as FIG. 7A, an electronic device (e.g., 101) displays (802a) a user interface that includes an individual user interface element (e.g., 708a) having a first appearance via a display generation component. In some embodiments, the individual user interface element is displayed in a three-dimensional environment (e.g., a computer-generated reality (CGR) environment such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment) that is generated, displayed, or otherwise made visible by the device. In some embodiments, displaying the individual user interface element with the first appearance includes displaying the individual user interface element with a first size, color, and / or transparency, and / or displaying the individual user interface element within a first individual virtual layer of the user interface. In some embodiments, the three-dimensional user interface includes a plurality of virtual layers that create the appearance of various virtual distances between the various user interface elements and the user. For example, displaying the individual user interface element with the first appearance includes displaying selectable options (e.g., buttons, sliders indicating one of a plurality of possible slider positions optionally set by the user) within the same virtual layer as the background behind the individual user interface element with the first size and the first color.

[0123] In some embodiments, such as FIG. 7B, while displaying an individual user interface element (e.g., 708a) having a first appearance, an electronic device (e.g., 101) detects (802b) that the attention of the device's user is directed to the individual user interface element via one or more input devices based on the posture (e.g., position, orientation, and / or grip) of the user's physical characteristic (e.g., eye or hand). In some embodiments, the electronic device detects that the user has looked at an individual user interface element for a predetermined threshold time (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 1 second, etc.) via an eye tracking device, or detects that the user is looking at an individual user interface element without considering the length of time the user has looked at the individual user interface element. In some embodiments, the electronic device detects that the user's hand is within a predetermined location for a predetermined threshold time (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 1 second, etc.) via a hand tracking device. For example, the predetermined location is a location corresponding to the virtual location of an individual user interface element, such as a location where the user's hand appears to overlap or be within a threshold distance (e.g., 1, 3, 10 inches) of the individual user interface element in a VR environment, an MR environment, or an AR environment, and / or a location corresponding to the physical location where the individual user interface element is displayed by a display generation component (e.g., when the hand is hovering above a touch-sensitive display). In some embodiments, the electronic device detects movement of an input device and moves the input focus of the electronic device and / or a cursor displayed in the user interface to the location of an individual user interface element within the user interface.

[0124] In some embodiments, such as FIG. 7B, in response to detecting that the attention of a user of the device is directed to an individual user interface element (e.g., 708a), and in accordance with a determination that one or more first criteria are met, an electronic device (e.g., 101) updates (802c) an individual user interface element (e.g., 708a) to have a second appearance different from the first appearance, by visually separating the individual user interface element (e.g., 708a) from a portion of the user interface (e.g., 706) having a predetermined spatial relationship to the individual user interface element (e.g., near, adjacent to, laterally adjacent to, horizontally adjacent to, vertically adjacent to the individual user interface element). In some embodiments, visually separating an individual user interface element from a portion of the user interface includes increasing the z-separation between the individual user interface element and the portion of the user interface, for example, by displaying the individual user interface element near the location of the user in a three-dimensional environment and / or displaying the portion of the user interface far from the location of the user in a three-dimensional environment. In some embodiments, updating an individual user interface element to have a second appearance different from the first appearance includes updating the size, color, position, and / or translucency in which the individual user interface element is displayed, and / or updating the virtual layer of the user interface in which the individual user interface element is displayed. For example, in response to a first user input, the electronic device updates the individual user interface element to be displayed in a first color in the same layer of the user interface as the background behind the user interface element, to display the individual user interface element in a second color within a virtual layer of the user interface that is above (e.g., in front of) the virtual layer of the user interface in which the background behind the user interface is displayed.In this example, in response to a first input, an electronic device changes the color of an individual user interface element and reduces the virtual distance between the individual user interface element and the user, such that the individual user interface element (e.g., a button) appears to pop out in front of the background plane on which the individual user interface element is displayed.

[0125] In some embodiments, such as in FIG. 7B, while an individual user interface element (e.g., 708a) has a second appearance, an electronic device (e.g., 101) detects (802d) a second user input corresponding to the activation of the individual user interface element (e.g., 708a) via one or more input devices based on the posture (e.g., position, orientation, and / or grip) of a physical characteristic (e.g., an eye or a hand) of the user. In some embodiments, the electronic device detects the posture via an eye tracking device, a hand tracking device, a touch sensing surface (e.g., a touch screen or a track pad), a keyboard, or a mouse. For example, in response to detecting via an eye tracking device that the user is looking at an individual user interface element, the electronic device updates the position of the user interface element to move it to a second layer that appears closer to the user than a first layer. In this example, in response to detecting via a hand tracking device that the user has tapped together the thumb and a finger (e.g., an index finger, a middle finger, a ring finger, or a little finger) of the same hand, the electronic device updates the individual user interface element from being displayed in a second layer to being displayed in a layer (e.g., the first layer, a layer between the first and second layers, a layer behind the first layer, etc.) that appears further from the user than the second layer.

[0126] In some embodiments, in response to detecting a second user input directed to an individual user interface element (e.g., 708a) (802e), and in accordance with a determination that the second user input meets one or more second criteria, an electronic device (e.g., 101) performs a selection operation associated with the individual user interface element (e.g., 708a) (802f), and as shown in FIG. 7C, updates the individual user interface element (e.g., 708a) by reducing the amount of separation between the individual user interface element (e.g., 708a) and a portion of the user interface (e.g., 706) that has a predetermined spatial relationship to the individual user interface element (e.g., 708a). In some embodiments, the one or more second criteria include criteria that are met when the electronic device detects, using an eye-tracking device, that the user has looked at the individual user interface element for a threshold time (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5 seconds, etc.) that is arbitrarily longer than a time threshold associated with a first user input. In some embodiments, the one or more second criteria include criteria that are met when the electronic device, while a gesture is being performed, simultaneously detects, via an eye-tracking device, that the user is looking at the individual user interface element and, via a hand-tracking device, that the user has performed a predetermined gesture (e.g., by hand). In some embodiments, the predetermined gesture includes the user tapping together their thumb and one of their fingers (e.g., index finger, middle finger, ring finger, little finger). In some embodiments, the one or more second criteria are met when the electronic device, via a hand-tracking device, detects that the location of the user's hand or the user's finger of the hand corresponds to a predetermined location, such as a predetermined virtual location within the user interface.For example, one or more second criteria are met when the electronic device detects that the user has moved their hand from a virtual location within a user interface where individual user interface elements are displayed in a first virtual layer of the user interface to a virtual location within a second virtual layer of the user interface (e.g., a virtual location within the second virtual layer corresponding to a virtual location behind the virtual location within the user interface where the individual user interface element was displayed in the first virtual layer) using a hand tracking device. In some embodiments, one or more second criteria are met in response to detecting a lift-off of a selection input, such as the release of a hardware key or button of an input device (e.g., keyboard, mouse, trackpad, remote, etc.) or a lift-off of contact on a touch-sensitive surface (e.g., touch sensor display, trackpad, etc.). In some embodiments, the electronic device reduces the amount of separation between an individual user interface element and a portion of the user interface having a predetermined spatial relationship to the individual user interface element, but the electronic device displays the individual user interface element in a third size, a third color, and / or a third transparency, and / or displays the individual user interface element in a third virtual location or a third virtual layer of the user interface. In some embodiments, the third size, color, transparency, virtual location, and / or third virtual layer are different from a second size, color, transparency, virtual location, and / or second virtual layer corresponding to a second appearance of the individual user interface element. In some embodiments, the third size, color, transparency, virtual location, and / or third virtual layer are the same as a second size, color, transparency, virtual location, and / or second virtual layer corresponding to a second appearance of the individual user interface element. In some embodiments, the third size, color, transparency, virtual location, and / or third virtual layer are different from a first size, color, transparency, virtual location, and / or first virtual layer corresponding to a first appearance of the individual user interface element.In some embodiments, the third size, color, translucency, virtual location, and / or third virtual layer is the same as the first size, color, translucency, virtual location, and / or first virtual layer corresponding to the first appearance of the individual user interface element. For example, while reducing the separation between an individual user interface element and a portion of the user interface having a predetermined spatial relationship to the individual user interface element to display the individual user interface element, the electronic device displays the individual user interface element in a third color different from the first color and the second color in a first layer of the user interface where the individual user interface element is displayed in the first appearance. In some embodiments, updating an individual user interface element from a first appearance to a second appearance includes transitioning from a first virtual layer of the user interface, where the individual user interface element is displayed, to a second virtual layer of the user interface that appears closer to the user than the first layer. In some embodiments, a second user input that meets one or more second criteria corresponds to an input magnitude that moves the individual user interface element from the second virtual layer to the first virtual layer, and the electronic device displays an animation that moves the user interface element from the second virtual layer to the first virtual layer while the second input is being received. In some embodiments, the first input and the second input are detected by different input devices. For example, detecting the first user input includes detecting the user's line of sight on the individual user interface element via an eye tracking device, while detecting an individual hand gesture performed by the user via a hand tracking device, and detecting the second user input includes detecting the user's line of sight for a period exceeding a threshold (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5 seconds, etc.) without detecting the input by the hand tracking device.

[0127] In some embodiments, as shown in FIG. 7B, in response to detecting a second user input directed to an individual user interface element (e.g., 708a) (802e), while it is still determined that the user's attention is directed to the individual user interface element (e.g., 708a), in accordance with a determination that the second user input does not meet one or more second criteria, the electronic device (e.g., 101) cancels the execution of a selection operation associated with the individual user interface element (e.g., 708a) (802h) without reducing the amount of separation between the individual user interface element (e.g., 708a) and a portion of the user interface having a predetermined spatial relationship to the individual user interface element (e.g., 706). In some embodiments, in accordance with a determination that the second user input does not meet one or more second criteria, the electronic device continues to display the individual user interface element in a second appearance and continues to visually separate the individual user interface element from a portion of the user interface having a predetermined spatial relationship to the individual user interface element. In some embodiments, in accordance with a determination that the second user input does not meet one or more second criteria, the electronic device displays the individual user interface element in a first appearance. The methods described above of updating the individual user interface element to have a second appearance in response to a first user input and updating the individual user interface element to be displayed further away from a portion of the user interface having a predetermined spatial relationship to the individual user interface element in response to a second user input that meets one or more second criteria provide an efficient way to provide feedback to the user that the first and second user inputs have been received, thereby simplifying the interaction between the user and the electronic device, improving the operability of the electronic device, making the interface between the user and the device more efficient, and thereby reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0128] In some embodiments, as shown in FIG. 7B, individual user interface elements (e.g., 708a) have a second appearance, but the electronic device (e.g., 101) detects (804a) that the user's attention is not directed to the individual user interface element (e.g., 708b) via one or more input devices based on the posture (e.g., position, orientation, and / or grip) of the user's physical characteristics (e.g., eyes or hands). In some embodiments, the electronic device detects the user's line of sight directed to a location within the user interface other than the individual user interface element via an eye tracking device, and / or via a hand tracking device, the electronic device detects that the user has released their hand from a predetermined location associated with the individual user interface element. In some embodiments, in response to detecting that the user's attention of the device (e.g., 101) is not directed to the individual user interface element (e.g., 708a), the electronic device (e.g., 101) updates the individual user interface element (e.g., 708a) by reducing the amount of separation between the individual user interface element (e.g., 708a) and a portion of the user interface (e.g., 706) that has a predetermined spatial relationship to the individual user interface element (e.g., 708a) (804b). In some embodiments, the electronic device cancels the execution of a selection operation associated with the individual user interface element. In some embodiments, the electronic device displays the individual user interface element in the same virtual layer as a portion of the user interface that has a predetermined spatial relationship to the individual user interface element.

[0129] In response to detecting that the attention of a user of an electronic device is not directed to an individual user interface element, the method described above of reducing the separation amount between the individual user interface element and a portion of the user interface having a predetermined spatial relationship to the individual user interface element provides an efficient way to restore the appearance of the portion of the user interface having the predetermined spatial relationship to the individual user interface element without requiring additional user input, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0130] In some embodiments, such as FIG. 7C, the second user input satisfies one or more second criteria (806a). In some embodiments, as shown in FIG. 7B, while detecting a second user input directed to an individual user interface element (e.g., 708a) and before the second user input satisfies one or more second criteria, the electronic device (e.g., 101) reduces the amount of separation between the individual user interface element (e.g., 708a) and a portion of the user interface (e.g., 706) that has a predetermined spatial relationship to the individual user interface element (e.g., 708a) according to the progress of the second user input to satisfy one or more second criteria, thereby updating the individual user interface element (806b). In some embodiments, the electronic device displays an animation of the individual user interface element returning to the same virtual layer as the portion of the user interface that has a predetermined spatial relationship to the individual user interface element while the second user input is being detected. For example, in response to detecting the user's line of sight on an individual user interface element, the electronic device animates the individual user interface element to gradually move towards a portion of the user interface that has a predetermined spatial relationship to the individual user interface element while the user's line of sight is held on the individual user interface element, completes the animation, and performs a selection action according to a determination that a predetermined time (e.g., 0.1, 0.2, 0.3, 0.4, 0.5, 1, 2 seconds, etc.) has elapsed while the user holds the line of sight on the individual user interface element.

[0131] While detecting a second user input before the second user input meets one or more second criteria, reducing the separation amount between an individual user interface element and a portion of the user interface having a predetermined spatial relationship with the individual user interface element, the method described above provides an efficient way to indicate to the user the progress of selecting an individual user interface element, thereby enabling the user to use the electronic device more quickly and efficiently, further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use.

[0132] In some embodiments, such as FIG. 7B, detecting that the attention of the user of the device is directed to an individual user interface element (e.g., 708a) based on the posture of the physical characteristics of the user includes detecting that the line of sight of the user (e.g., 712) is directed to an individual user interface element (e.g., 708a) via an eye tracking device that communicates with the electronic device (e.g., 101) (808a). In some embodiments, the electronic device widens the interval between an individual user interface element and a portion of the user interface that includes the individual user interface element according to a determination that the line of sight of the user has been held on the individual user interface element for a predetermined period of time (e.g., 0.1, 0.2, 0.3, 0.4, 0.5, 1 second, etc.). In some embodiments, the electronic device begins to separate an individual user interface element from a portion of the user interface that includes the individual user interface element in response to detecting the line of sight of the user on the individual user interface element for any period (e.g., immediately upon detecting the line of sight). In some embodiments, the electronic device begins to separate an individual user interface element from a portion of the user interface only in response to a line of sight input (e.g., without receiving additional input via an input device other than the eye tracking device).

[0133] The above-described method of detecting a user's attention based on the line of sight provides an efficient way to initiate the selection of individual user interface elements without input other than the user's line of sight, thereby enabling the user to use the electronic device more quickly and efficiently, further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use.

[0134] In some embodiments, such as FIG. 7B, detecting that the attention of the user of the device is directed to an individual user interface element (e.g., 708a) based on the posture of the physical characteristics of the user includes detecting that the user's line of sight (e.g., 712) is directed to the individual user interface element and that the user's hand (e.g., 704) is in a predetermined posture (810a) (e.g., gesture, location, movement) via an eye tracking device and a hand tracking device that communicate with the electronic device. In some embodiments, the electronic device increases the spacing between an individual user interface element and the portion of the user interface that includes the individual user interface element in response to detecting a non-line-of-sight input while detecting the user's line of sight on the individual user interface element. In some embodiments, the non-line-of-sight input is a hand gesture or position detected via a hand tracking device. For example, the hand gesture is a finger of the hand extended towards a location within a three-dimensional environment corresponding to an individual user interface element. As another example, the hand gesture is the user's thumb moving towards a finger of the same hand (e.g., index finger, middle finger, ring finger, little finger). As another example, the electronic device detects a hand at a location in a three-dimensional environment corresponding to an individual user interface element. For example, in response to detecting the user's line of sight on an individual user interface element while the user is extending a finger towards the individual user interface element in a three-dimensional environment, the electronic device begins to separate the individual user interface element from the portion of the user interface that includes the individual user interface element.

[0135] The above-described method of detecting a user's attention based on the user's gaze and the user's hand gesture provides an efficient way to initiate the selection of individual user interface elements in an intuitive way for the user, thereby enabling the user to use the electronic device more quickly and efficiently, further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use.

[0136] In some embodiments, such as FIG. 7B, detecting a second user input corresponding to the activation of an individual user interface element (e.g., 708a) based on the pose of the user's physical characteristics includes detecting a portion of the user's hand (e.g., 704) of the electronic device at a location corresponding to the individual user interface element (e.g., 708a) via a hand tracking device communicating with the electronic device (812a). In some embodiments, detecting the second user input further includes detecting a predetermined gesture made by the hand (e.g., touching the thumb to the finger, extending one or more fingers in a pointing gesture). In some embodiments, detecting the second user input includes detecting the user's hand at a location within a threshold distance (e.g., 1, 2, 3, 5, 10, 15, 20, 30, 50 centimeters) of the location of the individual user interface element in a three-dimensional environment while one or more fingers of the hand are extended (e.g., pointing with one or more fingers). For example, detecting the second input includes detecting that the user "pushes" each user interface element with one or more fingers at a location within the three-dimensional environment corresponding to the individual user interface element. In some embodiments, detecting the second user input includes detecting the tips of one or more fingers of the user within a threshold distance (e.g., 1, 2, 3, 5, 10, 15, 20, 30, 50 centimeters) of the location of the individual user interface element in a three-dimensional environment via a hand tracking device, and then detecting finger / hand / arm movement towards the individual user interface element while remaining at the location corresponding to the individual user interface element.

[0137] The above-described method of detecting a second user input based on the location of a part of the user's hand is intuitive for the user and provides an efficient way to receive an input that does not require the user to operate a physical input device, thereby enabling the user to use the electronic device more quickly and efficiently, further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use.

[0138] In some embodiments, such as FIG. 7B, detecting a second user input corresponding to the activation of an individual user interface element (e.g., 708a) based on the pose of the user's physical characteristics includes detecting an individual gesture (e.g., touching the thumb to the finger, extending one or more fingers in a pointing gesture) performed by the user's hand (e.g., 704) of the electronic device while there is a line of sight (e.g., 712) of the user of the electronic device (e.g., 101) directed at the individual user interface element (e.g., 708a) via an eye-tracking device and a hand-tracking device that communicate with the electronic device (e.g., 101) (814a). In some embodiments, detecting the second user input further includes detecting a predetermined gesture (e.g., touching the thumb to the finger, extending one or more fingers in a pointing gesture) performed by the hand while the hand is in a predetermined location (e.g., a location within a threshold distance (e.g., 5, 10, 20, 30, 45 centimeters, etc.) from the individual user interface element in a three-dimensional environment) while detecting, via the eye-tracking device, that the user's line of sight is directed at the individual user interface element. For example, detecting the second input includes detecting that the user taps another finger (e.g., index finger, middle finger, ring finger, little finger) of the same hand as the thumb while the user is looking at the individual user interface element.

[0139] The above-described method of detecting a second user input based on the location of a part of the user's hand is intuitive for the user and provides an efficient way to receive an input without the user having to operate a physical input device, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0140] In some embodiments, such as FIG. 7A, before detecting a second user input directed to an individual user interface element (e.g., 708a), the individual user interface element (e.g., 708a) is displayed with an individual visual characteristic having a first value (e.g., other than the distance between the individual user interface element and the portion of the user interface including the individual user interface element), while the individual user interface element (e.g., 708a) is visually separated (816a) from a portion of the user interface (e.g., 706). In some embodiments, the individual visual characteristic is the size, color, transparency, etc. of the individual user interface element. In some embodiments, such as FIG. 7C, performing a selection operation associated with an individual user interface element (e.g., 708a) includes displaying the individual user interface element with an individual visual characteristic having a second value different from the first value, while the amount of separation between the individual user interface element (e.g., 708a) and the portion of the user interface decreases (e.g., 706) (816b) (e.g., while the individual user interface element is not separated from the portion of the user interface). For example, before detecting the second input, the electronic device displays the individual user interface element in a first color, and in response to the second user input (e.g., selection), the electronic device performs a selection action while displaying the individual user interface element in a second color different from the first color.

[0141] The above-described method of updating visual characteristics as part of a selection operation provides an efficient way to confirm the selection of individual user interface elements, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0142] In some embodiments, such as FIG. 7C, a second user input meets one or more second criteria if the user's line of sight (e.g., 712) of an electronic device (e.g., 101) directed at an individual user interface element (e.g., 708a) is for a period longer than a time threshold (818a) (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 1, 2 seconds, etc.). In some embodiments, one or more second criteria are met in response to the line of sight being directed at an individual user interface element for a period longer than the time threshold without additional non-line-of-sight input. In some embodiments, the electronic device gradually reduces the amount of separation between the individual user interface element and a portion of the user interface while the user's line of sight is maintained on the individual user interface element over the time threshold. In some embodiments, one or more second criteria are met based solely on the user's line of sight without detecting additional input via an input device other than the line-of-sight tracking device.

[0143] The above-described method of selecting an individual user interface element in response to the user's line of sight being directed at the individual user interface element for a period of time threshold provides an efficient way to select an individual user interface element without the user having to operate a physical input device, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0144] In some embodiments, while an individual user interface element (e.g., 708a) has a second appearance, an electronic device (e.g., 101) detects (820a) that the hand (e.g., 704) of a user of the electronic device is at an individual location corresponding to a location for interacting with the individual user interface element (e.g., 708a) via a hand tracking device that communicates with the electronic device, as shown in FIG. 7B. In some embodiments, the hand is within a threshold distance (e.g., 5, 10, 15, 20, 25, 30, 40 centimeters, etc.) of an individual user interface element within a three-dimensional environment while in a predetermined posture (e.g., one or more fingers are extended in a pointing gesture and the thumb is within a threshold (e.g., 0.5, 1, 2 centimeters, etc.) of another finger in a pinching gesture). In some embodiments such as FIG. 7C, in response to detecting that the hand (e.g., 704) of a user of the electronic device (e.g., 101) is at an individual location, the electronic device (e.g., 101) updates (820b) the individual user interface element (e.g., 708a) to further visually separate the individual user interface element from a portion of the user interface (e.g., 706) that has a predetermined spatial relationship to the individual user interface element. In some embodiments, the electronic device updates the individual user interface in response to the user's line of sight being on the individual user interface element and / or the user's hand being in a predetermined posture (e.g., one or more fingers "pointing" at the individual user interface element) while the hand is at the individual location. In some embodiments, further visually separating the individual user interface element from a portion of the user interface that has a predetermined spatial relationship to the individual user interface element includes one or more of moving the individual user interface element towards the user's perspective within the three-dimensional environment and / or moving the user interface away from the user's perspective within the three-dimensional environment.

[0145] In response to detecting the user's hand at an individual location corresponding to a location for interacting with an individual user interface element, updating the individual user interface to further visually separate the individual user interface element from a portion of the user interface having a predetermined spatial relationship to the individual user interface element, the method described above makes it easier for the user to select a user interface element using hand movements or gestures, thereby reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0146] In some embodiments, such as FIG. 7B, an individual user interface element having a second appearance (e.g., 708a) is associated with a first hierarchical level within the user interface, and a portion of the user interface having a predetermined spatial relationship to the individual user interface element (e.g., 706) is associated with a second hierarchical level different from the first hierarchical level (822a). In some embodiments, the portion of the user interface having a predetermined spatial relationship to the individual user interface element is displayed within a virtual container (e.g., user interface, backplane, etc.) having a third hierarchical level above a second hierarchical level above the first hierarchical level. In some embodiments, the hierarchical level defines the distance of each user interface element from the user's perspective in a three-dimensional environment (e.g., z-depth). For example, an individual user interface element is displayed between the user's perspective and a portion of the user interface having a predetermined spatial relationship to the individual user interface element. In some embodiments, the dynamic range of an individual user interface element extends from a first hierarchical level to a second hierarchical level in response to a second user input. In some embodiments, the hierarchical level is a navigation level. For example, the currently displayed user interface is at a first hierarchical level, and the user interface to which the electronic device navigates from the current user interface is at a second hierarchical level.

[0147] The above-described method of associating an individual user interface element with a second appearance and a first hierarchical level and associating an individual portion of the user interface having a predetermined spatial relationship to the individual user interface element with a second hierarchical level provides an efficient way to direct the user's attention to the individual user interface element for interaction, reducing the user's cognitive burden, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0148] In some embodiments, such as FIG. 7B, detecting a second user input includes detecting a user's hand input to the electronic device corresponding to the movement of an individual user interface element (e.g., 708a) returning towards a portion of the user interface (e.g., 706) via a hand-tracking device communicating with the electronic device (824a). In some embodiments, the individual user interface elements are displayed between the user's perspective in a three-dimensional environment and the portion of the user interface. In some embodiments, the electronic device moves the individual user interface elements according to the hand input. In some embodiments, the electronic device moves the individual user interface element towards the portion of the user interface in response to a hand input corresponding to pushing the individual user interface element towards the portion of the user interface (e.g., one or more fingers extended from the hand touch the location corresponding to the individual user interface element and push the individual user interface element towards the portion of the user interface away from the user, or touch within its threshold distance (e.g., 0.5, 1, 2, 3, 5, 10, 20, 30 centimeters, etc.)). In some embodiments, such as FIG. 7C, in response to detecting a second user input (e.g., including hand input), the electronic device (e.g., 101) updates an individual user interface element (e.g., 708a) to reduce the amount of separation between the individual user interface element (e.g., 708a) and the portion of the user interface (e.g., 706) (824b). In some embodiments, the electronic device reduces the amount of separation between the individual user interface element and the portion of the user interface according to the characteristics of the hand input (e.g., distance or speed of movement, duration). For example, in response to detecting a hand moving towards the portion of the user interface by a first amount, the electronic device reduces the interval between the individual user interface element and the portion of the user interface by a second amount.In this example, in response to detecting a hand moving toward a portion of the user interface by only a third amount that is greater than the first amount, the electronic device decreases the spacing between an individual user interface element and the portion of the user interface by only a fourth amount that is greater than the second amount. In some embodiments, such as FIG. 7C, the second user input satisfies one or more second criteria (824c) when the hand input corresponds to the movement of an individual user interface element (e.g., 708a) within a threshold distance (e.g., 0, 0.5, 1, 2, 3, 5, 10 centimeters, etc.) from a portion of the user interface (e.g., 706). In some embodiments, the one or more second criteria include criteria that are satisfied when the individual user interface element reaches the portion of the user interface in accordance with the hand input. In some embodiments, prior to detecting the second input, the individual user interface element is displayed at a first hierarchical level, the portion of the user interface is displayed at a second hierarchical level, and there is no hierarchical level between the first hierarchical level and the second hierarchical level (e.g., no other user interface elements are displayed).

[0149] The method described above of updating the separation amount between an individual user interface element and a portion of the user interface to satisfy one or more second criteria when the hand input corresponds to the movement of the individual user interface element within a threshold distance from the portion of the user interface provides an efficient way to provide feedback to the user while the user provides the second input, thereby reducing power usage, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0150] In some embodiments, such as in FIG. 7D, after a second user input meets one or more second criteria, while an individual user interface element (e.g., 708a) is within a threshold distance (e.g., 0, 0.5, 1, 2, 3, 5, 10 centimeters, etc.) from a portion of the user interface (e.g., 706), the electronic device (e.g., 101) detects (826a) further hand input from the user of the electronic device corresponding to the movement of the individual user interface element (e.g., 708a) towards the portion of the user interface (e.g., 706) via the hand tracking device. In some embodiments, the hand movement continues beyond the amount by which one or more second criteria are met. In some embodiments, such as in FIG. 7D, in response to detecting further hand input, the electronic device moves (826b) the individual user interface element (e.g., 708a) and the portion of the user interface (e.g., 706) according to the further hand input (e.g., without changing the amount of separation between the individual user interface element and the portion of the user interface). In some embodiments, when an object at a first hierarchical level (e.g., an individual user interface element) is pushed against an object at a second hierarchical level (e.g., a portion of the user interface), the objects at both levels move together in response to further user input (e.g., as if both of them were included in the second hierarchical level). In some embodiments, further hand input causes the electronic device to push the individual user interface element and the portion of the user interface further to a third hierarchical level after the second hierarchical level (e.g., the second hierarchical level is between the first hierarchical level and the second hierarchical level). In some embodiments, the electronic device displays a backplane of the portion of the user interface at the third hierarchical level. In some embodiments, the electronic device returns the third hierarchical level to a fourth hierarchical level in response to further hand input, such as further pushing the first, second, and third hierarchical levels. In some embodiments, the movement of the individual user interface element and the portion of the user interface by further hand input is based on the speed, direction, and / or distance of the hand input movement.In some embodiments, in response to detecting hand movement towards the user's torso, the electronic device displays user interface elements and portions of the user interface that move towards the user in a three-dimensional environment with inertia.

[0151] The above-described method of moving individual user interface elements and portions of the user interface in accordance with further hand inputs in response to further hand inputs received while the individual user interface elements are within a threshold of the portion of the user interface provides an efficient way to confirm receipt of further hand inputs, reduces the user's cognitive burden, thereby further reducing power usage, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0152] In some embodiments, such as FIG. 7C, in response to detecting a second user input (828a), in accordance with a determination that the movement of an individual user interface element (e.g., 708a) corresponds to a return towards a portion of the user interface (e.g., 706) where the manual input is less than a threshold movement amount (e.g., 0.5, 1, 2, 3, 4, 5, 7, 10, 20 centimeters, etc., or the distance between an individual user interface element and a portion of the user interface), the electronic device (e.g., 101) moves the individual user interface element (e.g., 708a) towards the portion of the user interface (e.g., 706) according to the manual input (e.g., by an amount proportional to the following metric (e.g., duration, speed, distance)) without moving the portion of the user interface (e.g., 706) and reduces the separation amount between the individual user interface element (e.g., 708a) and the portion of the user interface (e.g., 706). In some embodiments, in response to a manual input corresponding to less than the threshold movement amount, the electronic device reduces the separation between the individual user interface element and the portion of the user interface without moving the portion of the user interface according to the manual input. In some embodiments, the electronic device moves the individual user interface element by an amount proportional to the metric of the manual input (e.g., duration, speed, distance). For example, in response to detecting a hand movement by a first amount, the electronic device moves the individual user interface element towards the portion of the user interface by a second amount. As another example, in response to detecting a hand movement by a third amount greater than the first amount, the electronic device moves the individual user interface element towards the portion of the user interface by a fourth amount greater than the second amount. In some embodiments, the first threshold corresponds to the movement amount at which the individual user interface element reaches the hierarchical level of the portion of the user interface. In some embodiments, the manual input satisfies one or more second criteria in response to a manual input corresponding to the threshold movement amount.In some embodiments, such as FIG. 7D, in response to detecting a second user input (828a), according to a determination that the movement of the hand input is greater than a threshold movement amount (e.g., less than a second threshold (e.g., 1, 2, 3, 4, 5, 7, 10, 20, 30, 40 centimeters, etc.) corresponding to a hierarchical level behind the hierarchical level of the portion of the user interface), the electronic device (e.g., 101) moves an individual user interface element (e.g., 708a) back towards a portion of the user interface (e.g., 706), and moves the portion of the user interface (e.g., 706) according to the movement of the hand (e.g., without changing the separation amount between the individual user interface element and the portion of the user interface). In some embodiments, when an object at the first hierarchical level (e.g., an individual user interface element) is pushed at least a threshold amount such that it is pushed against an object at the second hierarchical level (e.g., a portion of the user interface), the objects at both levels move together in response to a further user input (e.g., as if both of them were included in the second hierarchical level). In some embodiments, a further hand input (e.g., exceeding the second threshold) causes the electronic device to push the individual user interface element and the portion of the user interface further to a third hierarchical level behind the second hierarchical level (e.g., the second hierarchical level is between the first hierarchical level and the second hierarchical level). In some embodiments, the electronic device displays a backplane of a portion of the user interface at the third hierarchical level. In some embodiments, the electronic device returns the third hierarchical level to a fourth hierarchical level in response to a further user input, such as further pushing the first, second, and third hierarchical levels.

[0153] The above-described method of moving a portion of the user interface according to the manual input in response to the manual input exceeding a threshold provides an efficient way to provide feedback to the user when the manual input exceeds the threshold, thereby reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0154] In some embodiments, such as those of FIG. 7C, updating an individual user interface element (e.g., 708a) by reducing the amount of separation between an individual user interface (e.g., 708a) and a portion of the user interface (e.g., 706) includes moving the individual user interface element (e.g., 708a) and the portion of the user interface (e.g., 706) with inertia (e.g., simulated physical properties based on the speed of movement of the individual user interface element and the portion of the user interface) according to the movement component of a second user input (830a). In some embodiments, in response to detecting that the user stops the movement component of the second user input, or in response to detecting that the movement component of the second user input changes from moving towards the individual user interface element and the portion of the user interface to moving away from the individual user interface element and the portion of the user interface, the electronic device animates the continued progression, but deceleration, of the individual user interface element and the portion of the user interface and reduces the separation. For example, the second user input includes the movement of the user's hand towards the individual user interface element and the portion of the user interface. In this example, in response to the second input, the electronic device moves the individual user interface away from the user's perspective towards the portion of the user interface, and in response to detecting that the user has stopped moving their hand towards the individual user interface element and the portion of the user interface, even if the electronic device decelerates the movement of the individual user interface element towards the portion of the user interface, the individual user interface element continues to move towards the portion of the user interface with inertia for a period of time (e.g., 0.1, 0.3, 0.5 seconds). In some embodiments, when the portion of the user interface moves away from the user by the second input, the portion of the user interface also moves with inertia (e.g., continuing to move and decelerate for a period of time after the second input stops moving in the direction from the individual user interface element to the portion of the user interface).In some embodiments, if the second input does not move a portion of the user interface, after the second input stops moving in a direction from an individual user interface element to the portion of the user interface, the portion of the user interface does not move. In some embodiments, the electronic device (e.g., 101) detects (830b) the end of a second user input directed to an individual user interface element (e.g., 708a) of FIG. 7C. In some embodiments, the user stops looking at an individual user interface element, the user stops moving a hand towards an individual user interface element, the user stops performing a predetermined gesture (e.g., releasing a pinch gesture), and the user releases a hand from a predetermined location associated with the location of the individual user interface element, etc. In some embodiments, the end of the second user input is detected without the second user input satisfying one or more second criteria. In some embodiments, in response to detecting the end of a second user input directed to an individual user interface element (e.g., 708a) of FIG. 7C, the electronic device (e.g., 101) moves the individual user interface element (e.g., 708a) and the portion of the user interface (e.g., 706) in a direction opposite to the movement of the individual user interface element (e.g., 708a) and the portion of the user interface (e.g., 706) according to the second user input (830c). In some embodiments, in response to detecting the end of the second user input, the electronic device increases the spacing between the individual user interface element and the portion of the user interface. In some embodiments, one or more individual criteria include criteria that are satisfied when the second input does not satisfy one or more second criteria, when the user continues to look at an individual user interface element, and / or until an additional input directed to a different user interface element is received. In some embodiments, if the second input moves a portion of the user interface away from the user, the portion of the user interface moves in a direction towards the individual user interface element in response to the end of the second input.In some embodiments, an individual user interface element moves at a speed faster than and / or for a longer duration than a portion of the user interface moves such that the distance between the individual user interface element and the portion of the user interface increases in response to detection of the end of a second input. In some embodiments, if the second input does not cause the portion of the user interface to move, the portion of the user interface does not move after the electronic device detects the end of the second user input.

[0155] In response to detecting the end of a second user input, the method of moving an individual user interface element by inertia and moving the individual user interface element and a portion of the user interface in a direction opposite to the direction in response to the second user input as described above provides an efficient way to indicate to the user that the second user input did not meet one or more second criteria, thereby further reducing power usage, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0156] In some embodiments, such as FIG. 7B, detecting a second user input includes detecting a portion of a user's hand (e.g., 704) of an electronic device (e.g., 101) at a location corresponding to an individual user interface element (e.g., 708a) (832a) (e.g., detecting that the user "presses" an individual user interface element with one or more fingers of the hand). In some embodiments, while an individual user interface element (e.g., 708a) has a second appearance, the electronic device (e.g., 101), as in FIG. 7B, detects an individual input including an individual gesture (e.g., a pinch gesture in which the user touches the thumb with another finger (e.g., index finger, middle finger, ring finger, little finger) with the thumb hand) performed by the user's hand (e.g., 704) while the hand is in a location not corresponding to the individual user interface element (e.g., 708a) via a hand tracking device that communicates with the electronic device (832b). In some embodiments, the gesture is performed while the hand is not in a location within a three-dimensional environment corresponding to an individual user interface element (e.g., a location where the hand cannot "press" the individual user interface element to select it, rather a location away from the location of the individual user interface element). In some embodiments, the gesture is performed while the hand is in a location corresponding to an individual user interface element, and the response of the electronic device is the same as if the user had performed the gesture with the hand in a location not corresponding to the individual user interface element.In some embodiments, in response to the detection of an individual input (832c), based on a determination according to an individual gesture performed by the user's hand (e.g., 704) while the user's hand (e.g., 704) is in a location that does not correspond to an individual user interface element (e.g., 708a), if the individual input meets one or more third criteria, the electronic device (e.g., 101) updates the individual user interface element (832d) by reducing the separation amount between the individual user interface element (e.g., 708a) and a portion of the user interface, including moving the individual user interface element (e.g., 708a) and the inertia (e.g., 706) to move the individual user interface element (e.g., 708a) and a portion of the user interface as shown in FIG. 7D. In some embodiments, moving the individual user interface element and a portion of the user interface by inertia includes gradually increasing the moving speed of the individual user interface element and the portion of the user interface when the individual input is received, and gradually decreasing the moving speed of the individual user interface element in response to the end of the individual input. In some embodiments, in response to the detection of an individual input (e.g., the user performs a pinch gesture while looking at an individual user interface element), the individual user interface element and / or a portion of the user interface move towards each other, reducing the separation between the individual user interface element and the portion of the user interface. In some embodiments, the electronic device (e.g., 101) detects the end of the individual input (832e). In some embodiments, the user stops performing a gesture such as separating the thumb from the other fingers, or the hand tracking device stops detecting the user's hand because the user has removed the hand from the detection area of the hand tracking device.In some embodiments, in response to detecting the end of an individual input, an electronic device (e.g., 101) moves an individual user interface element (e.g., 708a) and a portion of the user interface (e.g., 706) in a direction opposite to the movement of the individual user interface element (e.g., 708a) and the portion of the user interface corresponding to the individual input (832f). In some embodiments, in response to detecting the end of an individual user input, the electronic device widens the gap between the individual user interface element and the portion of the user interface. In some embodiments, prior to detecting the end of an individual input, in accordance with a determination that the individual input meets one or more second criteria, the electronic device performs a selection operation associated with the individual user interface element and updates the individual user interface element by reducing the amount of separation between the individual user interface element and the portion of the user interface. In some embodiments, the electronic device responds to an individual input in the same manner as the electronic device responds to a second input.

[0157] The method described above of moving an individual user interface element by inertia in response to detecting the end of an individual user input and moving the individual user interface element and the portion of the user interface in a direction opposite to the direction corresponding to the individual input provides an efficient way to indicate to the user that the individual user input did not meet one or more second criteria, thereby further reducing power usage, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0158] In some embodiments, such as FIG. 7B, detecting a second user input includes detecting a portion of a user's hand (e.g., 704) of an electronic device (e.g., 101) at a location corresponding to an individual user interface element (e.g., 708a) (834a) (e.g., detecting that the user "presses" an individual user interface element with one or more fingers of the hand). In some embodiments, such as FIG. 7B, while an individual user interface element (e.g., 708a) has a second appearance, the electronic device (e.g., 101) detects an individual input including the user's line of sight (e.g., 712) directed at the individual user interface element (e.g., 708a) via an eye tracking device that communicates with the electronic device (834b). (e.g., while the user's hand is at a location not corresponding to an individual user interface element). In some embodiments, the electronic device detects the user's line of sight directed at the individual user interface element while detecting the user's hand at a location corresponding to the individual user interface element, and the response of the electronic device is the same as when the user's line of sight is detected on the individual user interface element by the hand at a location not corresponding to the individual user interface element. In some embodiments, such as FIG. 7B, in response to detecting an individual input (834c), the electronic device (e.g., 101) updates the individual user interface element (e.g., 708a) by reducing the amount of separation between the individual user interface element (e.g., 708a) and a portion of the user interface (e.g., 706), including moving the individual user interface element (e.g., 708a) and the portion of the user interface (e.g., 706) inertially, according to a determination based on the user's line of sight (e.g., 712) directed at the individual user interface element (e.g., 708a) (e.g., based on one or more parameters such as the direction and duration of the line of sight) that the individual input meets one or more third criteria (e.g., the user's line of sight is held on the individual user interface element for a predetermined time threshold (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5 seconds, etc.)) (834d).In some embodiments, moving individual user interface elements and portions of the user interface by inertia includes gradually increasing the speed of movement of the individual user interface elements and portions of the user interface when an individual input is received, and gradually decreasing the speed of movement of the individual user interface elements in response to the end of the individual input. In some embodiments, in response to the detection of an individual input, the individual user interface elements and / or portions of the user interface move towards each other, reducing the separation between the individual user interface elements and the portions of the user interface. In some embodiments, an electronic device (e.g., 101) detects the end of an individual input (834e) (e.g., detects the user's line of sight directed away from an individual user interface element, and / or detects that the user's eyes are closed for a predetermined time threshold (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 0.7, 1 second, etc.)). In some embodiments, in response to the detection of the end of an individual input, the electronic device (e.g., 101) moves the individual user interface element (e.g., 708a) and the portion of the user interface (e.g., 706) in a direction opposite to the movement of the individual user interface element (e.g., 708a) and the portion of the user interface corresponding to the individual input (834f). In some embodiments, in response to the detection of the end of an individual user input, the electronic device increases the spacing between the individual user interface element and the portion of the user interface. In some embodiments, before the end of an individual input is detected, in accordance with a determination that the individual input meets one or more second criteria, the electronic device performs a selection operation associated with the individual user interface element and updates the individual user interface element by reducing the amount of separation between the individual user interface element and the portion of the user interface. In some embodiments, the electronic device responds to an individual input in the same way the electronic device responds to a second input.

[0159] In response to detecting the end of an individual user input, moving an individual user interface element by inertia and moving the individual user interface element and a portion of the user interface in a direction opposite to the direction in response to the individual user input, the above-described method provides an efficient way to indicate to the user that the individual user input did not meet one or more second criteria, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0160] Figures 9A-9E illustrate examples of how an electronic device enhances its interaction with a slider user interface element according to some embodiments.

[0161] FIG. 9A shows an electronic device 101 that displays a three-dimensional environment 902 on a user interface via a display generation component 120. In some embodiments, it should be understood that the electronic device 101 may implement one or more of the techniques described herein with reference to FIGS. 9A-9E in a two-dimensional environment without departing from the scope of the present disclosure. As described above with reference to FIGS. 1-6, the electronic device 101 optionally includes a display generation component 120 (e.g., a touch screen) and a plurality of image sensors 314. The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor, and the electronic device 101 can be used to capture one or more images of the user or a part of the user while the user is interacting with the electronic device 101. In some embodiments, the display generation component 120 is a touch screen that can detect gestures and movements of the user's hand. In some embodiments, the user interface shown below can also be implemented on a head-mounted display that includes a display generation component that displays the user interface to the user, a sensor that detects the physical environment and / or the movement of the user's hand (e.g., an external sensor facing outward from the user), and / or a sensor that detects the user's line of sight (e.g., an internal sensor facing inward toward the user's face).

[0162] FIG. 9A shows an electronic device 101 that displays a dialog box 906 or a control element in a three-dimensional environment 902. The dialog box 906 includes a slider user interface element 908 having an indicator 910 of the current input state of the slider user interface element 908. The indicator 910 of the current input state of the slider user interface element 908 includes a selectable element 912 that, when selected, initiates one of the ways to change the current input state of the slider on the electronic device 101. As shown in FIG. 9A, the slider user interface element 908 controls the current volume level of the electronic device 101. In some embodiments, the electronic device 101 presents a slider user interface element similar to the slider 908 that controls other settings and / or operations of the electronic device 101.

[0163] As shown in FIG. 9A, the slider user interface element 908 is displayed without showing other available input states of the slider user interface element 908 other than the cursor or the current input state indicator 910. In some embodiments, the electronic device 101 detects that the user's line of sight has moved away from the slider user interface element 908 and / or the dialog box 906 (e.g., the user's line of sight is directed to another part of the user interface), and / or that the user's hand 904 is in a location that does not correspond to the location of the slider user interface element 908 within the three-dimensional environment 902, and presents the slider user interface element 908 as shown in FIG. 9A. As will be described in more detail below, in some embodiments, the electronic device 101 updates the slider user interface element 908 in response to an input that includes detecting a user who performs a predetermined gesture with the hand 904. In FIG. 9A, the user does not perform a predetermined gesture with the hand 904.

[0164] FIG. 9B shows an electronic device 101 that initiates a process of changing the current input state of a slider user interface element 908, such as by displaying the cursor 916 and / or indications 914a - g of the available input states of the slider user interface element 908 in response to detection of the user's line of sight on the dialog box 906 or the slider user interface element 908. In some embodiments, it should be understood that the electronic device 101 displays the indications 914a - g of the available input states without displaying the cursor 916. In some embodiments, the electronic device 101 displays the cursor 916 without displaying the indications 914a - g. In some embodiments, the electronic device 101 displays the cursor 916 and / or the indications 914a - g in response to detection of the user's line of sight 918a at any location within the dialog box 906. In some embodiments, the electronic device 101 does not display the cursor 916 and / or the indications 914a - g unless the user's line of sight 918b is directed towards the slider user interface element 908. In some embodiments, the electronic device 101 presents a display of the location of the user's line of sight while facilitating interaction with the slider user interface element 908 shown in FIGS. 9A - 9E. For example, displaying an indication of the line of sight includes increasing the brightness or lightness of the region of the three - dimensional environment 902 where the user's line of sight is detected.

[0165] In some embodiments, the electronic device 101 displays the cursor 916 at a location along the slider 908 based on one or more of the user's line of sight 918b and / or the location of the user's hand 904. For example, the electronic device 101 first displays the cursor 916 at a location along the slider user interface element 908 towards which the user's line of sight 918b is directed. In some embodiments, the electronic device 101 updates the location of the cursor 916 in response to detection of movement of the user's hand 904. For example, in response to detection of movement of the hand to the left, the electronic device 101 moves the cursor 916 to the left along the slider 908, and in response to detection of movement of the hand 904 to the right, the electronic device 101 moves the cursor 916 to the right along the slider 908. As will be described below with reference to FIG. 9E, in response to detection that the user has performed a predetermined gesture by touching the hand 904 (e.g., the thumb) with another finger of the same hand (e.g., index finger, middle finger, ring finger, little finger) (e.g., pinch gesture), the electronic device 101 moves the current input indicator 12 of the slider and the cursor 916 in accordance with further movement of the hand 904 while the hand is maintaining the pinch gesture. In some embodiments, while the pinch gesture is not detected, the electronic device 101 moves the cursor 916 in accordance with the movement of the hand without updating the current input state of the slider. In some embodiments, the cursor 916 moves only in the dimension towards which the slider user interface element 908 is directed (e.g., horizontally in FIG. 9B, vertically in a vertical slider user interface element). Thus, in some embodiments, the electronic device 101 moves the cursor 916 in accordance with the horizontal component of the movement of the user's hand 904 regardless of the vertical component of the movement of the user's hand 904. Further details regarding the cursor 916 are described below with reference to FIG. 9E.

[0166] In some embodiments, while indications 914a - g are being displayed, electronic device 101 updates the current input state of slider 908 in response to an input directed at indication 910 of the current input state of slider 908. As shown in FIG. 9B, electronic device 101 detects user's line of sight 918z on end 912 of indication 910 of the current input state of slider 908. In some embodiments, while user's line of sight 918z is detected at end 912 of slider user interface element 908, in response to detecting a user who performs a predetermined gesture with hand 904, electronic device 101 starts a process of moving current input state indicator 910 of the slider according to the movement of hand 904 while the gesture is being maintained. For example, the predetermined gesture is that the user touches the thumb with another finger of the same hand as the thumb (e.g., index finger, middle finger, ring finger, little finger) (e.g., pinch gesture). In some embodiments, in response to detection of a pinch gesture, electronic device 101 stops displaying indications 914a - g. In some embodiments, electronic device 101 continues to display indications 914a - g while, in response to detection of a pinch gesture and while the user moves current input state indicator 910 of slider user interface element 908 according to the movement of hand 904 while maintaining the pinch gesture. In some embodiments, current input state indicator 910 of slider user interface element 908 snaps to one of indications 914a - g. In some embodiments, it is possible to move current input state indicator 910 of slider user interface element 908 to a location between indications 914a - g. In some embodiments, in response to detecting that the user has stopped performing the pinch gesture with hand 904 (e.g., detecting that the thumb and finger are separated from each other), electronic device 101 updates the current input state of slider user interface element 908 and maintains the display of current input state indicator 910 of slider user interface element 908 corresponding to the updated state.

[0167] In some embodiments, as shown in FIGS. 9C-9D, the electronic device 101 updates the current input state of the slider user interface element 908 and moves the indication 910 of the current input state of the slider in response to detecting a selection of one of the other input state indications 914a-g of the slider. FIG. 9C shows the electronic device 101 detecting a selection of one of the input state indications 914e of the slider user interface element 908. For example, in some embodiments, while displaying the slider user interface element 908, the electronic device 101 detects the user's line of sight 918c directed at one of the input state indications 914e of the slider user interface element 908. In response to detecting the user's line of sight 908c directed at the indication 914e, the electronic device 101 gradually increases the size of the indication 914e.

[0168] When the line of sight 918c is detected for a predetermined threshold time (e.g., 0.1, 0.2, 0.5, 1, 5, 10, 30 seconds, etc.), the electronic device 101 updates the current input state of the slider user interface element 908 and the indication 910 of the current input state to the location corresponding to the indication 914e, as shown in FIG. 9D. In some embodiments, while the user's line of sight 918c is directed at the indication 914e for less than the threshold time, the electronic device 101 detects that the user is performing a pinch gesture with the hand 904. In some embodiments, in response to detecting a pinch gesture performed by the user's hand 904 while the user's line of sight 918c is directed at the indication 914e, the electronic device 101 updates the current input state of the slider to correspond to the indication 914e, as shown in FIG. 9D, regardless of the time the line of sight 918c is detected on the indication 914e.

[0169] In some embodiments, in response to detecting the user's line of sight 918c directed at the instruction 914e in FIG. 9C, while displaying the instruction 914e in a size larger than the other instructions 914a-d and 914f-g for a time less than a threshold time, the electronic device 101 detects the user's line of sight for a different one of the instructions 914a-d or 914f-g. For example, the electronic device 101 detects the user's line of sight directed at the instruction 914f. In this example, in response to detecting the user's line of sight directed at the instruction 914f, the electronic device 101 updates the instruction 914e to be the same size as the instructions 914a-d and 914g that the user is not looking at, and gradually increases the size of the instruction 914f. In some embodiments, the electronic device 101, as described above, in response to the line of sight continuing to be directed at the instruction 914f (e.g., for only the aforementioned time threshold), or regardless of whether the line of sight has been held for the time threshold, updates the current input state of the slider to correspond to the instruction 914f in response to detecting a pinch gesture while the line of sight is held at the instruction 914f. It should be understood that the electronic device 101 behaves similarly in response to detecting the user's line of sight on any of the other instructions 914a-914d and 914g.

[0170] FIG. 9E shows a user updating the current input state of the slider user interface element 908 while the electronic device 101 is displaying the cursor 916. As described above with reference to FIG. 9B, while the user's line of sight is directed at the end 912 of the slider user interface element 908, as shown in FIG. 9B, the electronic device 101, in response to detecting a pinch gesture, maintains the pinch gesture while updating the position of the cursor 916 along the slider user interface element 908 in accordance with the movement of the user's hand 904. In FIG. 9E, the electronic device 101 moves the indicator 910 of the current input state of the slider user interface element 908 using the cursor 916 in accordance with the movement of the user's hand 904 while maintaining the pinch gesture. In some embodiments, the electronic device 101 stops displaying the instructions 914a - g shown in FIGS. 9B - 9D in response to detecting a pinch gesture. In some embodiments, the instructions 914a - g continue to be displayed while the user is manipulating the current input state indicator 910 of the slider user interface element 908 while moving the hand 904 while maintaining the pinch gesture.

[0171] In some embodiments, in response to detecting that the user has stopped performing the pinch gesture with the hand 904, the electronic device 101 updates the current input state of the slider user interface element 908 to a value corresponding to the position of the indicator 910 of the current input state of the slider user interface element 908 when the pinch gesture was stopped. For example, in response to detecting that the user has stopped performing the pinch gesture while the slider user interface element 908 is being displayed as shown in FIG. 9E, the electronic device 101 updates the current input state of the slider user interface element 908 to the position of the indicator 910 shown in FIG. 9E and maintains the display of the indicator 910 as shown in FIG. 9E.

[0172] Figures 10A - 10J are flowcharts showing a method for enhancing interaction with slider user interface elements, according to some embodiments. In some embodiments, method 1000 is executed in a computer system (e.g., computer system 101 of FIG. 1 such as a tablet, smartphone, wearable computer, or head - mounted device) that includes a display generation component (e.g., display generation component 120 of FIGS. 1, 3, and 4) (e.g., a head - up display, a display, a touch screen, a projector, etc.) and one or more cameras (e.g., a camera facing downward with the user's hand (e.g., a color sensor, an infrared sensor, and other depth - sensing cameras), or a camera facing forward from the user's head). In some embodiments, method 1000 is stored on a non - transitory computer - readable storage medium and is executed by instructions executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A). Some of the operations of method 1000 may be arbitrarily combined and / or the order of some of the operations may be arbitrarily changed.

[0173] In some embodiments, such as FIG. 9A, method 1000 is executed on an electronic device that communicates with a display generation component and an eye tracking device (e.g., a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer). In some embodiments, the display generation component is a display integrated with the electronic device (optionally a touch screen display), an external display such as a monitor, projector, television, or a hardware component (optionally integrated or external) for projecting a user interface or making the user interface visible to one or more users. In some embodiments, the eye tracking device is a camera and / or a motion sensor capable of determining the direction and / or location of the user's line of sight. In some embodiments, the electronic device communicates with a hand tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (touch screen, trackpad)). In some embodiments, the hand tracking device is a wearable device such as a smart glove. In some embodiments, the hand tracking device is a handheld input device such as a remote control or a stylus.

[0174] In some embodiments, such as FIG. 9A, an electronic device (e.g., 101) displays (1002a) a slider user interface element (e.g., 908) via a display generation component 120. In some embodiments, the slider user interface element includes a current representation of an input point corresponding to the current input state of the slider user interface element and an individual representation of an input point corresponding to an individual input state of the slider that is different from the current input state of the slider (or a plurality of respective representations of input points each corresponding to an individual input state). In some embodiments, the electronic device first displays a slider user interface element having a slider bar at a position corresponding to the current input state without displaying the individual representation of the input point. In some embodiments, the electronic device displays the individual representation of the input point according to a determination that one or more first criteria are met. The one or more first criteria optionally include a criterion that is met in response to detecting, via an eye tracking device, that the user is looking at the slider user interface element, or in some embodiments, in response to detecting that the user has looked at the slider user interface element for longer than a predetermined period (e.g., 0.1, 0.2, 0.3, 0.4 seconds, etc.). In some embodiments, while displaying the slider user interface element along with the current representation of the input point and the individual representation of the input point, the electronic device displays the slider user interface element along with a plurality of respective representations of the input point at various locations along the length of the slider user interface element (e.g., at various predetermined input positions along the slider, such as the 10%, 20%, 30%, etc. positions along the slider). In some embodiments, the individual representation of the input point is a marking that is overlaid and displayed on the slider user interface element. For example, the user interface element of the slider includes a slider bar extending from one end of the slider to an individual location corresponding to the current input state of the slider user interface element, and the individual representation of the input point is overlaid on the slider bar or on a portion of the slider user interface element other than the slider bar and is displayed.In some embodiments, the user's line of sight does not coincide with the individual representation of the input point, but the individual representation of the input point is displayed with a first size, a first color, and / or a first transparency.

[0175] In some embodiments, while displaying a slider user interface element (e.g., 908), an electronic device (e.g., 101) detects (1002b) that the user's line of sight (e.g., 918b) is directed towards the slider user interface element (e.g., 908) via an eye tracking device, as shown in FIG. 9B. In some embodiments, the user's line of sight is detected by the eye tracking device as being directed towards the slider user interface element. In some embodiments, the electronic device detects that the user's line of sight is directed towards the slider user interface element for a period between a first time threshold (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5 seconds, etc.) and a second time threshold greater than the first time threshold (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5 seconds, etc.) via the eye tracking device.

[0176] In some embodiments, such as those of FIG. 9B, in response to detecting that the user's line of sight (e.g., 918b) is directed towards a slider user interface element (e.g., 908), the electronic device (e.g., 101) displays (1002c), via a display generation component, a representation of an input point (e.g., 914e) having a first appearance at a location on the slider user interface element (e.g., 908) determined based on the direction of the user's line of sight (e.g., initially display a representation of an input point having a first appearance, or change the appearance of the representation of the input point from a previous appearance to a first appearance different from the first appearance). In some embodiments, the representation of the input point is one of a plurality of respective locations along the slider corresponding to individual input states of the slider. In some embodiments, the electronic device displays a visual indication of each location along the slider corresponding to an individual input state of the slider. In some embodiments, the visual indication is displayed in response to detecting the user's line of sight on the slider user interface element. In some embodiments, the representation of the input point is an indication of the current input state of the slider. In some embodiments, the electronic device updates (e.g., the size, color, opacity, etc. of an additional visual indication, or add an additional visual indication) an indication of the current input state of the slider in response to detecting the user's line of sight on the slider user interface element and / or an indication of the current input state of the slider. In some embodiments, updating the individual representation of the input point includes one or more of updating the size, color, translucency, and / or opacity, or updating the virtual layer of the user interface on which the individual representation of the input point is displayed (e.g., popping out the individual representation of the input point in front of the rest of the slider element and / or other individual representations of the input point).

[0177] In some embodiments, as shown in FIG. 9B, after displaying a representation of an input point (e.g., 914e) having a first appearance (1002d), the electronic device (e.g., 101) determines that the user's line of sight (e.g., 918c) meets one or more first criteria, including a criterion that is satisfied when the user's line of sight (e.g., 918c) is directed at the representation of the input point (e.g., 914e) for a period longer than a time threshold (e.g., a second time threshold such as 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 0.6, 1 second, etc.) as shown in FIG. 9C. According to this determination, the electronic device sets the current input state of the slider user interface element (e.g., 908) to an individual input state corresponding to the representation of the input point (e.g., 914e) as shown in FIG. 9D (1002e). In some embodiments, the electronic device sets the current input state of the slider user interface element to an individual input state corresponding to the individual representation of the input point without additional input in response to a request to update the slider (e.g., without input detected via a hand tracking device), according to the determination that the user's line of sight meets one or more first criteria. In some embodiments, setting the current input state of the slider user interface element to an individual input state corresponding to the individual representation of the input point includes updating the slider user interface element to display the slider bar at a location corresponding to the individual input state (rather than the previous input state). In some embodiments, according to the determination that the user's line of sight does not meet one or more criteria, the electronic device cancels the update of the current input state of the slider.

[0178] In response to detecting that the user's gaze is directed towards an individual representation of an input point, update the individual representation of the input point to have a second appearance, and set the current input state of the slider user interface element to an individual input state according to a determination that the user's gaze meets one or more first criteria. The method described above provides an efficient way to provide the user with feedback that the user's gaze changes and / or actually changes the input state of the slider, thereby simplifying the interaction between the user and the electronic device, improving the operability of the electronic device, making the user-device interface more efficient, and enabling the user to use the electronic device more quickly and efficiently. By doing so, the power consumption is further reduced, the battery life of the electronic device is improved, and errors during use are reduced.

[0179] In some embodiments, such as FIG. 9B, in response to detecting that the user's line of sight (e.g., 918) is directed at a slider user interface element (e.g., 908), the electronic device (e.g., 101) displays, via a display generation component, a plurality of representations of input points (e.g., 914e) including representations of input points (e.g., 914a - g) at different locations on the slider user interface element (1004a). In some embodiments, each of the representations of the input points corresponds to an individual input state of the slider. In some embodiments, the slider includes additional input states between visual indicators. In some embodiments, the slider does not include additional input states between visual indicators (e.g., all possible input states are marked by indicators). In some embodiments, as in FIG. 9B, while displaying a slider user interface element that includes a plurality of representations of input points (e.g., 914a - 914g) (e.g., while the user's line of sight is directed at the slider user interface element), the electronic device (e.g., 101) detects user input including individual gestures performed by the user's hand (e.g., 704) via a hand - tracking device that communicates with the electronic device, as in FIG. 9B (1004b). In some embodiments, after a gesture, the user's hand movement continues while maintaining the individual gesture, and the magnitude of the hand movement corresponds to an individual location on the slider user interface element that does not correspond to one of the plurality of representations of the input points. In some embodiments, the individual gesture performed by the hand is for the user to touch the thumb with another finger of the same hand (e.g., index finger, middle finger, ring finger, little finger). In some embodiments, the electronic device detects a user holding the thumb on a finger while moving the hand and / or arm in the direction the indicator is pointing. For example, the electronic device detects a horizontal movement of the hand and changes the input state of a horizontal slider. In some embodiments, the electronic device updates the current input state of the slider according to the magnitude and / or speed and / or duration of the movement.For example, in response to detecting that the user has moved their hand by a first amount, the electronic device moves an indication of the current input state of the slider by a second amount, and if the user moves their hand by a third amount, the electronic device moves an indication of the current input state of the slider by a fourth amount. In some embodiments, such as those of FIG. 9D, in response to detecting user input, the electronic device (e.g., 101) sets the current input state of the slider user interface element (e.g., 908) to a second distinct input state corresponding to one of a plurality of representations of the input point (1004c). In some embodiments, the representation of the input point corresponding to the second distinct input state is the representation of the input point closest to the distinct location corresponding to the magnitude of the hand movement. In some embodiments, the slider includes a visual indicator disposed at a location along the slider corresponding to the current input state of the slider. In some embodiments, in response to detecting that the location on the slider corresponding to the user's hand movement does not correspond to one of the representations of the input point, the electronic device moves the current input state of the slider to the representation of the input point closest to the location corresponding to the hand movement.

[0180] The method described above of setting the current input state of the slider to an input state corresponding to a representation of the input point in response to user movement corresponding to a location not including a representation of the input point provides an efficient way of selecting the input state corresponding to the representation of the input point, thereby further reducing power usage, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0181] In some embodiments, in response to detecting that the user's line of sight (e.g., 918) is directed at a slider user interface element (e.g., 908), the electronic device (e.g., 101) displays (1005a), via a display generation component, a plurality of representations of input points (e.g., 914e), including representations of input points (e.g., 914a - g), at different locations on the slider user interface element (e.g., 908), as shown in FIG. 9B. In some embodiments, each representation of an input point corresponds to an individual input state of the slider. In some embodiments, the slider includes additional input states between visual indicators. In some embodiments, the slider does not include additional input states between visual indicators (e.g., all possible input states are marked by indicators). In some embodiments, as shown in FIG. 9B, while displaying a slider user interface element (e.g., 908) that includes a plurality of representations of input points (e.g., 914a - g) (e.g., while the user's line of sight is directed at the slider user interface element), the electronic device (e.g., 101) detects (1005b) user input, via a hand - tracking device that communicates with the electronic device, including an individual gesture performed by the user's hand (e.g., 904), as shown in FIG. 9B, and subsequent movement of the user's hand (e.g., 904) while maintaining the individual gesture, where the magnitude of the hand movement corresponds to an individual location on the slider user interface element (e.g., 908) that does not correspond to one of the plurality of representations of input points (e.g., 914a - g). In some embodiments, the individual gesture performed by the hand is the user touching the thumb to another finger of the same hand (e.g., index finger, middle finger, ring finger, little finger). In some embodiments, the electronic device detects a user who holds the thumb against a finger while moving the hand and / or arm in the direction the indicator is pointing. For example, the electronic device detects a horizontal movement of the hand and changes the input state of a horizontal slider. In some embodiments, the electronic device updates the current input state of the slider according to the magnitude and / or speed and / or duration of the movement.For example, in response to detecting that the user has moved their hand by a first amount, the electronic device moves an indication of the current input state of the slider by a second amount, and if the user moves their hand by a third amount, the electronic device moves an indication of the current input state of the slider by a fourth amount. In some embodiments, while detecting hand gestures and movements, the electronic device stops displaying multiple representations of the input point. In some embodiments such as FIG. 9E, in response to detecting user input, the electronic device (e.g., 101) sets the current input state of the slider user interface element (e.g., 908) to a second individual input state corresponding to an individual location on the slider user interface element (e.g., 908) (1005c). In some embodiments, the second individual input state does not correspond to one of the multiple representations of the input point. Thus, in some embodiments, the user can set the current input state of the slider to any state within the slider when using hand gestures as described above. In some embodiments, the second individual input state of the slider user interface element is based on the location of the user's line of sight when a predetermined hand gesture is detected and the direction, distance, and speed of the user's hand movement while maintaining the predetermined gesture.

[0182] The method described above of setting the current input state to a location corresponding to a user hand movement that does not correspond to one of the multiple representations of the input point efficiently provides the user with the ability to fine-tune the input state of the slider to an input state between the multiple representations of the input point or otherwise to an input state that does not correspond to the multiple representations of the input point, thereby further reducing power usage, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0183] In some embodiments, such as FIG. 9B, in response to detecting that the user's line of sight (e.g., 918b) is directed towards a slider user interface element, the electronic device displays, via a display generation component, a control element (e.g., 916) (e.g., a cursor) on the slider user interface element (e.g., 908) that indicates a location on the slider user interface element (e.g., 908) corresponding to the current location of a predetermined portion of the user's hand (e.g., 904) (e.g., one or more of the user's fingers and / or the user's thumb) (1006a). In some embodiments, the electronic device first positions the cursor at a location corresponding to the location of the user's hand. For example, if the user's hand is to the left of a predetermined region of the three-dimensional environment where the slider is displayed, the electronic device displays the cursor to the left of the current input state of the slider. As another example, if the user's hand is to the right of the predetermined region, the electronic device displays the cursor to the right of the current input state of the slider. In some embodiments, the electronic device displays the cursor in response to detecting a user performing a predetermined gesture (e.g., index finger, middle finger, ring finger, little finger) with the same hand as the thumb via a hand tracking device. In some embodiments, such as FIG. 9E, while detecting the movement of the user's hand (e.g., 904) while maintaining an individual gesture, the electronic device (e.g., 101) moves a control element (e.g., 916) on the slider user interface element (e.g., 908) in accordance with the movement of the user's hand (e.g., 904) (1006b). In some embodiments, the electronic device moves the cursor in accordance with the movement of the hand in the dimension in which the slider is oriented. For example, in response to detecting movement of the hand upwards and to the right, the electronic device updates the current input state of the horizontal slider by moving the horizontal slider to the right or updates the current input state of the vertical slider by moving the vertical slider upwards. In some embodiments, the electronic device moves the cursor without updating the current input state in response to detecting movement of the hand without detecting that the hand is performing a predetermined gesture (e.g., touching the thumb to another finger (e.g., index finger, middle finger, ring finger, little finger) of the same hand as the thumb).

[0184] The above-described method of displaying and updating the control element of the slider provides an efficient way to show the user how the input state of the slider is updated in response to hand detection-based input, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0185] In some embodiments, as shown in FIG. 9C, after displaying the representation of the input point (e.g., 914e) in the first appearance, while the user's line of sight (e.g., 918c) is directed towards the representation of the input point (e.g., 914e), but before the user's line of sight (e.g., 918c) is directed towards the representation of the input point (e.g., 914e) for a longer time than the time threshold (such as 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 0.6, 1 second, etc.), as shown in FIG. 9B, the electronic device (e.g., 101) detects (1008a), via a hand-tracking device that communicates with the electronic device, an individual gesture (e.g., touching the thumb to another finger (e.g., index finger, middle finger, ring finger, little finger) of the same hand and extending one or more fingers towards a slider user interface element) performed by the user's hand (e.g., 904). In some embodiments such as FIG. 9D, in response to the detection of the individual gesture, the electronic device (e.g., 101) sets the current input state of the slider user interface element (e.g., 908) to an individual input state corresponding to the representation of the input point (e.g., 914e) (e.g., before the user's line of sight is directed towards the representation of the input point for a threshold amount of time (such as 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 0.6, 1 second, etc.)). In some embodiments, the electronic device sets the current input state of the slider user interface element to an individual input state corresponding to the individual representation of the input point in response to detecting a gesture while the user's line of sight is directed towards the individual representation of the input point. In some embodiments, the electronic device updates the input state of the slider in response to detecting, via the hand-tracking device, that the user's hand is in a predetermined location and / or has performed a predetermined gesture. For example, the predetermined location corresponds to a virtual location where the slider user interface element and / or the individual representation of the input point are displayed in the user interface. As another example, the predetermined gesture is that the user taps the thumb and a finger (e.g., index finger, middle finger, ring finger, little finger) together.In some embodiments, in accordance with a determination that one or more first criteria are not met and / or the electronic device does not detect a gesture, the electronic device halts updating the current input state of the slider user interface element.

[0186] The above-described method of updating the input state of the slider before reaching the threshold time in response to a gesture provides an efficient way to interact with the slider in a time shorter than the threshold time, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0187] In some embodiments, such as those of FIG. 9A, a slider user interface element (e.g., 908) includes a current state indicator (e.g., 910) that indicates the current input state of the slider user interface element (e.g., 908) (1010a). In some embodiments, the slider includes a bar with one end aligned with the end of the slider and the other end (e.g., the current state indicator) aligned with the current input state of the slider. In some embodiments, the indicator is a visual indication displayed at a location corresponding to the current input state of the slider. In some embodiments, as shown in FIG. 9B, after displaying a representation of an input point having a first appearance (e.g., 914e) (1010b), and as shown in FIG. 9C, in accordance with a determination that the user's line of sight (e.g., 918c) meets one or more first criteria (e.g., the line of sight is held over the representation of the input point for a threshold time (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 0.6, 1 second, etc.)), the electronic device (e.g., 101) moves the current state indicator (e.g., 910) to a location on the slider user interface element (e.g., 908) corresponding to the representation of the input point (e.g., 914e), as shown in FIG. 9D (1010c). In some embodiments, as shown in FIG. 9B, after displaying a representation of an input point having a first appearance (e.g., 914e) (1010b), and as shown in FIG. 9C, in accordance with a determination that the user's line of sight (e.g., 918c) does not meet one or more first criteria and an individual gesture of the user's hand (e.g., 904) is detected (e.g., via a hand tracking device) while the user's line of sight (e.g., 918c) is directed at the representation of the input point (e.g., 914e), the electronic device (e.g., 101) moves the current state indicator (e.g., 910) to a location on the slider user interface element (e.g., 908) corresponding to the representation of the input point (e.g., 914e), as shown in FIG. 9D (1010d).In some embodiments, the electronic device moves the current state indicator to a location on a slider element corresponding to the representation of the input point in response to detecting a hand gesture while the line of sight has been held for less than a threshold time (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 0.6, 1 second, etc.). In some embodiments, the gesture includes the user touching the thumb with another finger of the same hand (e.g., index finger, middle finger, ring finger, little finger). In some embodiments, as shown in FIG. 9B, in accordance with the determination that an individual gesture of the user's hand (e.g., 904) is detected, followed by movement of the user's hand while maintaining the individual gesture (e.g., via a hand tracking device), after displaying the representation of the input point (e.g., 914e) in a first appearance (1010b), the magnitude of the movement of the user's hand (e.g., distance, speed, duration) corresponds to a location on a slider user interface element (e.g., 908) corresponding to the representation of the input point, and the electronic device (e.g., 101) moves the current state indicator (e.g., 910) to a location on the slider user interface element (e.g., 908) corresponding to the representation of the input point in accordance with the movement of the user's hand (e.g., 904), as shown in FIG. 9D (1010e). In some embodiments, in accordance with the determination that the magnitude of the movement of the user's hand corresponds to a second location on a slider corresponding to a second representation of the input point, the electronic device moves the current state indicator to a second location on a slider element corresponding to the second representation of the input point in accordance with the movement of the hand. Thus, in some embodiments, the electronic device updates the current input state of the slider in response to detecting any of (1) the user's line of sight being on the representation of the input point for a threshold time, (2) the user's line of sight being on the representation of the input point while a hand gesture is detected, or (3) the user making a gesture with the hand and moving the hand in the direction the slider is facing (e.g., while holding the gesture).

[0188] The above-described method of updating a slider according to a line-of-sight input and a non-line-of-sight input provides different fast and efficient ways to update the input state of the slider, enables the user to provide convenient and accessible input, and enables the user to use the electronic device more quickly and efficiently, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use.

[0189] In some embodiments, as shown in FIG. 9B, while displaying the representation of the input point (e.g., 914e) in the first appearance, the electronic device (e.g., 101) detects (1012a) via the eye tracking device that the user's line of sight is directed towards the second representation (e.g., 914b) of the second input point at a second location on the slider user interface element (e.g., 908). In some embodiments, while the user's line of sight is directed towards the representation of the input point, the second representation of the second input point is displayed in an appearance different from the first appearance (e.g., size, color, opacity, translucency, virtual layer, distance from the user's viewpoint). In some embodiments, the slider includes a plurality of representations of the input point, and one or more or all of the representations of the input point other than the representation of the input point that the user is currently viewing are displayed in an appearance other than the first appearance, and the representation of the input point is displayed in the first appearance. For example, while the user's line of sight is directed towards the representation of the input point, the other representations of the input point are displayed in a size smaller than the representation of the input point that the user is viewing. In some embodiments, in response to detecting that the user's line of sight is directed towards the second representation (e.g., 914b) of the input point, the electronic device (e.g., 101) displays (1012b) the second representation (e.g., 914b) of the second input point having the first appearance at the second location of the slider user interface element (e.g., 908) (e.g., updates the representation of the input point displayed in an appearance other than the first appearance). For example, in response to detecting the user viewing the second representation of the input point, the electronic device displays the representation of the input point (e.g., and one or more or all of the other representations of the input point on the slider) in a size smaller than the second representation of the input point. Thus, in some embodiments, when the user first views the first representation of the input point and then moves their line of sight to the second representation of the input point, the electronic device updates the appearance of the second representation of the input point in response to the user's line of sight towards the second representation of the input point (e.g., to change the appearance of the first representation of the input point).

[0190] In response to the user's line of sight being directed towards the second representation of the second input point, the method described above for updating the appearance of the second representation of the second input point provides an efficient way to change the representation of the selected input point, thereby enabling the user to use the electronic device more quickly and efficiently, further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use.

[0191] In some embodiments, such as FIG. 9B, in response to detecting that the user's line of sight (e.g., 918b) is directed towards a slider user interface element (e.g., 908), the electronic device (e.g., 101) displays, via a display generation component, a control element (e.g., 916) (e.g., a cursor, a representation of the user's hand) indicating a location on the slider user interface element (e.g., 908) corresponding to the current location of a predetermined portion (e.g., one or more fingers) of the user's hand (e.g., 904) on the slider user interface element (e.g., 908) (1014a). For example, in accordance with the determination that the user's hand is located to the left of a predetermined region in space, the electronic device displays the control element to the left of the slider. In some embodiments, such as FIG. 9E, while detecting movement of a predetermined portion of the user's hand (e.g., 904), the electronic device (e.g., 101) moves the control element (e.g., 916) on the slider user interface element (e.g., 908) in accordance with the movement of the predetermined portion of the user's hand (e.g., 904) (1014b). For example, in response to detecting a hand moving to the right, the electronic device moves the control element to the right. In some embodiments, the control element moves at a speed proportional to the speed of the hand movement and / or by a distance proportional to the distance the hand has moved. Thus, in some embodiments, the electronic device provides visual feedback to the user while the user is controlling the input state of the slider with a hand input.

[0192] While the user is controlling the slider with the movement of their hand, the above-described method of displaying and moving the control element on the slider user interface provides an efficient way to show the user how the slider changes in response to the hand input while the hand input is being provided, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0193] In some embodiments, such as FIG. 9A, prior to detecting that the user's line of sight is directed towards a slider user interface element (e.g., 908), the representation of the input point (e.g., 914e) of FIG. 9B is not displayed (1016a) on the slider user interface element (e.g., 908). In some embodiments, prior to detecting the user's line of sight on the slider user interface element, the electronic device displays the slider user interface element with the current input state display of the slider without displaying an indication of other individual input states of the slider (e.g., without displaying the representation of the input point on the slider).

[0194] The above-described method of displaying the representation of the input point of the slider in response to detecting the user's line of sight on the slider provides an efficient way to show the user that the input state of the slider is variable while the user is looking at the slider, reduces the user's visual clutter and cognitive burden while the user's line of sight is not directed towards the slider, and further reduces power consumption, improves the battery life of the electronic device, and reduces errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0195] In some embodiments, as shown in FIG. 9B, while displaying a slider user interface element (e.g., 908) that includes representations of input points (e.g., 914e) (e.g., multiple individual representations of input points corresponding to multiple individual input states of a slider), an electronic device (e.g., 101) detects (1018a) user input that includes an individual gesture performed by a user's hand (e.g., 904) via a hand-tracking device that communicates with the electronic device (e.g., 101), as shown in FIG. 9E, and subsequent movement of the user's hand (e.g., 904) while maintaining the individual gesture. In some embodiments, the individual gesture is the user touching their thumb to another finger of the same hand (e.g., index finger, middle finger, ring finger, little finger), and the electronic device detects movement of the hand and / or the arm of the hand while the thumb is touching the finger. In some embodiments, the electronic device updates the location of an indication of the current input state of the slider according to the movement of the hand while the gesture is being held, and determines the input state of the slider in response to detecting that the user has released the gesture (e.g., moves the thumb and finger away from each other). In some embodiments such as FIG. 9E, while detecting user input, the electronic device (e.g., 101) stops (1018b) the display of the representation of the input point in FIG. 9B on the slider user interface element (e.g., 908) (e.g., stops the display of multiple representations of input points on the slider). In some embodiments, after stopping the display of the representation of the input point, the electronic device continues to display an indication of the current input state of the slider. In some embodiments such as FIG. 9E, in response to detecting user input (e.g., in response to detecting that the user has stopped the gesture, moves the thumb and finger away from each other), the electronic device (e.g., 101) sets (1018c) the current input state of the slider user interface element (e.g., 908) to a second individual input state according to the movement of the user's hand (e.g., 904). In some embodiments, the current input state of the slider moves according to the distance, speed, and / or duration of the movement of the user's hand and / or arm.For example, the current input state moves by an amount greater than the amount by which the current input state moves in response to hand movements with a relatively long distance and / or duration and / or high speed, in response to hand movements with a relatively short distance and / or duration and / or low speed. In some embodiments, after the input for updating the current input state of the slider has ended, the electronic device displays a slider user interface element having a representation of the input point. Thus, in some embodiments, the electronic device stops displaying the representation of the input point of the slider (e.g., and multiple representations of the input point) while detecting an input for changing the current input state of the slider that includes the movement of the user's arm.

[0196] The above-described method of stopping the display of the representation of the input point while an input including hand movement is being detected provides an efficient way to reduce the user's visual confusion and cognitive burden while interacting with the slider, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0197] In some embodiments, an electronic device (e.g., 101) detects (1020a) the movement of a user's hand (e.g., 904) via a hand tracking device, as shown in FIG. 9E. In some embodiments, the electronic device (e.g., 101) updates (1020b) the current input state of a slider user interface element (e.g., 908) according to the movement of the user's hand (e.g., 904), as shown in FIG. 9E, according to a determination that one or more criteria are met while the movement of the user's hand (e.g., 904) is detected and the user's line of sight is directed at the slider user interface element (e.g., 908). In some embodiments, the one or more criteria further include detecting a predetermined gesture (e.g., a pinch gesture) being performed by the user's hand. In some embodiments, updating the current input state of the slider user interface element according to the movement of the user's hand (while maintaining, e.g., a pinch gesture) includes updating the current input state of the slider according to the component of the hand movement that is in the direction of the slider user interface element. For example, moving the hand upward and to the right causes a horizontal slider to move to the right or a vertical slider to move upward. In some embodiments, the electronic device (e.g., 101) stops (1020c) updating the current input state of the slider user interface element (e.g., 908) according to the movement of the user's hand (e.g., 904) according to a determination that the movement of the user's hand (e.g., 904) is detected while one or more criteria, including criteria that are met when the user's line of sight is directed at the slider user interface element (e.g., 908), are not met. In some embodiments, the criteria are not met in response to detecting the user's line of sight directed at a control element including the slider unless the line of sight is directed at the slider itself.

[0198] The method described above of ceasing to update the current input state of the slider when the line of sight is not directed towards the slider provides an efficient way to prevent the user from accidentally updating the input state of the slider, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0199] In some embodiments, the slider user interface element is included in a control region of the user interface (e.g., a region of the user interface that, when operated, causes the electronic device to change and / or activate settings and / or perform another action, and includes a plurality of user interface elements). In some embodiments, the control region of the user interface is visually distinguishable from the rest of the user interface (e.g., the control region is a visual container within the user interface). In some embodiments, while displaying the slider user interface element without an indication of the input point, the electronic device detects, via an eye tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touch screen, trackpad)), that the user's line of sight is directed towards the control region of the user interface rather than towards the slider user interface element. In some embodiments, in response to detecting that the user's line of sight is directed towards the control region of the user interface rather than towards the slider user interface element, the electronic device maintains the display of the slider user interface element without an indication of the input point. In some embodiments, the user's line of sight is directed towards a portion of the control region that does not include a user interface element (e.g., a portion of the background of the control region). In some embodiments, the user's line of sight is directed towards another user interface element within the control region. In some embodiments, the electronic device does not display an indication of the input point (e.g., or any other indication of the input point of the slider) unless the user's line of sight is directed towards the slider (e.g., or within a threshold distance of the slider (e.g., 0.5, 1, 2, 3 centimeters, etc.)). In some embodiments, in response to detecting, via a hand tracking device, that the user performs a predetermined hand gesture (e.g., touching the thumb to another finger of the same hand (e.g., index finger, middle finger, ring finger, little finger)) and moves the hand along the direction of the slider user interface element without detecting the user's line of sight on the slider user interface element, the electronic device cancels the update of the current input state of the slider user interface element.

[0200] In some embodiments, such as FIG. 9A, a slider user interface element (e.g., 908) is included (1022a) in a control region of the user interface (e.g., 906) (e.g., a region of the user interface that includes a plurality of user interface elements that, when operated, cause the electronic device to change and / or activate settings and / or perform another action). In some embodiments, the control region of the user interface is visually distinct from the rest of the user interface (e.g., the control region is a visual container within the user interface). In some embodiments, as shown in FIG. 9A, while displaying the slider user interface element (e.g., 908) without the representation of the input point of FIG. 9B (e.g., 914e), the electronic device (e.g., 101) detects (1022b) via an eye tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touch screen, trackpad)) that the user's line of sight (e.g., 918a) is directed towards the control region of the user interface (e.g., 906) as shown in FIG. 9B (e.g., regardless of whether the user's line of sight is directed towards the slider user interface element). In some embodiments, the user's line of sight is directed towards a part of the control region that does not include user interface elements (e.g., a part of the background of the control region). In some embodiments, the user's line of sight is directed towards another user interface element within the control region. In some embodiments, in response to detecting that the user's line of sight (e.g., 918a) is directed towards the control region of the user interface (e.g., 906), the electronic device (e.g., 101) displays (1022c) the slider user interface element that includes the representation of the input point (e.g., 914e) as shown in FIG. 9B. In some embodiments, in accordance with the determination that the user's line of sight is not directed towards the control region (e.g., the user is looking at a different part of the user interface, has their eyes closed, or has their eyes away from the display generation component), the electronic device displays the slider user interface element without the representation of the input point.In some embodiments, via a hand tracking device, in response to detecting that a user performs a predetermined hand gesture (e.g., touching the thumb to another finger of the same hand (e.g., index finger, middle finger, ring finger, little finger)) and moves the hand along the direction of a slider user interface element while detecting the user's line of sight on a control area, the electronic device updates the current input state of the slider user interface element according to the movement of the user's hand. In some embodiments, in response to detecting the user's line of sight on the control area, the electronic device updates one or more other user interface elements within the control area (e.g., updating a second slider to include a representation of an input point, updating the color, size, virtual distance from the user of one or more other user interface elements).

[0201] The above-described method of displaying a representation of an input point in response to detecting the user's line of sight on a part of the control area other than the slider provides an efficient way to indicate the input state of the slider without the user having to wait to look at the slider, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0202] In some embodiments, in response to detecting that the user's line of sight is directed at a slider user interface element (e.g., 908), the electronic device (e.g., 101) displays (1024a) a visual indication of the line of sight at a location on the slider user interface element (e.g., 908) based on the direction of the user's line of sight and / or a portion of the user interface element at the location of the user's line of sight on the slider user interface element (e.g., 908). In some embodiments, in response to detecting movement of the user's line of sight from a first discrete portion of the slider to a second discrete portion of the slider, the electronic device updates the location of the visual indication to correspond to the second discrete portion of the slider. In some embodiments, the visual indication is one of a change in the icon and / or appearance (e.g., color, translucency, opacity, etc.) of the slider at the location the user is looking at.

[0203] The above-described method of displaying a visual indication of the user's line of sight provides an efficient way of indicating that the user's line of sight is being detected and / or that the electronic device can update the slider according to sight-based input, thereby further reducing power usage, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0204] In some embodiments, as shown in FIG. 9B, after displaying a representation of an input point having a first appearance (e.g., 914e), in response to detecting a user's line of sight (e.g., 918c) directed at the representation of the input point having the first appearance for less than a time threshold (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 0.6, 1 second, etc.), the electronic device (e.g., 101) displays (1024b), as shown in FIG. 9C, a representation of an input point having a second appearance different from the first appearance (e.g., size, color, opacity, translucency, distance from the user's perspective in the three-dimensional environment where the slider is displayed). In some embodiments, the electronic device gradually updates the appearance of the representation of the input point as long as the user's line of sight remains on the representation of the input point until the threshold time is reached. For example, before detecting the user's line of sight on the representation of the input point, the electronic device displays the representation of the input point in a first size, and in response to detecting the user's line of sight on the representation of the input point, while continuously detecting the user's line of sight on the representation of the input point, the electronic device gradually increases the size of the representation of the input point until the threshold time is reached, and the electronic device updates the input state of the slider to correspond to the individual representation of the input point. As another example, the electronic device gradually changes the color of the representation of the input point while the user's line of sight is on the representation of the input point.

[0205] The method described above of updating the appearance of the representation of the input point while the user's line of sight is on the representation of the input point for less than the threshold time provides an efficient way to show the user that the input state of the slider is updated to correspond to the representation of the input point when the user continues to look at the representation of the input point, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0206] FIGS. 11A - 11D show examples of how an electronic device moves virtual objects and facilitates access to actions associated with the virtual objects in a three-dimensional environment according to some embodiments.

[0207] FIG. 11A shows an electronic device 100 that displays a three-dimensional environment 1102 on a user interface via a display generation component 120. As described above with reference to FIGS. 1-6, the electronic device 101 optionally includes a display generation component 120 (e.g., a touch screen) and a plurality of image sensors 314. The image sensors 314 optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor, and the electronic device 101 can be used to capture one or more images of the user or a part of the user while the user is interacting with the electronic device 101. In some embodiments, the display generation component 120 is a touch screen that can detect gestures and movements of the user's hand. In some embodiments, the user interface shown below can also be implemented on a head-mounted display that includes a display generation component that displays the user interface to the user, a sensor that detects the physical environment and / or the movement of the user's hand (e.g., an external sensor facing outward from the user), and / or a sensor that detects the user's line of sight (e.g., an internal sensor facing inward toward the user's face).

[0208] In FIG. 11A, the electronic device 101 displays an application within a three-dimensional environment 1102, a two-dimensional object 1106b, and a representation 1106a of a three-dimensional object 1106c. In some embodiments, the representation 1106a of the application includes a user interface of the application that includes selectable options, content, and the like. In some embodiments, the two-dimensional object 1106b is a file or item of content such as a document, an image, or video content. In some embodiments, the two-dimensional object 1106b is an object associated with the application associated with the representation 1106a or a different application. For example, the representation 1106a is a representation of an email application, and the two-dimensional object 1106b is an attachment to an email that the electronic device 101 displays outside of the representation 1106a in response to an input to separately display the two-dimensional object 1106b. In some embodiments, the three-dimensional object 1106c is a virtual object or three-dimensional content. In some embodiments, the three-dimensional object 1106c is associated with the same application or a different application associated with the representation 1106a.

[0209] In some embodiments, in response to detecting the user's line of sight on an individual virtual object (e.g., representation 1106a, 2D object 1106b, 3D object 1106c), in some embodiments, for a predetermined threshold time (e.g., 0.1, 0.2, 0.5, 1 second, etc.), the electronic device 101 displays user interface elements 1180a - 1180c proximate to the individual virtual object. In some embodiments, the electronic device 101 displays the user interface elements 1180a - c in response to detecting a user performing a gesture with the hand 1104b while the user's line of sight is on an individual virtual object (e.g., representation 1106a, 2D object 1106b, 3D object 1106c), regardless of the duration of the line of sight (e.g., even if the line of sight is detected for less than the threshold time). In some embodiments, the gesture is a pinch gesture that includes the user touching the thumb with another finger (e.g., index finger, middle finger, ring finger, little finger) of the same hand as the thumb. For example, in response to detecting the user's line of sight 1110c on the 3D object 1106c over the threshold time and / or simultaneously with the detection of the pinch gesture, the electronic device 101 displays the user interface element 1180c. As another example, in response to detecting the user's line of sight 1110b on the 2D object 1106b during the threshold time and / or simultaneously with the detection of the pinch gesture, the electronic device 101 displays the user interface element 1180b. As another example, in response to detecting the user's line of sight on the application representation 1106a during the threshold time and / or simultaneously with the detection of the pinch gesture, the electronic device 101 displays the user interface element 1180a. The user interface element 1180a associated with the application representation 1106a is optionally larger than the user interface elements 1180b - c associated with the objects within the 3D environment 1102. In some embodiments, user interface elements similar to the user interface element 1180a are displayed at a first size when displayed in relation to the representation of an application and / or virtual object that is displayed independently from other objects within the 3D environment.In some embodiments, user interface elements similar to user interface element 1180b are displayed in a second, smaller size when displayed in relation to a representation of an object initially displayed within another object or application user interface. For example, representation 1106a is a representation of an email application displayed independently, and object 1106b is an attachment of an email initially displayed within representation 1106a. In some embodiments, each user interface element 1180 associated with a virtual object within the three-dimensional environment 1102 is the same for virtual objects associated with different applications. For example, two-dimensional object 1106b and three-dimensional object 1106c are associated with different applications, but user interface element 1180b and user interface element 1180c are the same. In some embodiments, the electronic device 101 displays one user interface element 1180 at a time at the time corresponding to a virtual object towards which the user's line of sight is directed. In some embodiments, the electronic device 101 displays all user interface elements 1180 simultaneously (e.g., in response to detection of the user's line of sight towards one of the virtual objects within the three-dimensional environment 1102).

[0210] While the line of sight 1110b on the two-dimensional object 1106b is detected, the electronic device 101 displays a user interface element 1180b on the right side of the two-dimensional object 1106b in response to detecting the user's hand 1104a on the right side of the two-dimensional object 1106b. In some embodiments, instead of detecting the hand 1104a on the right side of the two-dimensional object 1106b, if the electronic device 101 detects the hand 1104a on the left side of the two-dimensional object 1106b, the electronic device 101 displays the user interface element 1180b on the left side of the two-dimensional object 1106b. In some embodiments, the electronic device 101 similarly displays a user interface element 1180c on the side of the three-dimensional object 1106c where the user's hand 1104a is detected while the user's line of sight 1110c is directed at the three-dimensional object 1106c, and similarly displays a user interface element 1180a on the side of the representation 1106a where the user's hand 1104a is detected while the user's line of sight is directed at the representation 1106a. In some embodiments, if the location of the hand 1104a changes while the user's line of sight 1110b is directed at the two-dimensional object 1106b, the electronic device 101 presents an animation of the user interface element 1180b that moves to a location corresponding to the updated position of the hand.

[0211] In some embodiments, the electronic device 101 is in a predetermined hand posture, such as the start of a pinch gesture where the thumb is less than a threshold distance (e.g., 0.1, 0.5, 1, 2, 5, 10, 30 centimeters, etc.) from another finger, and when the hand 1104a is within a threshold distance (e.g., 1, 5, 10, 30, 50, 100 centimeters, etc.) of the two-dimensional object 1106b in the three-dimensional environment 1102, it only displays the user interface element 1180b to the right of the two-dimensional object 1106b according to the position of the hand 1104a. In some embodiments, in response to detecting a pinch gesture while the user's hand 1104a is within a threshold distance (e.g., 1, 5, 10, 30, 50, 100 centimeters, etc.) of the object 1106b in the three-dimensional environment 1102, the electronic device 101 moves the user interface element 1180b to the location corresponding to the pinch. For example, if the electronic device 101 detects the hand 1104a in a pinch gesture towards the lower right corner of the two-dimensional object 1106b, the electronic device 101 displays the user interface element 1180b at the lower right corner of the two-dimensional object 1106b (or, in some embodiments, at the center on the right side of the two-dimensional object 1106b). Instead of detecting the hand 1104a at the start of the pinch gesture, if the electronic device 101 detects the hand with a different gesture, such as a pointing gesture where one or more (but not all) fingers are extended while the hand 1104a is within a threshold distance (e.g., 1, 5, 10, 20, 50, 100 centimeters, etc.) of the two-dimensional object 1106b in the three-dimensional environment 1102, the electronic device 101 updates the user interface element 1180b to include one or more selectable options related to the two-dimensional object 1106b, similar to how the user interface element 1180a includes selectable options 1112a - e related to the representation 1106a of the application, as described with reference to FIG. 11B.

[0212] While the electronic device 101 is detecting the user's line of sight 1110b on the two-dimensional object 1106b, if the electronic device 101 does not detect the user's hand 1104a on one side of the two-dimensional object 1106b, the electronic device 101 optionally displays a user interface element 1180b along the bottom edge of the two-dimensional object 1106b in a manner similar to the way the electronic device 101 displays a user interface element 1180a along the bottom edge of the application representation 1106a in FIG. 11A. Since the three-dimensional object 1106c is three-dimensional, the electronic device 101 optionally displays a user interface element 1180c along the front bottom edge of the three-dimensional object 1106c (e.g., as opposed to the back (e.g., further) bottom edge of the three-dimensional object 1106c).

[0213] In some embodiments, in response to detecting an input directed to any of the user interface elements 1180, the electronic device 101 updates the user interface element 1180 to include one or more selectable options that, when selected, cause the electronic device 101 to perform an action directed to a virtual object associated with the user interface element 1180. In some embodiments, the input directed to the user interface element 1180 includes a gaze input. As shown in FIG. 11A, the electronic device 101 detects a user's gaze 1110a directed to the user interface element 1180a associated with the representation 1106a of the application 1106a. In some embodiments, the user's gaze 1106a on the user interface element 1180a is detected over a threshold duration (e.g., 0.1, 0.2, 0.5, 1, 5, 10, 30, 50 seconds, etc.). In some embodiments, the electronic device 101 detects a user performing a pinch gesture with the hand 1104b while detecting the user's gaze 1106a on the user interface element 1180a for a period less than the threshold time. In response to detecting the user's gaze 1110a on the user interface element 1180a during the threshold time or during any period while simultaneously detecting a pinch gesture performed by the hand 1104b, the electronic device 101 updates the user interface element 1180a to include a plurality of selectable options 1112a - e associated with the representation 1106a of the application, as shown in FIG. 11B.

[0214] FIG. 11B shows an electronic device 101 that displays selectable options 1112a - e for an extended user interface element 1180a in response to detection of one of the inputs directed to the user interface element 1180a described with reference to FIG. 11A. In some embodiments, in response to detection of selection of option 1112a, the electronic device 101 stops the display of representation 1106a. In some embodiments, in response to detection of selection of option 1112b, the electronic device 101 initiates a process for sharing the representation 1106a with another user within the three - dimensional environment 1102. In some embodiments, in response to detection of selection of option 1112c, the electronic device 101 initiates a process for updating the location of the representation 1106a within the three - dimensional environment 1102. In some embodiments, in response to detection of selection of option 1112d, the electronic device 101 displays the representation 1106a in a full - screen / immersive mode that includes, for example, stopping the display of other objects 1106b - c (e.g., virtual objects and / or real objects) within the three - dimensional environment 1102. In some embodiments, in response to detection of selection of option 1112e, the electronic device 101 displays all objects and representations associated with the same application associated with the representation 1106a that are within a threshold distance (e.g., 1, 5, 10, 30, 50 centimeters, etc.) of each other within the three - dimensional environment 1102. In some embodiments, the selection of options 1112a - e is detected in response to detecting the user's line of sight directed to an individual option while the user is performing a pinch gesture with hand 1104c.

[0215] In some embodiments, the options 1112a - e included in the user interface object 1180a are customized for the application corresponding to the representation 1106a. Thus, in some embodiments, the options displayed for different applications may be different from the options 1112a - e displayed for the application associated with the representation 1106a. For example, a content editing application may optionally include markup options, while an internet browser may optionally not include markup options. Further, in some embodiments, the electronic device 101 displays different options depending on whether the option is associated with an application's representation or a virtual object (e.g., 2D object 1106b, 3D object 1106c). For example, the options displayed in the representation 1180c associated with the 3D object 1106c or the options displayed in the representation 1180b associated with the 2D object 1106b may optionally be different from the options 1112a - e. For example, the options associated with the 2D object 1106b include an option to stop displaying the 2D object 1106b, an option to share the 2D object with other users who can access the 3D environment 1102, an option to move the object 1106b, an option to display the object 1106b in full - screen or immersive mode, and an option to edit the 2D object 1106b (e.g., via a markup or text - editing application).

[0216] As shown in FIG. 11B, the electronic device 101 detects the user's line of sight 1110d directed to option 1112c and moves the representation 1106a within the three-dimensional environment 1102. While detecting the user's line of sight 1110d at option 1112c, the electronic device 101 also detects that the user has performed a pinch gesture with hand 1104c. In some embodiments, while the user is maintaining the pinch gesture with hand 1104c, the electronic device 101 moves the application representation 1106a within the three-dimensional environment 101 in accordance with the movement of the hand. For example, in response to detecting the movement of hand 1104d towards the user, the electronic device 101 moves the representation 1106a towards the user within the three-dimensional environment 1102 as shown in FIG. 11D.

[0217] Referring to FIG. 11D, while the electronic device 101 moves the application representation 1106a according to the movement of the hand 1104f while maintaining the pinch gesture, the electronic device 101 displays the user interface element 1180a without selectable options 1112a - e (e.g., the device 101 folds the element 1180a to the state shown in FIG. 11A). In some embodiments, in response to detecting the end of the pinch gesture, the electronic device 101 maintains the display of the application representation 1106a at the updated location within the 3D environment 1102 and resumes the display of the selectable options 1112a - e (e.g., the device 101 automatically re - expands the element 1180a to the state shown in FIG. 11B). In some embodiments, the electronic device 101 continues to display the selectable options 1112a - e while moving the application representation 1106a according to the movement of the hand 1104f performing the pinch gesture (e.g., the device 101 maintains the element 1180a in its expanded state shown in FIG. 11B). In some embodiments, the user can select the selectable options 1112a - e while moving the representation 1106a. For example, while the hand is in a pointing gesture with (although not all) one or more fingers extended, the electronic device 101 moves the hand to a location within the 3D environment 1102 within a threshold distance (e.g., 1, 5, 10, 30 centimeters, etc.) of the selected option within the 3D environment 1102, and detects the selection of the options 1112a - e by detecting that the user "pushes" one of the options 1112a - e with the other hand. The fingers of the user's other hand arbitrarily select one of the options 1112a - e (e.g., as described with reference to method 800).

[0218] Accordingly, the electronic device 101 moves an object within the three-dimensional environment 1102 in response to user input directed to the user interface element 1180. Next, the movement of the three-dimensional object 1106c when an input to the user interface element 1180c is detected will be described. Referring again to FIG. 11B, the electronic device 101 detects a user performing a pinch gesture with the hand 1104d while detecting the user's line of sight 1110e on the user interface element 1180c associated with the three-dimensional object 1106c. In response to the detection of the user's line of sight 1110e and the pinch gesture on the user interface element 1180c, the electronic device 101 begins a process of moving the three-dimensional object 1106c within the three-dimensional environment 1102 in accordance with the movement of the hand 1104d while the pinch gesture is maintained. In response to the input for moving the three-dimensional object 1106c, the electronic device 101 shifts up the position of the three-dimensional object 1106c within the three-dimensional environment 1102 (as shown, for example, in the transition from FIG. 11A to FIG. 11B), and in response to the end of the input for moving the object, displays an indication 1114 of the footprint of the three-dimensional object 1106c corresponding to the location where the three-dimensional object will be displayed (e.g., positioned). In some embodiments, the indication 1114 has the same shape as the bottom surface of the three-dimensional object 1106c. In some embodiments, the indication 1114 is displayed so as to appear to be on a surface within the three-dimensional environment (e.g., a representation of a virtual or physical surface in the physical environment of the electronic device 101).

[0219] FIG. 11C shows the electronic device 101 continuing to move the three-dimensional object 1106c in accordance with the movement of the hand 1104e maintaining the pinch gesture. As shown in FIG. 11C, in response to the detection of the end of the movement input, the electronic device 101 continues to display an indication 1114 of the location where the three-dimensional object 1106c will be displayed (e.g., positioned) in the three-dimensional environment 1102.

[0220] In some embodiments, the electronic device 101 shifts the 3D object 1106c upward and displays the instruction 1114 only when the 3D object 1106c is "snapped" to a surface within the 3D environment 101. For example, in FIG. 11A, the object 1106c is "snapped" to the floor of the 3D environment 101 corresponding to the floor of the physical environment of the electronic device 101, and the electronic device 101 displays the 3D object 1106c such that it appears as if the 3D object 1106c is placed on the floor of the 3D environment 1102. In some embodiments, the virtual object "snaps" to the representation of the user's hand within the 3D environment 1102 while the object is being moved. In some embodiments, the representation of the user's hand is either a realistic representation of the hand displayed via the display generation component 120 or a view of the hand through the transparent portion of the display generation component 120. In some embodiments, the 3D object 1106c snaps only to a particular object within the 3D environment 1102, such as the user's hand and / or a surface (e.g., a flat surface or a vertical surface), when the 3D object 1106c is within a predetermined threshold of the object (e.g., 1, 10, 50, 100 centimeters, etc.).

[0221] In some embodiments, while moving the 3D object 1106c, the electronic device 101 updates the appearance of the user interface element 1180c, such as by changing the size or color of the user interface element 1180c. In some embodiments, while moving the 3D object 1106c, the electronic device 101 tilts the 3D object 1106c in response to detecting a rotation of the hand 1104e (e.g., relative to the user's arm). In some embodiments, in response to detecting the end of the movement input, the electronic device 101 displays the 3D object 1106c at the angle (e.g., upright) shown in FIG. 11C even if the user ends the movement input while the 3D object 1106c is tilted in response to the rotation of the hand 1104e. In some embodiments, the end of the movement input is detected in response to detecting that the user stops performing a pinch gesture with the hand 1104e (e.g., by moving the thumb and finger away from each other).

[0222] In some embodiments, while displaying the 3D object 1106c as shown in FIG. 11C, the electronic device 101 detects the end of the movement input. In response thereto, the electronic device 101 displays the 3D object 1106c as shown in FIG. 11D. In FIG. 11D, the electronic device 101 displays the 3D object 1106c at a location within the 3D environment 1102 corresponding to the location of the indication 1114 in FIG. 11C. Thus, in response to detecting the end of the movement input in FIG. 11C, the electronic device 101 optionally lowers the representation 1106c of the 3D environment 1102 by the same amount that the electronic device 101 lifted the representation 1106c in FIG. 11B in response to the start of the movement input, to the area of the footprint 1114.

[0223] Unless otherwise specified, it is understood that any one of the characteristics of any one of the above-described elements 1180a - c is optionally equally applicable to any of the elements 1180a - c.

[0224] Figures 12A - 12O are flowcharts showing a method for moving virtual objects in a 3D environment and facilitating access to actions associated with the virtual objects, according to some embodiments. In some embodiments, method 1200 is executed in a computer system (e.g., computer system 101 of FIG. 1 such as a tablet, smartphone, wearable computer, or head - mounted device) that includes a display generation component (e.g., display generation component 120 of FIGS. 1, 3, and 4) (e.g., a head - up display, a display, a touch screen, a projector, etc.) and one or more cameras (e.g., a camera facing downward with the user's hand (e.g., a color sensor, an infrared sensor, and other depth - sensing cameras), or a camera facing forward from the user's head). In some embodiments, method 1200 is stored on a non - transitory computer - readable storage medium and is executed by instructions executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A). Some operations of method 1200 may be arbitrarily combined and / or the order of some operations may be arbitrarily changed.

[0225] In some embodiments, such as in FIG. 11A, method 1200 is executed on an electronic device 101 that communicates with a display generation component 120 and one or more input devices (e.g., a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer). In some embodiments, the display generation component is an integrated display (optionally a touchscreen display) with the electronic device, an external display such as a monitor, projector, television, or a hardware component (optionally integrated or external) that projects a user interface or makes the user interface visible to one or more users. In some embodiments, the one or more input devices include an electronic device or component that can receive user input (e.g., capture user input, detect user input, etc.) and transmit information related to the user input to the electronic device. Examples of input devices include a touchscreen, a mouse (e.g., external), a trackpad (optionally integrated or external), a touchpad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the electronic device), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, an eye tracking device, and / or a motion sensor (e.g., a hand tracking device, a hand motion sensor). In some embodiments, the electronic device communicates with a hand tracking device (e.g., one or more cameras, a depth sensor, a proximity sensor, a touch sensor (e.g., a touchscreen, a trackpad)). In some embodiments, the hand tracking device is a wearable device such as a smart glove. In some embodiments, the hand tracking device is a handheld input device such as a remote control or a stylus.

[0226] In some embodiments, such as FIG. 11A, an electronic device (e.g., 101) displays (1202a) a user interface (e.g., a computer-generated reality (CGR) environment such as a three-dimensional environment, a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment) that includes, via a display generation component, a first virtual object (e.g., 1106a) (e.g., an application, a window of an application, a virtual object such as a virtual clock, etc.) and a corresponding individual user interface element (e.g., 1180a) that is separate from the first virtual object (e.g., 1106a) and is displayed in relation to the first virtual object (e.g., 1106a). In some embodiments, the corresponding individual user interface element is a horizontal bar that is displayed at a predetermined location relative to the first virtual object, such as below the horizontal center of the first virtual object or overlaid on the first virtual object. The user interface optionally includes a plurality of virtual objects (e.g., applications, windows, etc.), and each virtual object is optionally displayed in relation to its own individual user interface element.

[0227] In some embodiments, such as in FIG. 11A, while displaying a user interface, an electronic device (e.g., 101) detects (1202b) a first user input directed to an individual user interface element (e.g., 1180a) via one or more input devices. In some embodiments, the electronic device detects the first input via an eye tracking device, a hand tracking device, a touch-sensitive surface (e.g., a touch screen or a track pad), a mouse, or a keyboard. For example, the electronic device detects, via an eye tracking device, that the user's line of sight is directed at an individual user interface element. As another example, the electronic device detects, via a hand tracking device, that the user has performed a predetermined gesture (e.g., touching the thumb to another finger of the same hand (e.g., index finger, middle finger, ring finger, little finger)), while detecting, via an eye tracking device, that the user has looked at an individual user interface element. In some embodiments, such as shown in FIG. 11B, in response to the detection of a first user input directed to an individual user interface element (e.g., 1180a) (1202c), and according to a determination that the first user input corresponds to a request to move a first virtual object (e.g., 1106a) within the user interface, the electronic device (e.g., 101) moves (1202d) the first virtual object (e.g., 1106a) and the individual user interface element (e.g., 1180a) within the user interface according to the first user input, as shown in FIG. 11D. In some embodiments, the first user input corresponds to a request to move a first virtual object within the user interface if the user input includes a selection of an individual user interface element followed by a direction input. In some embodiments, the direction input is detected within a threshold time (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 1 second, etc.) for detecting a selection of an individual user interface element.For example, an electronic device detects, via an eye-tracking device, that a user is looking at an individual user interface element, and detects, via a hand-tracking device, that the user is moving their hand while performing a pinch gesture while looking at the individual user interface element. In this example, in response to user input, the electronic device moves a first virtual object and an individual user interface element in accordance with the movement of the user's hand. In some embodiments, the electronic device maintains the position of the individual user interface element relative to the first virtual object while moving the first virtual object and the individual user interface element. In some embodiments, in accordance with a determination that a first user input meets one or more criteria, including criteria that are met when the first user input is an input other than an input for moving a first virtual object (e.g., 1106a) within the user interface in response to detection of the first user input (1202c) directed to an individual user interface element (e.g., 1180a) such as in FIG. 11A, the electronic device (e.g., 101) updates the display of the individual user interface element (e.g., 1180a) to include one or more selectable options (e.g., 1110a-e) that are selectable to perform one or more corresponding actions associated with the first virtual object (e.g., 1106a) such as in FIG. 11B (1202e). In some embodiments, when the first user input does not include a direction input, the first user input meets one or more criteria. In some embodiments, detecting the first user input includes detecting, via an eye-tracking device, that the user is looking at an individual user interface element, and in some embodiments, without detecting a direction input (e.g., without detecting a hand gesture corresponding to a request to move a user interface element via a hand-tracking device), for a period of time longer than a predetermined amount of time (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 1 second), detecting that the user is looking at the individual user interface element.In some embodiments, detecting a first user input includes detecting that the user is looking at an individual user interface element while detecting a predetermined gesture via an eye tracking device or via a hand tracking device. In some embodiments, the predetermined gesture corresponds to a request to select an individual user interface element. In some embodiments, detecting the predetermined gesture includes detecting via the hand tracking device that the user has tapped the thumb and a finger (e.g., index finger, middle finger, ring finger, little finger) together. In some embodiments, updating the display of an individual user interface element to include one or more selectable options includes updating the appearance of the individual user interface element (e.g., increasing the size of the individual user interface element, changing the color, transparency, etc.) and displaying one or more selectable options within the individual user interface element. In some embodiments, one or more of the selectable options, when selected, include an option to cause the electronic device to close a first virtual object (e.g., stop displaying the first virtual object). In some embodiments, one or more of the selectable options, when selected, include an option to cause the electronic device to initiate a process of moving the first virtual object. In some embodiments, one or more of the selectable options, when selected, include an option to cause the electronic device to initiate a process of sharing the first virtual object using an application or sharing protocol accessible to the electronic device (e.g., sharing with / visible to another user in a three-dimensional environment).

[0228] In accordance with the determination that the first user input corresponds to a request to move a first virtual object within the user interface, move the first virtual object and the individual user interface elements, and in accordance with the determination that the first user input meets one or more criteria, update the display of the individual user interface elements to include one or more selectable options. The method described above provides an efficient way to either move an object or gain access to options related to the object, thereby simplifying the interaction between the user and the electronic device (e.g., by reducing the amount of input and time required to move the first virtual object or access options related to the object), improving the operability of the electronic device, making the user-device interface more efficient, and thereby enabling the user to use the electronic device more quickly and efficiently, further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use.

[0229] In some embodiments, such as FIG. 11A, one or more criteria include a criterion that is met when the first user input includes the line of sight (e.g., 1110a) of the user of the electronic device directed at an individual user interface element (e.g., 1180a) for a period longer than a time threshold (1204a) (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 1 second, etc.). In some embodiments, the first user input meets one or more criteria when the electronic device detects the line of sight of the user on the individual user interface element for a period longer than the time threshold without detecting non-line-of-sight input via an eye tracking device or another input device. In some embodiments, in accordance with the determination that the line of sight of the user has become directed away from the individual user interface element while displaying one or more selectable options, the electronic device stops displaying the one or more selectable options. Thus, in some embodiments, in response to detecting the line of sight of the user on the individual user interface element for a period longer than the time threshold, the electronic device displays one or more selectable options.

[0230] The above-described method of displaying one or more selectable options in response to a user's gaze on individual user interface elements provides an efficient way to display individual user interface elements with reduced visual clutter until the user views the individual user interface, and to quickly display selectable options when the user is likely to want to interact with the selectable options based on the user's gaze, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0231] In some embodiments, such as FIG. 11A, one or more criteria are met when a first user input includes an individual gesture that is performed by the hand 704 of a user of an electronic device while the user's line of sight (e.g., 1110a) is directed at an individual user interface element (e.g., 1180a) (1205a). In some embodiments, the gesture is detected using a hand-tracking device. In some embodiments, the gesture is the user touching their thumb with another finger of the same hand (e.g., index finger, middle finger, ring finger, little finger). In some embodiments, one or more criteria are met when the electronic device detects the gesture while detecting the user's line of sight on the individual user interface element for at least a predetermined time threshold (e.g., 0.1, 0.2, 0.3, 0.4, 0.5, 1 second, etc.). In some embodiments, one or more criteria are met when the electronic device detects the gesture while detecting the user's line of sight on the individual user interface element for any amount of time (e.g., less than the time threshold). In some embodiments, the electronic device displays one or more selectable options in response to detecting the user's line of sight on the individual user interface element for a predetermined threshold time (e.g., an input that meets one or more first criteria), or in response to detecting the user's line of sight on the individual user interface element while detecting the gesture (e.g., an input that meets one or more second criteria). In some embodiments, the electronic device displays an option in response to either detecting the user's line of sight on the individual user interface element for a predetermined threshold time or detecting the user's line of sight on the individual user interface element while detecting the gesture, but not both. Thus, in some embodiments, the electronic device displays one or more selectable options in response to detecting the user's line of sight on the individual user interface element while detecting a user performing a predetermined gesture.

[0232] The above-described method of displaying one or more selectable options in response to detecting a user's line of sight on an individual user interface element while the user is performing a gesture displays the one or more selectable options without waiting for a threshold time (e.g., displaying the options only after the user's line of sight has been held on an individual user interface element for a predetermined threshold time), thereby providing an efficient method that allows the user to use the electronic device more quickly and efficiently, improving the battery life of the electronic device and reducing errors during use.

[0233] In some embodiments, such as FIG. 11B, while displaying an individual user interface element (e.g., 1180a) that includes one or more selectable options (e.g., 1112a - e), an electronic device (e.g., 101) detects a second user input (1206a) via one or more input devices (e.g., a hand - tracking device) that includes an individual gesture (e.g., in some embodiments, the gesture is the user touching their thumb with another finger of the same hand (e.g., index finger, middle finger, ring finger, little finger)) performed by the hand (e.g., 1104c) of the user of the electronic device while the user's line of sight (e.g., 1110d) is directed at an individual selectable option (e.g., 1112c) of the one or more selectable options. In some embodiments, in response to detecting the second user input (1206b), such as in FIG. 11B, and according to a determination that the second user input meets one or more first criteria, the electronic device (e.g., 101) performs an individual operation corresponding to the individual selectable option (e.g., 1112c) (1206c). In some embodiments, the one or more first criteria are met according to a determination that the user performs a gesture while looking at the individual selectable option, regardless of the time the user's line of sight is held on the individual selectable option. In some embodiments, in response to detecting the user's line of sight to an individual selectable option, prior to detecting the gesture, the electronic device updates visual characteristics (e.g., size, color, etc.) of the individual selectable option. For example, in response to detecting the user's line of sight to a first selectable option, the electronic device highlights the first selectable option, and in response to detecting the user's line of sight to a second selectable option, the electronic device highlights the second selectable option. Thus, in some embodiments, the electronic device performs an operation corresponding to an individual selectable option in response to detecting the user's line of sight to the individual selectable option while detecting that the user is performing a predetermined gesture by hand.In some embodiments, in accordance with a determination that the second user input does not meet one or more first criteria, the electronic device ceases to perform individual operations corresponding to individual selectable options.

[0234] The above-described method of performing operations related to individual selectable options in response to line-of-sight and non-line-of-sight inputs provides an efficient way to cause the electronic device to perform operations, thereby further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use by enabling the user to use the electronic device more quickly and efficiently.

[0235] In some embodiments, the first user input corresponds to a request to move a first virtual object (e.g., 1106c) within the user interface according to a determination that the first user input includes an individual gesture performed by the hand (e.g., 1104d) of the user of the electronic device while the user's line of sight (e.g., 1110e) is directed at an individual user interface element (e.g., 1180c), as shown in FIG. 11B. Thereafter, movement (1208a) of the hand (e.g., 1104d) of the user of the electronic device continues within a time threshold (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 1 second, etc.) for detecting an individual gesture performed by the hand (e.g., 1104d) of the user. In some embodiments, according to a determination that the user has looked at an individual user interface element for the time threshold without performing a predetermined hand gesture (e.g., touching the thumb to another finger (e.g., index finger, middle finger, ring finger, little finger) of the same hand), the electronic device displays one or more selectable options. In some embodiments, according to a determination that the user has looked at an individual user interface element for the time threshold while performing a predetermined hand gesture (e.g., touching the thumb to another finger (e.g., index finger, middle finger, ring finger, little finger) of the same hand) without moving the hand, the electronic device displays one or more selectable options. In some embodiments, in response to detecting hand movement while performing a gesture within a threshold time for detecting a line of sight on an individual user interface element, the electronic device moves the first virtual object and the individual user interface element according to the movement of the user's hand without displaying one or more selectable options. In some embodiments, the first user input corresponds to a request to move a first virtual object within the user interface according to a determination that the first user input includes an individual gesture performed by the hand of the user of the electronic device while the user's line of sight is directed at an individual user interface element within a time threshold (e.g., 0.1, 0.2, 0.3, 0.4, 0.5, 1 second, etc.) for detecting an individual gesture performed by the hand of the user.In some embodiments, the first user input corresponds to a request to move a first virtual object within the user interface according to a determination that the first user input includes an individual gesture performed by the hand of the user of the electronic device while the user's line of sight is directed at an individual user interface element, and then the movement of the hand of the user of the electronic device continues within a time threshold (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 1 second, etc.) for detecting the user's line of sight on the individual user interface element.

[0236] The above-described method of detecting a request to move a first virtual object in response to detecting hand movement within a threshold period for detecting the user's line of sight on an individual user interface element provides an efficient way to move the first virtual object in accordance with the hand movement without an intermediate step to start the movement of the first virtual object, thereby allowing the user to use the electronic device more quickly and efficiently, further reducing power consumption, improving the battery life of the electronic device, and reducing errors during use.

[0237] In some embodiments, as shown in FIG. 11B, while displaying an individual user interface element that includes one or more selectable options (e.g., 1112a - e), an electronic device (e.g., 101) detects (1210a), via one or more input devices, a second user input corresponding to a request to move a first virtual object (e.g., 1106a) within the user interface, as shown in FIG. 11B. In some embodiments, the one or more selectable options include an option to initiate a process for moving the first virtual object within the user interface. In some embodiments, the request to move the first virtual object is a selection of an option to initiate a process for moving the first virtual object. In some embodiments, while detecting a user performing a hand gesture (e.g., touching the thumb to another finger (e.g., index finger, middle finger, ring finger, little finger) of the same hand as the thumb), in response to detecting the user's line of sight to an option to move the first virtual object, the electronic device selects the option to move the first virtual object. In some embodiments, the request to move the first virtual object within the user interface is not a selection of an option to move the object. For example, the request to move the first virtual object is a user looking at an area of an individual user interface element that has no selectable options displayed for a predetermined time threshold (e.g., 0.02, 0.05, 0.1, 0.2, 0.25, 0.3, 0.5, 1 second, etc.), and / or a user looking at an area of an individual user interface element that has no selectable options displayed while performing a hand gesture (e.g., touching the thumb to another finger (e.g., index finger, middle finger, ring finger, little finger) of the same hand as the thumb). In some embodiments, initiating a process for moving the first virtual object includes moving the first virtual object in accordance with hand movements of the user detected by a hand - tracking device.In some embodiments, in response to the detection of a second user input (1210b), the electronic device (e.g., 101) stops the display of one or more selectable options (e.g., 1112a - e) of FIG. 11B while maintaining the display of an individual user interface element (e.g., 1180a), as shown in FIG. 11D (1210c). In some embodiments, as shown in FIG. 11D, in response to the detection of a second user input (1210b), the electronic device (e.g., 101) moves a first virtual object (e.g., 1106a) and an individual user interface element (e.g., 1106a) within the user interface according to the second user input (1210d). In some embodiments, the second user input includes a movement component (e.g., movement of the hand or arm, movement of the user's eyes, direction input such as selection of a direction key (e.g., arrow keys)), and the electronic device moves the individual user interface element and the first virtual object according to the magnitude of the movement of the second input (e.g., distance, duration, speed). In some embodiments such as FIG. 11D, the electronic device (e.g., 101) detects the end of the second user input via one or more input devices (1210e). In some embodiments, the end of the second user input is that the user stops providing the input. In some embodiments, the end of the second user input is that the user stops performing a predetermined gesture by hand. For example, the predetermined gesture includes the user touching the thumb with another finger of the same hand as the thumb (e.g., index finger, middle finger, ring finger, little finger) (e.g., pinching), and detecting the end of the gesture includes detecting that the user separates the thumb from the finger (e.g., non - pinching). In some embodiments, the end of the second user input is an individual user interface element and / or the first virtual object and / or the user looking outside the user interface.In some embodiments, as shown in FIG. 11B, in response to detecting the end of a second user input, the electronic device (e.g., 101) automatically (e.g., without input to display one or more selectable options) executes one or more corresponding operations associated with the first virtual object (e.g., 1106a) to update the display of an individual user interface element (e.g., 1180a) to include one or more selectable options (e.g., 1112a - e) that are selectable. For example, the electronic devi...

Claims

1. A method, in an electronic device that communicates with a display generation component and one or more input devices, displaying, via the display generation component, a user interface including individual user interface elements having a first appearance, wherein selection of the individual user interface elements causes the electronic device to perform individual operations; while displaying the individual user interface elements having the first appearance, detecting, via the one or more input devices, a first user input including the user's attention directed to the electronic device towards the individual user interface elements based on the posture of the user's physical characteristics; in response to detecting that the user's attention of the electronic device is directed to the individual user interface elements, updating the individual user interface elements so as to visually separate the individual user interface elements in a depth direction from a part of the user interface having a predetermined spatial relationship with the individual user interface elements so as to have a second appearance different from the first appearance, wherein updating the individual user interface elements includes moving the individual user interface elements towards the location corresponding to the user and away from the part of the user interface by reducing the distance between the individual user interface elements and the location corresponding to the user and increasing the distance between the individual user interface elements and the part of the user interface; while the individual user interface elements have the second appearance, detecting, via the one or more input devices, a second user input corresponding to a gesture performed by the user's hand corresponding to activation of the individual user interface elements based on the posture of the user's physical characteristics; in response to detecting the second user input directed to the individual user interface elements, according to a determination that the second user input satisfies one or more second criteria indicating that the second user input corresponds to selection of the individual user interface elements Performing the individual operation corresponding to the selection of the individual user interface element, Updating the individual user interface element by reducing a separation amount in a depth direction between the individual user interface element and the part of the user interface having the predetermined spatial relationship with respect to the individual user interface element, where updating the individual user interface element includes moving the individual user interface element away from the location corresponding to the user and toward the part of the user interface by increasing a distance between the individual user interface element and the location corresponding to the user and decreasing a distance between the individual user interface element and the part of the user interface, Ceasing to perform the individual operation corresponding to the selection of the individual user interface element according to a determination that the second user input does not meet the one or more second criteria while it is still determined that the user's attention is directed to the individual user interface element, without reducing the separation amount in the depth direction between the individual user interface element and the part of the user interface having the predetermined spatial relationship with respect to the individual user interface element, the method comprising:

2. Detecting, via the one or more input devices, that the user's attention of the electronic device is not directed to the individual user interface element based on a posture of the physical characteristics of the user while the individual user interface element has the second appearance, Updating the individual user interface element by reducing a separation amount between the individual user interface element and the part of the user interface having the predetermined spatial relationship with respect to the individual user interface element in response to detecting that the user's attention of the electronic device is not directed to the individual user interface element, The method according to claim 1, further comprising:

3. The second user input meets the one or more second criteria, and the method further, While detecting the second user input directed to the individual user interface element and before the second user input satisfies the one or more second criteria, in accordance with the progress of the second user input towards satisfying the one or more second criteria, updating the individual user interface element by reducing the separation amount in the depth direction between the individual user interface element and the portion of the user interface having the predetermined spatial relationship with the individual user interface element. The method according to any one of claims 1 to 2, including.

4. While displaying the individual user interface element having the first appearance, detecting the first user input including the attention of the user of the electronic device directed to the individual user interface element based on the posture of the physical characteristics of the user, through an eye tracking device communicating with the electronic device, detecting that the user's line of sight is directed to the individual user interface element. The method according to any one of claims 1 to 3, including.

5. While displaying the individual user interface element having the first appearance, detecting the first user input including the attention of the user of the electronic device directed to the individual user interface element based on the posture of the physical characteristics of the user, through an eye tracking device and a hand tracking device communicating with the electronic device, detecting that the user's line of sight is directed to the individual user interface element and the user's hand is in a predetermined posture. The method according to any one of claims 1 to 4, including.

6. While displaying the individual user interface element having the second appearance, detecting the second user input corresponding to the activation of the individual user interface element based on the posture of the physical characteristics of the user, through a hand tracking device communicating with the electronic device, detecting a part of the user's hand of the electronic device at the location corresponding to the individual user interface element. The method according to any one of claims 1 to 5, including.

7. While displaying the individual user interface element having the second appearance, detecting the second user input corresponding to the activation of the individual user interface element based on the posture of the physical characteristics of the user, including detecting an individual gesture performed by the hand of the user of the electronic device while the line of sight of the user of the electronic device is directed at the individual user interface element via an eye tracking device and a hand tracking device that communicate with the electronic device. The method according to any one of claims 1 to 6. [

8. ] Before detecting the second user input directed at the individual user interface element, the individual user interface element is displayed by an individual visual characteristic having a first value while the individual user interface element is visually separated from the part of the user interface. Updating the individual user interface element includes displaying the individual user interface element having an individual visual characteristic having a second value different from the first value while the separation amount between the individual user interface element and the part of the user interface is reduced. The method according to any one of claims 1 to 7, including. [

9. ] If the second user input includes the line of sight of the user of the electronic device directed at the individual user interface element for a time longer than a time threshold, the second user input satisfies the one or more second criteria. The method according to any one of claims 1 to 8. [

10. ] While the individual user interface element has the second appearance, detecting, via a hand tracking device that communicates with the electronic device, that the hand of the user of the electronic device is in an individual location corresponding to a location for interacting with the individual user interface element; In response to detecting that the hand of the user of the electronic device is in the individual location, updating the individual user interface element to further visually separate the individual user interface element from the part of the user interface having the predetermined spatial relationship to the individual user interface element. The method according to any one of claims 1 to 9, further comprising **Claim 11** wherein the individual user interface element having the second appearance is associated with a first hierarchical level within the user interface, and a part of the user interface having the predetermined spatial relationship with respect to the individual user interface element is associated with a second hierarchical level different from the first hierarchical level The method according to any one of claims 1 to 10. **Claim 12** detecting the second user input includes detecting a user input from the user of the electronic device corresponding to a movement of the individual user interface element back towards the part of the user interface via a hand tracking device communicating with the electronic device, and the method further comprises updating the individual user interface element to reduce a separation amount between the individual user interface element and the part of the user interface in response to detection of the second user input the second user input satisfies the one or more second criteria when the hand input corresponds to a movement of the individual user interface element within a threshold distance from the part of the user interface The method according to any one of claims 1 to 11. **Claim 13** after the second user input satisfies the one or more second criteria and while the individual user interface element is within the threshold distance from the part of the user interface, detecting a further user input from the user of the electronic device corresponding to the movement of the individual user interface element back towards the part of the user interface via the hand tracking device; and moving the individual user interface element and the part of the user interface according to the further user input in response to detection of the further user input The method according to claim 12, further comprising **Claim 14** in response to detecting the second user input In accordance with the determination that the manual input corresponds to the movement of the individual user interface element towards the part of the user interface that is less than the threshold movement amount, moving the individual user interface element towards the part of the user interface according to the manual input without moving the part of the user interface, and reducing the separation amount between the individual user interface element and the part of the user interface, In accordance with the determination that the manual input corresponds to the movement of the individual user interface element returning towards the part of the user interface that is greater than the threshold movement amount, moving the individual user interface element and moving the part of the user interface according to the manual input, The method according to any one of claims 12 to 13, further comprising.

15. Updating the individual user interface element by reducing the separation amount between the individual user interface element and the part of the user interface includes moving the individual user interface element and the part of the user interface inertially according to the movement component of the second user input, and the method further comprises Detecting the end of the second user input directed to the individual user interface element, In response to detecting the end of the second user input directed to the individual user interface element, moving the individual user interface element and the part of the user interface in a direction opposite to the movement of the individual user interface element and the part of the user interface according to the second user input. The method according to any one of claims 1 to 14, comprising.

16. Detecting the second user input includes detecting a part of the user's hand of the electronic device at a location corresponding to the individual user interface element, and the method further comprises While the individual user interface element has the second appearance, detecting an individual input that includes an individual gesture made by the user's hand while the user's hand is in a location that does not correspond to the individual user interface element via a hand-tracking device that communicates with the electronic device; In response to detecting the individual input; Updating the individual user interface element by reducing a separation amount between the individual user interface element and the portion of the user interface, including moving the individual user interface element and the portion of the user interface inertially according to a determination that the individual input based on the individual gesture made by the user's hand while the user's hand is in the location that does not correspond to the individual user interface element meets one or more third criteria; Detecting an end of the individual input; In response to detecting the end of the individual input, moving the individual user interface element and the portion of the user interface in a direction opposite to the movement of the individual user interface element and the portion of the user interface according to the individual input, the method according to claim 15.

17. Detecting the second user input includes detecting a part of the user's hand of the electronic device at a location corresponding to the individual user interface element, and the method further includes While the individual user interface element has the second appearance, detecting an individual input that includes the user's line of sight directed at the individual user interface element via an eye-tracking device that communicates with the electronic device; In response to detecting the individual input; Updating the individual user interface element by reducing a separation amount between the individual user interface element and the portion of the user interface, including moving the individual user interface element and the portion of the user interface inertially according to a determination that the individual input based on the user's line of sight directed at the individual user interface element meets one or more third criteria; detecting the end of the individual input; in response to detecting the end of the individual input, moving the individual user interface element and the portion of the user interface in a direction opposite to the movement of the individual user interface element and the portion of the user interface according to the individual input, the method according to any one of claims 15 to 16. **Claim 18** one or more processors; a memory; one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the method according to any one of claims 1 to 17, an electronic device. **Claim 19** a program including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • projector

    JP2005215144A

  • Control apparatus, electronic apparatus, control method, and program

    JP2013196158A

  • Devices, methods, and graphical user interfaces for sharing content objects within a document.

    JP2015515040A

  • Information processing apparatus, virtual object operation method, computer program, and storage medium

    JP2017027206A

  • Terminal device, and program

    JP2018005516A