Methods for interacting with objects in environment

The computer system addresses inefficiencies in augmented and virtual reality interactions by using touch-sensitive and gaze-tracking components to reduce user inputs and enhance interaction efficiency, improving battery life.

JP2025124636APending Publication Date: 2025-08-26APPLE INC

Patent Information

Application Number
JP2025072340
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-23
Filing Date
2025-04-24
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Existing methods and interfaces for interacting with augmented, mixed, and virtual reality environments are cumbersome, inefficient, and complex, leading to a significant cognitive burden on users and inefficient use of battery-operated devices.

Method used

A computer system with enhanced input methods and interfaces that reduce the number and type of user inputs by utilizing touch-sensitive displays, eye-tracking, hand-tracking, and gaze-tracking components, along with visual feedback, to enhance interaction efficiency and intuitiveness.

Benefits of technology

The system provides more efficient and intuitive interaction with virtual objects by reducing the complexity of user inputs, enhancing interaction with user interface elements in three-dimensional environments, and improving battery life through reduced power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025124636000001_ABST
    Figure 2025124636000001_ABST
Patent Text Reader

Abstract

To provide methods for making interaction with computer systems more efficient and intuitive for a user.SOLUTION: A method comprises, in an electronic device: displaying user interface objects in a three-dimensional environment through a display generation component; detecting an individual input including movement of a predetermined portion of a user of the electronic device while displaying the user interface objects; displaying a visual indication at a first location in the three-dimensional environment corresponding to a first position while detecting the individual input and in accordance with determination that a first portion satisfies one or more criteria and that the predetermined portion of the user is at the first position; and displaying the visual indication at a second location in the three-dimensional environment corresponding to a second position in accordance with determination that the first portion satisfies the one or more criteria and that the predetermined portion of the user is at the second position.SELECTED DRAWING: Figure 18A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 139,566, filed January 20, 2021, and U.S. Provisional Patent Application No. 63 / 261,559, filed September 23, 2021, the contents of which are incorporated herein by reference in their entirety for all purposes.

[0002] It generally relates to a computer system having a display generation component and one or more input devices that present a graphical user interface, including but not limited to electronic devices that present interactive user interface elements via the display generation component. [Background technology]

[0003] The development of computer systems for augmented reality has progressed significantly in recent years. Exemplary augmented reality environments include at least some virtual elements that replace or augment the physical world. Input devices such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays for computer systems and other electronic computing devices are used to interact with the virtual / augmented reality environment. Exemplary virtual elements include virtual objects, including digital images, video, text, icons, and control elements such as buttons and other graphics.

[0004] However, methods and interfaces for interacting with environments (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) that include at least some virtual elements are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems in which manipulating virtual objects is complex and error-prone create a significant cognitive burden for users and detract from the experience of the virtual / augmented reality environment. In addition, these methods are unnecessarily time-consuming, thereby wasting energy. This latter consideration is particularly important in battery-operated devices. Summary of the Invention

[0005] Therefore, there is a need for a computer system having improved methods and interfaces for providing users with computer-generated experiences that make interaction with the computer system more efficient and intuitive for the user. Such methods and interfaces can optionally complement or replace conventional methods of providing users with computer-generated reality experiences. Such methods and interfaces reduce the number, extent, and / or type of inputs from the user by helping the user understand the connection between the input provided and the device response to that input, thereby creating a more efficient human-machine interface.

[0006] The above-mentioned deficiencies and other problems associated with user interfaces for computer systems having a display generation component and one or more input devices are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a "touch screen" or "touchscreen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to the display generation component, the output devices including one or more tactile output generators and one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in the memory for performing a plurality of functions. In some embodiments, a user interacts with the GUI through stylus and / or finger contacts and gestures on the touch-sensitive surface, the movement of the user's eyes and hands in space relative to the GUI or the user's body as captured by cameras and other movement sensors, and voice input as captured by one or more audio input devices.In some embodiments, the functions performed through the interactions optionally include image editing, drawing, presenting, word processing, spreadsheet creation, game playing, making phone calls, video conferencing, emailing, instant messaging, training support, digital photography, digital videography, web browsing, digital music playback, note taking, and / or digital video playback, and executable instructions to perform those functions are optionally contained on a non-transitory computer-readable storage medium or other computer program product configured to be executed by one or more processors.

[0007] There is a need for electronic devices with improved methods and interfaces for interacting with objects in a three-dimensional environment. Such methods and interfaces can complement or replace conventional methods for interacting with objects in a three-dimensional environment. Such methods and interfaces reduce the number, extent, and / or type of input from a user, creating a more efficient human-machine interface.

[0008] In some embodiments, the electronic device performs or does not perform an action in response to a user input depending on whether detecting a user readiness state precedes the user input. In some embodiments, the electronic device processes the user input based on a zone of attention associated with the user. In some embodiments, the electronic device enhances interaction with user interface elements at different distances and / or angles relative to the user's line of sight in a three-dimensional environment. In some embodiments, the electronic device enhances interaction with user interface elements for mixed direct and indirect interaction modes. In some embodiments, the electronic device manages input from both hands of a user. In some embodiments, the electronic device presents a visual indication of the user input. In some embodiments, the electronic device enhances interaction with user interface elements in a three-dimensional environment using a visual indication of such interaction. In some embodiments, the electronic device redirects input from one user interface element to another according to a movement involved in the input.

[0009] It should be noted that the various embodiments described above can be combined with any other embodiment described herein. The features and advantages described herein are not exhaustive, and many additional features and advantages will become apparent to those skilled in the art, particularly in light of the drawings, specification, and claims. Furthermore, it should be noted that the language used herein has been selected solely for the purposes of readability and explanation, and not to define or limit the subject matter of the present invention.

[0010] For a better understanding of the various described embodiments, reference should be made to the following Detailed Description of the Invention in conjunction with the following drawings, in which like reference numerals refer to corresponding parts throughout: [Brief explanation of the drawings]

[0011] [Figure 1]FIG. 1 is a block diagram illustrating a computer system operating environment for providing a CGR experience, according to some embodiments.

[0012] [Figure 2] FIG. 1 is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's CGR experience, according to some embodiments.

[0013] [Figure 3] FIG. 1 is a block diagram illustrating display generation components of a computer system configured to provide a user with visual components of a CGR experience, according to some embodiments.

[0014] [Figure 4] FIG. 1 is a block diagram illustrating a hand tracking unit of a computer system configured to capture a user's gesture input, according to some embodiments.

[0015] [Figure 5] FIG. 1 is a block diagram illustrating an eye-tracking unit of a computer system configured to capture a user's gaze input, according to some embodiments.

[0016] [Figure 6A] 1 is a flowchart illustrating a glint-assisted gaze tracking pipeline, according to some embodiments.

[0017] [Figure 6B] 1 illustrates an exemplary environment for an electronic device for providing a CGR experience, according to some embodiments.

[0018] [Figure 7A] 1 illustrates an exemplary way in which an electronic device performs or does not perform an action in response to a user input depending on whether detecting a user ready state precedes the user input, according to some embodiments. [Figure 7B] 1 illustrates an exemplary way in which an electronic device performs or does not perform an action in response to a user input depending on whether detecting a user ready state precedes the user input, according to some embodiments. [Figure 7C] 1 illustrates an exemplary way in which an electronic device performs or does not perform an action in response to a user input depending on whether detecting a user ready state precedes the user input, according to some embodiments.

[0019] [Figure 8A] 1 is a flowchart illustrating a method for performing or not performing an action in response to a user input depending on whether the user input is preceded by detecting a user ready state, according to some embodiments. [Figure 8B] 1 is a flowchart illustrating a method for performing or not performing an action in response to a user input depending on whether the user input is preceded by detecting a user ready state, according to some embodiments. [Figure 8C] 1 is a flowchart illustrating a method for performing or not performing an action in response to a user input depending on whether the user input is preceded by detecting a user ready state, according to some embodiments. [Figure 8D] 1 is a flowchart illustrating a method for performing or not performing an action in response to a user input depending on whether the user input is preceded by detecting a user ready state, according to some embodiments. [Figure 8E] 1 is a flowchart illustrating a method for performing or not performing an action in response to a user input depending on whether the user input is preceded by detecting a user ready state, according to some embodiments. [Figure 8F] 1 is a flowchart illustrating a method for performing or not performing an action in response to a user input depending on whether the user input is preceded by detecting a user ready state, according to some embodiments. [Figure 8G]1 is a flowchart illustrating a method for performing or not performing an action in response to a user input depending on whether the user input is preceded by detecting a user ready state, according to some embodiments. [Figure 8H] 1 is a flowchart illustrating a method for performing or not performing an action in response to a user input depending on whether the user input is preceded by detecting a user ready state, according to some embodiments. [Figure 8I] 1 is a flowchart illustrating a method for performing or not performing an action in response to a user input depending on whether the user input is preceded by detecting a user ready state, according to some embodiments. [Figure 8J] 1 is a flowchart illustrating a method for performing or not performing an action in response to a user input depending on whether the user input is preceded by detecting a user ready state, according to some embodiments. [Figure 8K] 1 is a flowchart illustrating a method for performing or not performing an action in response to a user input depending on whether the user input is preceded by detecting a user ready state, according to some embodiments.

[0020] [Figure 9A] 1 illustrates an example method by which an electronic device processes user input based on a zone of attention associated with the user, according to some embodiments. [Figure 9B] 1 illustrates an example method by which an electronic device processes user input based on a zone of attention associated with the user, according to some embodiments. [Figure 9C] 1 illustrates an example method by which an electronic device processes user input based on a zone of attention associated with the user, according to some embodiments.

[0021] [Figure 10A] 1 is a flowchart illustrating a method for processing user input based on a zone of attention associated with a user, according to some embodiments. [Figure 10B] 1 is a flowchart illustrating a method for processing user input based on a zone of attention associated with a user, according to some embodiments. [Figure 10C] 1 is a flowchart illustrating a method for processing user input based on a zone of attention associated with a user, according to some embodiments. [Figure 10D] 1 is a flowchart illustrating a method for processing user input based on a zone of attention associated with a user, according to some embodiments. [Figure 10E] 1 is a flowchart illustrating a method for processing user input based on a zone of attention associated with a user, according to some embodiments. [Figure 10F] 1 is a flowchart illustrating a method for processing user input based on a zone of attention associated with a user, according to some embodiments. [Figure 10G] 1 is a flowchart illustrating a method for processing user input based on a zone of attention associated with a user, according to some embodiments. [Figure 10H] 1 is a flowchart illustrating a method for processing user input based on a zone of attention associated with a user, according to some embodiments.

[0022] [Figure 11A] 10A-10C illustrate examples of how an electronic device can enhance interaction with user interface elements at different distances and / or angles relative to a user's line of sight in a three-dimensional environment, according to some embodiments. [Figure 11B] 10A-10C illustrate examples of how an electronic device can enhance interaction with user interface elements at different distances and / or angles relative to a user's line of sight in a three-dimensional environment, according to some embodiments. [Figure 11C] 10A-10C illustrate examples of how an electronic device can enhance interaction with user interface elements at different distances and / or angles relative to a user's line of sight in a three-dimensional environment, according to some embodiments.

[0023] [Figure 12A] 1 is a flowchart illustrating a method for enhancing interaction with user interface elements at different distances and / or angles relative to a user's line of sight in a three-dimensional environment, according to some embodiments. [Figure 12B] 1 is a flowchart illustrating a method for enhancing interaction with user interface elements at different distances and / or angles relative to a user's line of sight in a three-dimensional environment, according to some embodiments. [Figure 12C] 1 is a flowchart illustrating a method for enhancing interaction with user interface elements at different distances and / or angles relative to a user's line of sight in a three-dimensional environment, according to some embodiments. [Figure 12D] 1 is a flowchart illustrating a method for enhancing interaction with user interface elements at different distances and / or angles relative to a user's line of sight in a three-dimensional environment, according to some embodiments. [Figure 12E] 1 is a flowchart illustrating a method for enhancing interaction with user interface elements at different distances and / or angles relative to a user's line of sight in a three-dimensional environment, according to some embodiments. [Figure 12F] 1 is a flowchart illustrating a method for enhancing interaction with user interface elements at different distances and / or angles relative to a user's line of sight in a three-dimensional environment, according to some embodiments.

[0024] [Figure 13A] 10 illustrates an example of how an electronic device enhances interaction with user interface elements for mixed direct and indirect interaction modes, according to some embodiments. [Figure 13B] 10 illustrates an example of how an electronic device enhances interaction with user interface elements for mixed direct and indirect interaction modes, according to some embodiments. [Figure 13C]10 illustrates an example of how an electronic device enhances interaction with user interface elements for mixed direct and indirect interaction modes, according to some embodiments.

[0025] [Figure 14A] 1 is a flowchart illustrating a method for enhancing interaction with a user interface element for mixed direct and indirect interaction modes, according to some embodiments. [Figure 14B] 1 is a flowchart illustrating a method for enhancing interaction with a user interface element for mixed direct and indirect interaction modes, according to some embodiments. [Figure 14C] 1 is a flowchart illustrating a method for enhancing interaction with a user interface element for mixed direct and indirect interaction modes, according to some embodiments. [Figure 14D] 1 is a flowchart illustrating a method for enhancing interaction with a user interface element for mixed direct and indirect interaction modes, according to some embodiments. [Figure 14E] 1 is a flowchart illustrating a method for enhancing interaction with a user interface element for mixed direct and indirect interaction modes, according to some embodiments. [Figure 14F] 1 is a flowchart illustrating a method for enhancing interaction with a user interface element for mixed direct and indirect interaction modes, according to some embodiments. [Figure 14G] 1 is a flowchart illustrating a method for enhancing interaction with a user interface element for mixed direct and indirect interaction modes, according to some embodiments. [Figure 14H] 1 is a flowchart illustrating a method for enhancing interaction with a user interface element for mixed direct and indirect interaction modes, according to some embodiments.

[0026] [Figure 15A]1 illustrates an exemplary method by which an electronic device manages input from both hands of a user, according to some embodiments. [Figure 15B] 1 illustrates an exemplary method by which an electronic device manages input from both hands of a user, according to some embodiments. [Figure 15C] 1 illustrates an exemplary method by which an electronic device manages input from both hands of a user, according to some embodiments. [Figure 15D] 1 illustrates an exemplary method by which an electronic device manages input from both hands of a user, according to some embodiments. [Figure 15E] 1 illustrates an exemplary method by which an electronic device manages input from both hands of a user, according to some embodiments.

[0027] [Figure 16A] 1 is a flowchart illustrating a method for managing input from both hands of a user, according to some embodiments. [Figure 16B] 1 is a flowchart illustrating a method for managing input from both hands of a user, according to some embodiments. [Figure 16C] 1 is a flowchart illustrating a method for managing input from both hands of a user, according to some embodiments. [Figure 16D] 1 is a flowchart illustrating a method for managing input from both hands of a user, according to some embodiments. [Figure 16E] 1 is a flowchart illustrating a method for managing input from both hands of a user, according to some embodiments. [Figure 16F] 1 is a flowchart illustrating a method for managing input from both hands of a user, according to some embodiments. [Figure 16G] 1 is a flowchart illustrating a method for managing input from both hands of a user, according to some embodiments. [Figure 16H] 1 is a flowchart illustrating a method for managing input from both hands of a user, according to some embodiments. [Figure 16I]1 is a flowchart illustrating a method for managing input from both hands of a user, according to some embodiments.

[0028] [Figure 17A] 1 illustrates various ways in which an electronic device may present a visual indication of a user input, according to some embodiments. [Figure 17B] 1 illustrates various ways in which an electronic device may present a visual indication of a user input, according to some embodiments. [Figure 17C] 1 illustrates various ways in which an electronic device may present a visual indication of a user input, according to some embodiments. [Figure 17D] 1 illustrates various ways in which an electronic device may present a visual indication of a user input, according to some embodiments. [Figure 17E] 1 illustrates various ways in which an electronic device may present a visual indication of a user input, according to some embodiments.

[0029] [Figure 18A] 1 is a flowchart illustrating a method for presenting a visual indication of a user input, according to some embodiments. [Figure 18B] 1 is a flowchart illustrating a method for presenting a visual indication of a user input, according to some embodiments. [Figure 18C] 1 is a flowchart illustrating a method for presenting a visual indication of a user input, according to some embodiments. [Figure 18D] 1 is a flowchart illustrating a method for presenting a visual indication of a user input, according to some embodiments. [Figure 18E] 1 is a flowchart illustrating a method for presenting a visual indication of a user input, according to some embodiments. [Figure 18F] 1 is a flowchart illustrating a method for presenting a visual indication of a user input, according to some embodiments. [Figure 18G]1 is a flowchart illustrating a method for presenting a visual indication of a user input, according to some embodiments. [Figure 18H] 1 is a flowchart illustrating a method for presenting a visual indication of a user input, according to some embodiments. [Figure 18I] 1 is a flowchart illustrating a method for presenting a visual indication of a user input, according to some embodiments. [Figure 18J] 1 is a flowchart illustrating a method for presenting a visual indication of a user input, according to some embodiments. [Figure 18K] 1 is a flowchart illustrating a method for presenting a visual indication of a user input, according to some embodiments. [Figure 18L] 1 is a flowchart illustrating a method for presenting a visual indication of a user input, according to some embodiments. [Figure 18M] 1 is a flowchart illustrating a method for presenting a visual indication of a user input, according to some embodiments. [Figure 18N] 1 is a flowchart illustrating a method for presenting a visual indication of a user input, according to some embodiments. [Figure 18O] 1 is a flowchart illustrating a method for presenting a visual indication of a user input, according to some embodiments.

[0030] [Figure 19A] 1 illustrates an example of how an electronic device can enhance interaction with user interface elements in a three-dimensional environment using visual indications of such interaction, according to some embodiments. [Figure 19B] 1 illustrates an example of how an electronic device can enhance interaction with user interface elements in a three-dimensional environment using visual indications of such interaction, according to some embodiments. [Figure 19C]1 illustrates an example of how an electronic device can enhance interaction with user interface elements in a three-dimensional environment using visual indications of such interaction, according to some embodiments. [Figure 19D] 1 illustrates an example of how an electronic device can enhance interaction with user interface elements in a three-dimensional environment using visual indications of such interaction, according to some embodiments.

[0031] [Figure 20A] 1 is a flowchart illustrating a method for enhancing interaction with user interface elements in a three-dimensional environment using visual indications of such interaction, according to some embodiments. [Figure 20B] 1 is a flowchart illustrating a method for enhancing interaction with user interface elements in a three-dimensional environment using visual indications of such interaction, according to some embodiments. [Figure 20C] 1 is a flowchart illustrating a method for enhancing interaction with user interface elements in a three-dimensional environment using visual indications of such interaction, according to some embodiments. [Figure 20D] 1 is a flowchart illustrating a method for enhancing interaction with user interface elements in a three-dimensional environment using visual indications of such interaction, according to some embodiments. [Figure 20E] 1 is a flowchart illustrating a method for enhancing interaction with user interface elements in a three-dimensional environment using visual indications of such interaction, according to some embodiments. [Figure 20F] 1 is a flowchart illustrating a method for enhancing interaction with user interface elements in a three-dimensional environment using visual indications of such interaction, according to some embodiments.

[0032] [Figure 21A]10A-10C illustrate examples of how an electronic device redirects input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments. [Figure 21B] 10A-10C illustrate examples of how an electronic device redirects input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments. [Figure 21C] 10A-10C illustrate examples of how an electronic device redirects input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments. [Figure 21D] 10A-10C illustrate examples of how an electronic device redirects input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments. [Figure 21E] 10A-10C illustrate examples of how an electronic device redirects input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments.

[0033] [Figure 22A] 1 is a flowchart illustrating a method for redirecting input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments. [Figure 22B] 1 is a flowchart illustrating a method for redirecting input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments. [Figure 22C] 1 is a flowchart illustrating a method for redirecting input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments. [Figure 22D] 1 is a flowchart illustrating a method for redirecting input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments. [Figure 22E] 1 is a flowchart illustrating a method for redirecting input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments. [Figure 22F] 1 is a flowchart illustrating a method for redirecting input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments. [Figure 22G] 1 is a flowchart illustrating a method for redirecting input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments. [Figure 22H] 1 is a flowchart illustrating a method for redirecting input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments. [Figure 22I] 1 is a flowchart illustrating a method for redirecting input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments. [Figure 22J] 1 is a flowchart illustrating a method for redirecting input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments. [Figure 22K] 1 is a flowchart illustrating a method for redirecting input from one user interface element to another in response to detecting a movement contained in the input, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0034] The present disclosure relates to a user interface that provides a computer-generated reality (CGR) experience to a user, according to some embodiments.

[0035] The systems, methods, and GUIs described herein provide improved ways for electronic devices to interact with and manipulate objects in a three-dimensional environment, which optionally includes one or more virtual objects, one or more representations of real objects in the physical environment of the electronic device (e.g., displayed as photo-realistic (e.g., "pass-through") representations of the real objects or visible to the user through transparent portions of a display generating component), and / or a representation of a user within the three-dimensional environment.

[0036] In some embodiments, the electronic device automatically updates the orientation of the virtual object within the three-dimensional environment based on the user's viewpoint within the three-dimensional environment. In some embodiments, the electronic device moves the virtual object according to user input and displays the object at the updated location upon completion of the user input. In some embodiments, the electronic device automatically updates the orientation of the virtual object at the updated location (e.g., and / or as the virtual object moves to the updated location) so that the virtual object is oriented toward the user's viewpoint within the three-dimensional environment (e.g., throughout and / or at the end of its movement). Automatically updating the orientation of a virtual object within the three-dimensional environment allows a user to view and interact with the virtual object more naturally and efficiently, without requiring the user to manually adjust the object's orientation.

[0037] In some embodiments, the electronic device automatically updates the orientation of the virtual object in the three-dimensional environment based on the viewpoints of multiple users in the three-dimensional environment. In some embodiments, the electronic device moves the virtual object according to user input and displays the object at the updated location upon completion of the user input. In some embodiments, the electronic device automatically updates the orientation of the virtual object at the updated location (e.g., and / or as the virtual object moves to the updated location) so that the virtual object is oriented toward the viewpoints of multiple users in the three-dimensional environment (e.g., throughout and / or at the end of its movement). Automatically updating the orientation of a virtual object in a three-dimensional environment allows a user to view and interact with the virtual object more naturally and efficiently, without requiring the user to manually adjust the orientation of the object.

[0038] In some embodiments, the electronic device modifies the appearance of real objects that are between the virtual object and the user's viewpoint in the three-dimensional environment. The electronic device optionally blurs, darkens, or otherwise modifies portions of real objects (e.g., displayed as photorealistic (e.g., "pass-through") representations of the real objects or visible to the user through transparent portions of the display generation components) that are between the user's viewpoint and the virtual object in the three-dimensional environment. In some embodiments, the electronic device modifies portions of real objects that are within a threshold distance (e.g., 5, 10, 30, 50, 100 centimeters, etc.) from the boundary of the virtual object without modifying portions of the real object that are beyond the threshold distance from the boundary of the virtual object. Modifying the appearance of real objects allows the user to view and interact with virtual objects more naturally and efficiently. Furthermore, modifying the appearance of real objects reduces cognitive burden on the user.

[0039] In some embodiments, the electronic device automatically selects a location for the user within a three-dimensional environment that includes one or more virtual objects and / or other users. In some embodiments, the user gains access to a three-dimensional environment that already includes one or more other users and one or more virtual objects. In some embodiments, the electronic device automatically selects a location to associate the user with (e.g., a location to place the user's viewpoint) based on the locations and orientations of the virtual objects and other users within the three-dimensional environment. In some embodiments, the electronic device selects the user's location to allow the user to view other users and virtual objects within the three-dimensional environment without obstructing the other users' views of the user and virtual objects. Automatically placing the user within the three-dimensional environment based on the locations and orientations of the virtual objects and other users within the three-dimensional environment allows the user to efficiently view and interact with virtual objects and other users within the three-dimensional environment without having to manually select locations within the three-dimensional environment with which the user should be associated.

[0040] In some embodiments, the electronic device redirects the input from one user interface element to another according to movement included in the input. In some embodiments, the electronic device presents multiple interactive user interface elements and accepts input directed to a first of the multiple user interface elements via one or more input devices. In some embodiments, after detecting a portion of the input (e.g., without detecting the entire input), the electronic device detects a movement portion of the input that corresponds to a request to redirect the input to a second user interface element. In response, in some embodiments, the electronic device directs the input to the second user interface element. In some embodiments, in response to movement meeting one or more criteria (e.g., based on speed, duration, distance, etc.), the electronic device cancels the input instead of redirecting the input. Allowing a user to redirect or cancel an input after providing a portion of the input allows a user to efficiently interact with the electronic device with fewer inputs (e.g., to undo an unintended action and / or to direct the input to a different user interface element).

[0041] 1-6 provide a description of an exemplary computer system for providing a CGR experience to a user (as described below with respect to methods 800, 1000, 1200, 1400, 1600, 1800, 2000, and 2200). In some embodiments, the CGR experience is provided to a user via an operating environment 100 that includes a computer system 101, as shown in FIG. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a touch sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., a home appliance, a wearable device, etc.). In some embodiments, one or more of the input device 125, the output device 155, the sensor 190, and the peripheral device 195 are integrated with the display generation component 120 (e.g., within a head-mounted or handheld device).

[0042] The processes described below enhance the usability of a device and streamline the user-device interface (e.g., by helping the user provide appropriate inputs when operating / interacting with the device and reducing user errors) through various techniques, including providing improved visual feedback to the user, reducing the number of inputs required to perform an action, providing additional control options without cluttering the user interface with additional controls that are displayed, performing an action without requiring further user input when a set of conditions is met, and / or other techniques. These techniques also reduce power usage and improve the device's battery life by allowing the user to use the device more quickly and efficiently.

[0043] When describing a CGR experience, various terms are used to individually refer to several related, but distinct, environments that a user senses and / or with which the user can interact (e.g., using inputs detected by computer system 101 that cause the computer system generating the CGR experience to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to computer system 101 generating the CGR experience). The following is a subset of these terms:

[0044] Physical Environment: The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. A physical environment, such as a physical park, includes physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through their senses, such as sight, touch, hearing, taste, and smell.

[0045] Computer-Generated Reality: In contrast, a computer-generated reality (CGR) environment refers to a wholly or partially mimicked environment that people sense and / or interact with via electronic systems. In a CGR, a subset of a person's physical movements or representations thereof are tracked, and one or more properties of one or more virtual objects simulated within the CGR environment are adjusted accordingly to behave according to at least one law of physics. For example, a CGR system may detect a person's head rotation and adjust the graphical content and sound field presented to the person accordingly, in a manner similar to how such views and sounds change in a physical environment. In some circumstances (e.g., for accessibility reasons), adjustments to the property(ies) of a virtual object(s) in a CGR environment may be made in response to a representation of a physical movement (e.g., a voice command). A person may sense and / or interact with a CGR object using any one of these senses, including sight, hearing, touch, taste, and smell. For example, a person may sense and / or interact with audio objects that create a 3D or spatially expansive audio environment, providing the perception of a point sound source in 3D space. In another example, audio objects may enable audio transparency, selectively incorporating ambient sounds from the physical environment, with or without computer-generated audio. In some CGR environments, a person may sense and / or interact with only audio objects.

[0046] Examples of CGR include virtual reality and mixed reality.

[0047] Virtual Reality: A virtual reality (VR) environment refers to an emulated environment designed to be based entirely on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment through a simulation of the person's presence in the computer-generated environment and / or through a simulation of a subset of the person's physical movements in the computer-generated environment.

[0048] Mixed Reality: A mixed reality (MR) environment refers to a mimetic environment designed to incorporate sensory input from or representations of a physical environment in addition to including computer-generated sensory input (e.g., virtual objects), as opposed to a VR environment designed to be based entirely on computer-generated sensory input. On the virtual continuum, a mixed reality environment is anywhere between, but not including, a fully physical environment at one end and a virtual reality environment at the other. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Some electronic systems for presenting MR environments may also track location and / or orientation relative to the physical environment to allow virtual objects to interact with real objects (i.e., physical items from the physical environment or their representations). For example, the system may account for movement so that a virtual tree appears stationary relative to the physical ground.

[0049] Examples of mixed reality include augmented reality and augmented virtuality.

[0050] Augmented reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on a physical environment or a representation thereof. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on the transparent or translucent display, whereby a person using the system perceives the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system composites the images or videos with virtual objects and presents the composite on the opaque display. The person uses the system to indirectly view the physical environment through the images or videos of the physical environment and perceive the virtual objects superimposed on the physical environment. As used herein, video of a physical environment shown on an opaque display is referred to as "pass-through video," meaning that the system captures images of the physical environment using one or more image sensors and uses those images in presenting the AR environment on the opaque display. Alternatively, the system may include a projection system that projects virtual objects, e.g., as holograms, into the physical environment or onto a physical surface, such that a person using the system perceives the virtual objects superimposed on the physical environment. An augmented reality environment also refers to a mimicked environment in which a representation of the physical environment is transformed by computer-generated sensory information. For example, when providing pass-through video, the system may distort one or more sensor images to impose a selected perspective (e.g., viewpoint) different from the perspective captured by the imaging sensor. As another example, the representation of the physical environment may be distorted by graphically altering (e.g., magnifying) a portion thereof, thereby rendering the altered portion a non-photorealistic, altered version of the originally captured image. As a further example, the representation of the physical environment may be distorted by graphically removing or obscuring a portion thereof.

[0051] Augmented Virtual: An augmented virtual (AV) environment refers to a mimicking environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory inputs may be representations of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, while people with faces are realistically recreated from images taken of physical people. As another example, virtual objects may adopt the shape or color of physical items imaged by one or more imaging sensors. As a further example, virtual objects may adopt shadows that match the position of the sun in the physical environment.

[0052] Hardware: There are many different types of electronic systems that enable a person to sense and / or interact with various CGR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed over a person's eyes (e.g., similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speaker(s) and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system may have a transparent or translucent display rather than an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed to a person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser-scanned light source, or any combination of these technologies. The medium may be a light guide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. The projection-based system may employ retinal projection technology to project a graphical image onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or as physical surfaces.In some embodiments, controller 110 is configured to manage and coordinate the user's CGR experience. In some embodiments, controller 110 includes a suitable combination of software, firmware, and / or hardware. Controller 110 is described in more detail below with reference to FIG. 2. In some embodiments, controller 110 is a computing device that is local or remote to scene 105 (e.g., the physical environment). For example, controller 110 is a local server located within scene 105. In another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., an HMD, a display, a projector, a touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., BLUETOOTH, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, the controller 110 is contained within the housing (e.g., physical housing) of, or shares the same physical housing or support structure as, one or more of the display generation component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or one or more of the peripheral devices 195.

[0053] In some embodiments, display generation component 120 is configured to provide a CGR experience (e.g., at least a visual component of the CGR experience) to a user. In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Display generation component 120 is described in more detail below with reference to FIG. 3. In some embodiments, functionality of controller 110 is provided by and / or combined with display generation component 120.

[0054] According to some embodiments, the display generation component 120 provides a CGR experience to the user while the user is virtually and / or physically present within the scene 105.

[0055] In some embodiments, the display generation component is worn on a part of the user's body (e.g., on their head, their hand, etc.). Thus, display generation component 120 includes one or more CGR displays provided for displaying CGR content. For example, in various embodiments, display generation component 120 surrounds the user's field of view. In some embodiments, display generation component 120 is a handheld device (e.g., a smartphone or tablet) configured to present CGR content, where the user holds the device with a display pointed toward the user's field of view and a camera pointed toward scene 105. In some embodiments, the handheld device is optionally located within a housing worn on the user's head. In some embodiments, the handheld device is optionally located on a support (e.g., a tripod) in front of the user. In some embodiments, display generation component 120 is a CGR chamber, housing, or room configured to present CGR content without the user wearing or holding display generation component 120. Many user interfaces described with reference to one type of hardware for displaying CGR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying CGR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interactions with CGR content triggered based on interactions occurring in the space in front of a handheld or tripod-mounted device may be implemented similarly to an HMD in which the interactions occur in the space in front of the HMD and the CGR content responses are displayed via the HMD. Similarly, a user interface showing interactions with CGR content triggered based on movement of a handheld or tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hands)) may be implemented similarly to an HMD in which the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hands)).

[0056] While relevant features of operating environment 100 are shown in FIG. 1, those skilled in the art will understand from this disclosure that various other features have not been shown for the sake of brevity so as not to obscure more pertinent aspects of the exemplary embodiments disclosed herein.

[0057] 2 is a block diagram of an example controller 110, according to some embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more pertinent aspects of the embodiments disclosed herein. Thus, by way of non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a central processing unit (CPU), a processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), BLUETOOTH, ZIGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.

[0058] In some embodiments, one or more communication buses 204 include circuitry that interconnects and controls communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0059] Memory 220 includes high-speed random-access memory, such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220, or the non-transitory computer-readable storage medium of memory 220, stores the following programs, modules, and data structures, or a subset thereof, including optional operating system 230 and CGR experience module 240:

[0060] Operating system 230 includes instructions for handling various basic system services and for performing hardware-dependent tasks. In some embodiments, CGR experience module 240 is configured to manage and coordinate one or more CGR experiences for one or more users (e.g., a single CGR experience for one or more users, or multiple CGR experiences for respective groups of one or more users). To that end, in various embodiments, CGR experience module 240 includes a data acquisition unit 242, a tracking unit 244, an adjustment unit 246, and a data transmission unit 248.

[0061] 1 , and optionally from one or more of input devices 125, output devices 155, sensors 190, and / or peripheral devices 195. To that end, in various embodiments, data acquisition unit 242 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0062] In some embodiments, tracking unit 244 is configured to map scene 105 and track the position / location of at least display generation component 120 relative to scene 105 of FIG. 1 , and optionally the position / location of one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, tracking unit 244 includes instructions and / or logic therefor, as well as heuristics and metadata therefor. In some embodiments, tracking unit 244 includes hand tracking unit 243 and / or eye tracking unit 245. In some embodiments, hand tracking unit 243 is configured to track the position / location of one or more parts of a user's hand and / or the movement of one or more parts of a user's hand relative to scene 105 of FIG. 1 , relative to display generation component 120, and / or relative to a coordinate system defined relative to the user's hand. Hand tracking unit 243 is described in more detail below with respect to FIG. 4. In some embodiments, eye tracking unit 245 is configured to track the position and movement of a user's gaze (or, more broadly, the user's eyes, face, or head) relative to scene 105 (e.g., relative to the physical environment and / or the user (e.g., the user's hands)) or relative to CGR content displayed via display generation component 120. Eye tracking unit 245 is described in more detail below with respect to FIG. 5.

[0063] In some embodiments, adjustment unit 246 is configured to manage and adjust the CGR experience presented to the user by display generation component 120 and, optionally, by one or more of output devices 155 and / or peripheral devices 195. To that end, in various embodiments, adjustment unit 246 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0064] In some embodiments, data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least display generation component 120, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data transmission unit 248 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0065] Although the data acquisition unit 242, the tracking unit 244 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 are shown as being present on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 242, the tracking unit 244 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 may be located within separate computing devices.

[0066] Furthermore, Figure 2 is intended more to illustrate the functionality of various features that may be present in particular embodiments, as opposed to a structural overview of the embodiments described herein. As will be recognized by those skilled in the art, items shown separately may be combined and some items may be separated. For example, some functional modules shown separately in Figure 2 may be implemented within a single module, and various functions of a single functional block may be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of specific functions and how functions are allocated among them, will vary depending on implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0067] 3 is a block diagram of an example of a display generation component 120, according to some embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that, for the sake of brevity, various other features are not shown so as not to obscure more pertinent aspects of the embodiments disclosed herein. To that end, by way of non-limiting example, in some embodiments, the HMD 120 includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, BLUETOOTH, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more CGR displays 312, one or more optional inward-facing and / or outward-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.

[0068] In some embodiments, the one or more communication buses 304 include circuitry that interconnects and controls communication between the system components. In some embodiments, the one or more I / O devices and sensors 306 include at least one of an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.), etc.

[0069] In some embodiments, one or more CGR displays 312 are configured to provide a CGR experience to a user. In some embodiments, one or more CGR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emissive element display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, one or more CGR displays 312 correspond to a waveguide display, such as a diffractive, reflective, polarized, holographic, etc. For example, the HMD 120 includes a single CGR display. In another example, the HMD 120 includes a CGR display for each eye of the user. In some embodiments, one or more CGR displays 312 are capable of presenting MR or VR content. In some embodiments, one or more CGR displays 312 are capable of presenting MR or VR content.

[0070] In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as eye-tracking cameras). In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand(s) and optionally the user's arm(s) (and may be referred to as hand-tracking cameras). In some embodiments, the one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene viewed by the user when the HMD 120 is not present (and may be referred to as scene cameras). The one or more optional image sensors 314 may include one or more RGB cameras (e.g., with a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, one or more event-based cameras, and / or the like.

[0071] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320, or its non-transitory computer-readable storage medium, stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 330 and a CGR presentation module 340:

[0072] The operating system 330 includes instructions for handling various basic system services and for performing hardware-dependent tasks. In some embodiments, the CGR presentation module 340 is configured to present CGR content to a user via one or more CGR displays 312. To that end, in various embodiments, the CGR presentation module 340 includes a data acquisition unit 342, a CGR presentation unit 344, a CGR map generation unit 346, and a data transmission unit 348.

[0073] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 of Figure 1. To that end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0074] In some embodiments, the CGR presentation unit 344 is configured to present CGR content via one or more CGR displays 312. To that end, in various embodiments, the CGR presentation unit 344 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0075] In some embodiments, the CGR map generation unit 346 is configured to generate a CGR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed) based on the media content data. To that end, in various embodiments, the CGR map generation unit 346 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0076] In some embodiments, data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least controller 110, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data transmission unit 348 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0077] Although the data acquisition unit 342, the CGR presentation unit 344, the CGR map generation unit 346, and the data transmission unit 348 are shown as residing on a single device (e.g., the display generation component 120 of FIG. 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, the CGR presentation unit 344, the CGR map generation unit 346, and the data transmission unit 348 may be located within separate computing devices.

[0078] Furthermore, Figure 3 is intended more to illustrate the functionality of various features that may be present in particular implementations, as opposed to a structural overview of the embodiments described herein. As will be recognized by those skilled in the art, items shown separately can be combined and some items can be separated. For example, some functional modules shown separately in Figure 3 can be implemented within a single module, and various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of specific functions and how functions are allocated among them, will vary from implementation to implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0079] 4 is a schematic diagram of an example embodiment of a hand tracking device 140. In some embodiments, hand tracking device 140 (FIG. 1) is controlled by hand tracking unit 243 (FIG. 2) to track the position / location of one or more parts of a user's hand relative to scene 105 of FIG. 1 (e.g., relative to parts of the physical environment surrounding the user, relative to display generation component 120, or relative to parts of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system defined relative to the user's hand. In some embodiments, hand tracking device 140 is part of display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, hand tracking device 140 is separate from display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0080] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures hand images with sufficient resolution to allow for differentiation of the fingers and their respective positions. The image sensor 404 typically captures images of other parts of the user's body, or all of the body, and can have either zoom capabilities or a dedicated sensor with high magnification to capture hand images at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are processed as inputs to the controller 110.

[0081] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and possibly color image data) to the controller 110, which extracts high-level information from the map data. This high-level information is provided, typically via an application program interface (API), to an application running on the controller, which drives the display generation component 120 accordingly. For example, a user can interact with software running on the controller 110 by moving their hand 408 and changing the posture of their hand.

[0082] In some embodiments, the image sensor 404 projects a spot pattern onto a scene containing the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the pattern's spots. This approach is advantageous in that it does not require the user to hold or wear any type of beacon, sensor, or other marker. This provides depth coordinates of points in the scene relative to a predetermined reference plane at a specific distance from the image sensor 404. In this disclosure, the image sensor 404 is assumed to define an orthogonal set of x, y, and z axes such that the depth coordinate of a point in the scene corresponds to the z-component measured by the image sensor. Alternatively, the hand tracking device 440 can use other 3D mapping methods, such as stereoscopic imaging or time-of-flight measurements, based on single or multiple cameras or other types of sensors.

[0083] In some embodiments, the hand tracking device 140 captures and processes a time sequence of depth maps containing the user's hand while the user moves the hand (e.g., the entire hand or one or more fingers). Software running on the image sensor 404 and / or a processor in the controller 110 processes the 3D map data to extract patch descriptors of the hand in these depth maps. The software matches these descriptors with patch descriptors stored in the database 408, based on a previous learning process, to estimate the pose of the hand in each frame. The pose typically includes the 3D locations of the user's wrist joints and fingertips.

[0084] The software can also analyze hand and / or finger trajectories across multiple frames in a sequence to identify gestures. The pose estimation functionality described herein may be interleaved with motion tracking functionality, whereby patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to discover pose changes that occur across the remaining frames. The pose, motion, and gesture information is provided to an application program running on the controller 110 via the API described above. This program can, for example, move and modify an image presented on the display generation component 120 or perform other functions in response to the pose and / or gesture information.

[0085] In some embodiments, the software may be downloaded to the controller 110 in electronic form, for example, over a network, or alternatively may be provided on a tangible, non-transitory medium, such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively, or additionally, some or all of the described functionality of the computer may be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). While the controller 110 is shown in FIG. 4 as, by way of example, a separate unit from the image sensor 440, some or all of the processing functionality of the controller may be implemented by a suitable microprocessor and software, or by dedicated circuitry within the housing of the hand tracking device 402, or otherwise associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, handheld device, or head-mounted device) or using any other suitable computerized device, such as a game console or media player. The sensing function of the image sensor 404 may likewise be integrated into a computer or other computerized device that is controlled by the sensor output.

[0086] FIG. 4 also includes a schematic diagram of a depth map 410 captured by the image sensor 404, according to some embodiments. The depth map includes a matrix of pixels having respective depth values, as described above. A pixel 412 corresponding to the hand 406 is segmented from the background and wrist in this map. The intensity of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z-distance from the image sensor 404, with increasing gray levels corresponding to increasing depth. The controller 110 processes these depth values ​​to identify and segment components of the image (i.e., groups of adjacent pixels) that have characteristics of a human hand. These characteristics can include, for example, the overall size, shape, and frame-to-frame motion of the depth map sequence.

[0087] 4 also schematically illustrates a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406, according to some embodiments. In FIG. 4, the skeleton 414 is superimposed on a hand background 416 that was segmented from the original depth map. In some embodiments, key feature points on the hand (e.g., knuckles, fingertips, palm center, hand end corresponding to the wrist, etc.), and optionally the wrist or arm connected to the hand, are identified and positioned on the hand skeleton 414. In some embodiments, the location and movement of these key feature points over multiple image frames are used by the controller 110 to determine hand gestures performed by the hand or the current state of the hand, according to some embodiments.

[0088] FIG. 5 illustrates an exemplary embodiment of eye tracking device 130 (FIG. 1). In some embodiments, eye tracking device 130 is controlled by eye tracking unit 245 (FIG. 2) to track the position and movement of a user's gaze relative to scene 105 or relative to CGR content displayed via display generation component 120. In some embodiments, eye tracking device 130 is integrated with display generation component 120. For example, in some embodiments, if display generation component 120 is a head-mounted device, such as a headset, helmet, goggles, or glasses, or a handheld device disposed in a wearable frame, the head-mounted device includes both components for generating CGR content for viewing by the user and components for tracking the user's gaze relative to the CGR content. In some embodiments, eye tracking device 130 is separate from display generation component 120. For example, if the display generation component is a handheld device or a CGR chamber, eye tracking device 130 is optionally a device separate from the handheld device or the CGR chamber. In some embodiments, eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, head-mounted eye tracking device 130 is optionally used in conjunction with head-mounted or non-head-mounted display generation components. In some embodiments, eye tracking device 130 is not a head-mounted device, and is optionally used in conjunction with head-mounted display generation components. In some embodiments, eye tracking device 130 is not a head-mounted device, and is optionally part of non-head-mounted display generation components.

[0089] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that displays frames including left and right images in front of the user's eyes to provide a 3D virtual view to the user. For example, a head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display that allows the user to view the physical environment directly and display virtual objects on the transparent or translucent display. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects are projected, for example, onto a physical surface or as a hologram, allowing an individual using the system to observe the virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.

[0090] As shown in FIG. 5 , in some embodiments, the eye tracking device 130 includes at least one eye tracking camera (e.g., an infrared (IR) or near-IR (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera may be aimed at the user's eyes to receive reflected IR or NIR light from the light source directly from the eyes, or alternatively, may be aimed at a “hot” mirror positioned between the user's eyes and a display panel that reflects the IR or NIR light from the eyes to the eye tracking camera while allowing visual light to pass through. The eye tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the images to generate eye tracking information, and communicates the eye tracking information to the controller 110. In some embodiments, the user's eyes are tracked separately by their respective eye tracking cameras and illumination sources. In some embodiments, only one eye of the user is tracked by a separate eye-tracking camera and lighting source.

[0091] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the eye tracking device's parameters for the particular operating environment 100, such as the 3D geometric relationships and parameters of the LEDs, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at a factory or another facility before delivery of the AR / VR equipment to the end user. The device-specific calibration process may be an automatic or manual calibration process. The user-specific calibration process may include estimation of a particular user's eye parameters, such as pupil location, central visual location, optical axis, visual axis, eye spacing, etc. Once the device-specific and user-specific parameters for the eye tracking device 130 have been determined, images captured by the eye tracking camera can be processed using glint-assisted methods to determine the user's current visual axis and viewpoint relative to the display.

[0092] As shown in FIG. 5, eye tracking device 130 (e.g., 130A or 130B) includes an eyepiece(s) 520 and an eye tracking system including at least one eye tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking occurs and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye(s) 592. The eye tracking camera 540 may be positioned between the user's eye(s) 592 and the display 510 (e.g., the left or right display panel of a head-mounted display, or the display of a handheld device, a projector, etc.) and may be directed at a mirror 550 that reflects IR or NIR light from the eye(s) 592 while transmitting visible light (e.g., as shown at the top of FIG. 5), or may be directed at the user's eye(s) 592 to receive IR or NIR light reflected from the eye(s) 592 (e.g., as shown at the bottom of FIG. 5).

[0093] In some embodiments, controller 110 renders AR or VR frames 562 (e.g., left and right frames for left and right display panels) and provides frames 562 to display 510. Controller 110 uses gaze tracking input 542 from eye tracking camera 540 for various purposes, such as in processing frames 562 for display. Controller 110 optionally estimates the user's viewpoint on display 510 based on gaze tracking input 542 obtained from eye tracking camera 540, using a glint-assisted method or other suitable method. The viewpoint estimated from gaze tracking input 542 is optionally used to determine the user's current looking direction.

[0094] Some possible use cases of the user's current gaze direction are described below, but are not intended to be limiting. As an exemplary use case, the controller 110 can render virtual content differently based on the determined user's gaze direction. For example, the controller 110 may generate virtual content with higher resolution in a central visual area determined from the user's current gaze direction than in a peripheral area. As another example, the controller may position or move virtual content within a view based at least in part on the user's current gaze direction. As another example, the controller may display particular virtual content within a view based at least in part on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 can orient an external camera to capture the physical environment of the CGR experience and focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface within the environment the user is currently viewing on the display 510. As another exemplary use case, eyepiece 520 may be a focusable lens, and eye-tracking information is used by the controller to adjust the focus of eyepiece 520 so that the virtual object the user is currently looking at has the proper binocular coordination to match the convergence of the user's eyes 592. Controller 110 can utilize the eye-tracking information to orient and focus eyepiece 520 so that close objects the user is looking at appear at the correct distance.

[0095] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eyepieces (e.g., eyepiece(s) 520), an eye tracking camera (e.g., eye tracking camera(s) 540), and a light source (e.g., light source 530 (e.g., IR LED or NIR LED)) attached to the wearable housing. The light source emits light (e.g., IR or NIR light) toward the user's eye(s) 592. In some embodiments, the light sources may be arranged in a ring or circle around each lens, as shown in FIG. 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520, as an example. However, more or fewer light sources 530 may be used, and other arrangements and locations of the light sources 530 may be used.

[0096] In some embodiments, the display 510 emits light in the visible light range and not in the IR or NIR range, thereby not introducing noise into the eye-tracking system. Note that the location and angle of the eye-tracking camera(s) 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.

[0097] Embodiments of an eye tracking system such as that shown in FIG. 5 may be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a user with a computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experience.

[0098] FIG. 6A illustrates a glint-assisted eye tracking pipeline, according to some embodiments. In some embodiments, the eye tracking pipeline is implemented by a glint-assisted eye tracking system (e.g., eye tracking device 130 as shown in FIGS. 1 and 5). The glint-assisted eye tracking system can maintain a tracking state. Initially, the tracking state is off or "no." When in the tracking state, the glint-assisted eye tracking system tracks the pupil contour and glint in the current frame using prior information from the previous frame when analyzing the current frame. When not in the tracking state, the glint-assisted eye tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in the tracking state.

[0099] As shown in FIG. 6A, an eye-tracking camera can capture left and right images of a user's left and right eyes. The captured images are then input into an eye-tracking pipeline for processing beginning at 610. As indicated by the arrow returning to element 600, the eye-tracking system can continue to capture images of the user's eyes at a rate of, for example, 60-120 frames per second. In some embodiments, each set of captured images may be input into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.

[0100] At 610, if the tracking status is yes for the currently captured image, the method proceeds to element 640. If the tracking status is no at 610, the image is analyzed to detect the user's pupil and glint in the image, as shown at 620. If the pupil and glint are successfully detected at 630, the method proceeds to element 640. If not, the method returns to element 610 to process the next image of the user's eyes.

[0101] At 640, proceeding from element 410, the current frame is analyzed to track pupils and glints based in part on prior information from the previous frame. At 640, proceeding from element 630, a tracking state is initialized based on the detected pupils and glints in the current frame. The results of the processing at element 640 are checked to ensure that the tracking or detection results are reliable. For example, the results can be checked to determine whether a sufficient number of glints are successfully tracked or detected in the current frame to perform pupil and gaze estimation. At 650, if the results are not reliable, the tracking state is set to no and the method returns to element 610 to process the next image of the user's eyes. At 650, if the results are reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes) and pupil and glint information is passed to element 680 to estimate the user's gaze point.

[0102] 6A is intended to serve as an example of eye-tracking technology that may be used in particular implementations. As will be recognized by those skilled in the art, other eye-tracking technologies, now existing or developed in the future, may be used in place of or in combination with the glint-assisted eye-tracking technology described herein in computer system 101 to provide a user with a CGR experience according to various embodiments.

[0103] 6B illustrates an exemplary environment for electronic devices 101a and 101b for providing a CGR experience, according to some embodiments. In FIG. 6B, real-world environment 602 includes electronic devices 101a and 101b, users 608a and 608b, and a real-world object (e.g., table 604). As shown in FIG. 6B, electronic devices 101a and 101b are optionally tripod-mounted or otherwise secured within real-world environment 602 so that one or more hands of users 608a and 608b are free (e.g., users 608a and 608b are optionally not holding devices 101a and 101b with one or more hands). As described above, devices 101a and 101b optionally have one or more groups of sensors located on different sides of devices 101a and 101b, respectively. For example, devices 101a and 101b optionally include sensor groups 612-1a and 612-1b and sensor groups 612-2a and 612-2b (e.g., capable of capturing information from each side of devices 101a and 101b) located on the "back" and "front" sides of devices 101a and 101b, respectively. As used herein, the front side of device 101a is the side facing users 608a and 608b, and the back side of devices 101a and 101b is the side facing away from users 608a and 608b.

[0104] In some embodiments, sensor groups 612-2a and 612-2b include an eye tracking unit (e.g., eye tracking unit 245 described above with reference to FIG. 2) that includes one or more sensors for tracking the eyes and / or gaze of users such that the eye tracking unit can "see" users 608a and 608b and track the eye(s) of users 608a and 608b in the manner described above. In some embodiments, the eye tracking units of devices 101a and 101b can capture the movement, orientation, and / or gaze of the eyes of users 608a and 608b and treat the movement, orientation, and / or gaze as input.

[0105] In some embodiments, sensor groups 612-1a and 612-1b include hand tracking units (e.g., hand tracking units 243 described above with reference to FIG. 2) that can track one or more hands of users 608a and 608b held on the “back” sides of devices 101a and 101b, as shown in FIG. 6B. In some embodiments, hand tracking units are optionally included in sensor groups 612-2a and 612-2b so that users 608a and 608b can additionally or alternatively hold one or more hands on the “front” sides of devices 101a and 101b while devices 101a and 101b track the positions of the one or more hands. As described above, the hand tracking units of devices 101a and 101b can capture the movements, positions, and / or gestures of one or more hands of users 608a and 608b and treat the movements, positions, and / or gestures as inputs.

[0106] In some embodiments, sensor groups 612-1 a and 612-1 b optionally include one or more sensors (e.g., image sensor 404 described above with reference to FIG. 4 ) configured to capture images of real-world environment 602, including table 604. As described above, devices 101 a and 101 b can capture images of portions (e.g., part or all) of real-world environment 602 and present the captured portions of real-world environment 602 to the user via one or more display generation components of devices 101 a and 101 b (e.g., displays of devices 101 a and 101 b, optionally positioned on user-facing sides of devices 101 a and 101 b opposite sides facing the captured portions of real-world environment 602).

[0107] In some embodiments, the captured portion of the real-world environment 602 is used to provide the user with a CGR experience, e.g., a mixed reality environment in which one or more virtual objects are overlaid on a representation of the real-world environment 602.

[0108] Accordingly, the description herein describes several embodiments of three-dimensional environments (e.g., CGR environments) that include representations of real-world objects and representations of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table present in a physical environment that is captured and displayed within the three-dimensional environment (e.g., actively via a camera and display of the electronic device, or passively via a transparent or translucent display of the electronic device). As mentioned above, the three-dimensional environment is optionally a mixed reality system based on a physical environment, where the three-dimensional environment is captured by one or more sensors of the device and displayed via a display generation component. As a mixed reality system, the device can optionally display portions and / or objects of the physical environment such that each of the portions and / or objects of the physical environment appears to exist within the three-dimensional environment displayed by the electronic device. Similarly, the device can optionally display virtual objects in the three-dimensional environment such that the virtual objects appear to exist within the real world (e.g., the physical environment) by placing the virtual objects at respective locations in the three-dimensional environment that have corresponding locations in the real world. For example, the device optionally displays a vase in a manner that makes it appear as if the real vase were placed on a table in the physical environment. In some embodiments, each location in the three-dimensional environment has a corresponding location in the physical environment. Thus, when a device is described as displaying a virtual object at a location distinct from a physical object (e.g., at or near the location of a user's hand, or on or near a physical table, etc.), the device displays the virtual object at a particular location in the three-dimensional environment in a manner that makes it appear as if the virtual object were at or near the physical object in the physical world (e.g., the virtual object is displayed at a location in the three-dimensional environment that corresponds to the location in the physical environment where the virtual object would be displayed if the virtual object were a real object at that particular location).

[0109] In some embodiments, real-world objects present in the physical environment that are displayed in the three-dimensional environment can interact with virtual objects that exist only in the three-dimensional environment. For example, the three-dimensional environment can include a table and a vase placed on the table, where the table is a view (or representation) of the physical table in the physical environment and the vase is a virtual object.

[0110] Similarly, a user can optionally use one or more hands to interact with virtual objects in the three-dimensional environment as if the virtual objects were real objects in the physical environment. For example, as described above, one or more sensors of the device optionally capture one or more of the user's hands and display a representation of the user's hand in the three-dimensional environment (e.g., in a manner similar to displaying real-world objects in the three-dimensional environment described above), or in some embodiments, due to the transparency / translucency of the user interface, or the projection of the user interface onto a transparent / translucent surface, or the portion of the display generation component displaying the projection of the user interface to the user's eyes or field of view of the user's eyes, the user's hands are visible through the display generation component with the ability to see the physical environment through the user interface. Thus, in some embodiments, the user's hands are displayed at discrete locations in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment that can interact with virtual objects in the three-dimensional environment as if they were actual physical objects in the physical environment. In some embodiments, a user can move their hands to cause a representation of their hand in the three-dimensional environment to move in coordination with the movement of the user's hand.

[0111] In some of the embodiments described below, the device can optionally determine an “effective” distance between a physical object in the physical world and a virtual object in the three-dimensional environment, for example, to determine whether a physical object is interacting with a virtual object (e.g., whether a hand is touching, grabbing, holding, etc., a virtual object, or whether it is within a threshold distance from the virtual object). For example, when determining whether and / or how a user is interacting with a virtual object, the device determines the distance between the user's hand and the virtual object. In some embodiments, the device determines the distance between the user's hand and the virtual object by determining the distance between the location of the hand in the three-dimensional environment and the location of a target virtual object in the three-dimensional environment. For example, one or more of the user's hands are placed at specific positions in the physical world, which the device optionally captures and displays at specific corresponding positions in the three-dimensional environment (e.g., positions in the three-dimensional environment at which the hands are displayed, if the hands are virtual rather than physical hands). The positions of the hands in the three-dimensional environment are optionally compared to the positions of the target virtual objects in the three-dimensional environment to determine the distance between the user's one or more hands and the virtual object. In some embodiments, the device optionally determines the distance between a physical object and a virtual object by comparing positions in the physical world (e.g., as opposed to comparing positions in a three-dimensional environment). For example, when determining the distance between one or more of a user's hands and a virtual object, the device optionally determines the corresponding location in the physical world of the virtual object (e.g., the position at which the virtual object is located in the physical world if the virtual object is a physical object rather than a virtual object), and then determines the distance between the corresponding physical position and the user's one or more hands.In some embodiments, the same techniques are optionally used to determine the distance between any physical object and any virtual object. Thus, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the device optionally performs any of the techniques described above to map the location of the physical object to the three-dimensional environment and / or to map the location of the virtual object to the physical world.

[0112] In some embodiments, the same or similar techniques are used to determine where and what a user's gaze is directed at and / or where and what a physical stylus held by the user is directed at. For example, if a user's gaze is directed at a particular position in the physical environment, the device optionally determines a corresponding position in the three-dimensional environment, and if a virtual object is located at that corresponding virtual position, the device optionally determines that the user's gaze is directed at that virtual object. Similarly, the device can optionally determine where the physical stylus is pointing in the physical world based on the orientation of the physical stylus. In some embodiments, based on this determination, the device determines a corresponding virtual position in the three-dimensional environment that corresponds to the location in the physical world where the stylus is pointing, and optionally determines that the stylus is pointing to the corresponding virtual position in the three-dimensional environment.

[0113] Similarly, embodiments described herein may refer to the location of a user (e.g., a user of a device) and / or the location of a device within a three-dimensional environment. In some embodiments, a user of a device is holding, wearing, or otherwise located at or near the electronic device. Thus, in some embodiments, the location of the device is used as a proxy for the location of the user. In some embodiments, the location of the device and / or user within the physical environment corresponds to a distinct location within the three-dimensional environment. In some embodiments, the distinct location is a location from which a “camera” or “view” of the three-dimensional environment extends. For example, if a user stands at a location facing a distinct portion of the physical environment displayed by the display generation component, the location of the device is the location within the physical environment (and its corresponding location within the three-dimensional environment) at which the user would see objects within the physical environment in the same position, orientation, and / or size (e.g., absolutely and / or relative to each other) as the objects are displayed by the display generation component of the device. Similarly, if the virtual objects displayed in the three-dimensional environment were physical objects in the physical environment (e.g., the virtual objects are located in the same physical environment location and have the same physical environment size and orientation as in the three-dimensional environment), the location of the device and / or user is the position at which the user would see the virtual objects in the physical environment in the same position, orientation, and / or size (e.g., absolutely and / or relative to each other and to real-world objects) as displayed by the display generation component of the device.

[0114] In this disclosure, various input methods are described with respect to interaction with a computer system. Where one example is provided using one input device or input method and another example is provided using a different input device or input method, it should be understood that each example may be compatible with, and optionally utilize, the input device or input method described with respect to the other example. Similarly, various output methods are described with respect to interaction with a computer system. Where one example is provided using one output device or output method and another example is provided using a different output device or output method, it should be understood that each example may be compatible with, and optionally utilize, the output device or output method described with respect to the other example. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment via a computer system. Where an example is provided using interaction with a virtual environment and another example is provided using a mixed reality environment, it should be understood that each example may be compatible with, and optionally utilize, the method described with respect to the other example. Thus, the present disclosure discloses embodiments that are combinations of features of multiple examples, without exhaustively listing all features of the embodiments in the description of each exemplary embodiment.

[0115] Furthermore, for methods described herein in which one or more steps are conditioned on one or more conditions being satisfied, it should be understood that the described method can be repeated in multiple iterations, such that over the course of the iterations, all of the conditions on which the method steps are conditioned are satisfied in different iterations of the method. For example, if a method requires performing a first step if a condition is satisfied and a second step if the condition is not satisfied, one skilled in the art will understand that the steps recited in the claim are repeated in a particular order until the conditions are satisfied and then no longer satisfied. Thus, a method described with one or more steps that depend on one or more conditions being satisfied can be rewritten as a method that is repeated until each condition recited in the method is satisfied. However, this is not required for system or computer-readable medium claims in which the system or computer-readable medium includes instructions for performing a conditional action based on the satisfaction of the corresponding one or more conditions, and thus can determine whether a contingency is met without explicitly repeating the method steps until all conditions on which the method steps are conditioned are satisfied. Those skilled in the art will also understand that, as with methods having conditional steps, the system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all of the conditional steps have been performed. User Interface and Related Processes

[0116] We now turn our attention to embodiments of user interfaces (“UIs”) and associated processes that may be executed in a computer system, such as a portable multifunction device or a head-mounted device, equipped with a display generation component, one or more input devices, and (optionally) one or more cameras.

[0117] 7A-7C illustrate exemplary ways in which electronic device 101a or 101b performs or does not perform an action in response to a user input, depending on whether detecting a user ready state precedes the user input, according to some embodiments.

[0118] 7A illustrates electronic devices 101a and 101b displaying a three-dimensional environment via display generation components 120a and 120b. It should be understood that in some embodiments, electronic devices 101a and 101b utilize one or more of the techniques described with reference to FIGS. 7A-7C in a two-dimensional environment or user interface without departing from the scope of this disclosure. As described above with reference to FIGS. 1-6, electronic devices 101a and 101b optionally include display generation components 120a and 120b (e.g., touchscreens) and multiple image sensors 314a and 314b. The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor that electronic devices 101a and / or 101b can use to capture one or more images of a user or a portion of a user while the user is interacting with electronic devices 101a and / or 101b. In some embodiments, display generation components 120a and 120b are touchscreens capable of detecting the gestures and movements of a user's hands. In some embodiments, the user interface described below may also be implemented in a head-mounted display that includes display generation components that display the user interface to a user and sensors that detect the physical environment and / or the movements of the user's hands (e.g., external sensors facing outward from the user) and / or the user's line of sight (e.g., internal sensors facing inward toward the user's face).

[0119] 7A shows two electronic devices 101a and 101b displaying a three-dimensional environment including a representation 704 of a table (e.g., table 604 of FIG. 6B ) within the physical environment of electronic devices 101a and 101b, selectable options 707, and scrollable user interface elements 705. Electronic devices 101a and 101b are associated with different user viewpoints in the three-dimensional environment, and thus present the three-dimensional environment from different viewpoints in the three-dimensional environment. In some embodiments, representation 704 of the table is a photorealistic representation (e.g., a digital pass-through) displayed by display generation components 120a and / or 120b. In some embodiments, representation 704 of the table is a view of the table through a transparent portion of display generation components 120a and / or 120b (e.g., a physical pass-through). 7A, a user's gaze 701a of the first electronic device 101a is directed at a scrollable user interface element 705, which is within a user's zone of attention 703 of the first electronic device 101a. In some embodiments, the zone of attention 703 is similar to the zone of attention described in more detail below with reference to FIGS. 9A-10H.

[0120] In some embodiments, the first electronic device 101a displays objects in the three-dimensional environment that are not within the zone of attention 703 (e.g., representations of the table 704 and / or options 707) with a blurred and / or dimmed appearance (e.g., a de-emphasized appearance). In some embodiments, the second electronic device 101b blurs and / or dims (e.g., makes less prominent) portions of the three-dimensional environment based on a zone of attention of the user of the second electronic device 101b, which is optionally different from the zone of attention of the user of the first electronic device 101a. Thus, in some embodiments, the zone of attention and the blurring of objects outside the zone of attention are not synchronized between the electronic devices 101a and 101b. Rather, in some embodiments, the zones of attention associated with the electronic devices 101a and 101b are independent of one another.

[0121] 7A , the user's hand 709 of the first electronic device 101a is in an inactive hand state (e.g., hand state A). For example, the hand 709 is in a ready state or a hand shape that does not correspond to input, as described in more detail below. Because the hand 709 is in the inactive hand state, the first electronic device 101a displays the scrollable user interface element 705 without indicating that input is or will be directed at the scrollable user interface element 705. Similarly, the electronic device 101b also displays the scrollable user interface element 705 without indicating that input is or will be directed at the scrollable user interface element 705.

[0122] In some embodiments, the electronic device 101a displays an indication that the user's gaze 701a is on the user interface element 705 while the user's hand 709 is in an inactive state. For example, the electronic device 101a optionally changes the color, size, and / or position of the scrollable user interface element 705 in a manner different from how the electronic device 101a updates the scrollable user interface element 705 in response to detecting the user's ready state, as described below. In some embodiments, the electronic device 101a indicates the user's gaze 701a on the user interface element 705 by displaying a visual indication separate from updating the appearance of the scrollable user interface element 705. In some embodiments, the second electronic device 101b refrains from displaying an indication of the user's gaze of the first electronic device 101a. In some embodiments, the second electronic device 101b displays an indication to indicate the location of the user's gaze of the second electronic device 101b.

[0123] 7B, first electronic device 101a detects a user's ready state while user's gaze 701b is directed at scrollable user interface element 705. In some embodiments, the user's ready state is detected in response to detecting user's hand 709 in a direct ready hand state (e.g., hand state D). In some embodiments, the user's ready state is detected in response to detecting user's hand 711 in an indirect ready hand state (e.g., hand state B).

[0124] In some embodiments, the user's hand 709 of the first electronic device 101a is in a direct-ready state when the hand 709 is within a predetermined threshold distance (e.g., 0.5, 1, 2, 3, 4, 5, 10, 15, 20, 30 centimeters, etc.) of the scrollable user interface element 705, the scrollable user interface element 705 is within the user's attention zone 703, and / or the hand 709 is in a pointing hand shape (e.g., a hand shape with one or more fingers curled toward the palm and one or more fingers extended toward the scrollable user interface element 705). In some embodiments, the scrollable user interface element 705 does not need to be within the attention zone 703 for the direct-input readiness criteria to be met. In some embodiments, the user's gaze 701b does not need to be directed toward the scrollable user interface element 705 for the direct-input readiness criteria to be met.

[0125] In some embodiments, the user's hand 711 of the electronic device 101a is in an indirect ready state when the hand 711 is further than a predetermined threshold distance (e.g., 0.5, 1, 2, 3, 4, 5, 10, 15, 20, 30 centimeters, etc.) from the scrollable user interface element 705, the user's gaze 701b is directed toward the scrollable user interface element 705, and the hand 711 is in a pre-pinch hand shape (e.g., a hand shape with the thumb within a threshold distance (e.g., 0.1, 0.5, 1, 2, 3 centimeters, etc.) of the other fingers of the hand without touching the other fingers). In some embodiments, the indirect input ready state criteria is met when the scrollable user interface element 705 is within the user's attention zone 703, even if the gaze 701b is not directed toward the user interface element 705. In some embodiments, the electronic device 101a resolves ambiguity in determining the location of the user's gaze 701b, as described below with reference to FIGS. 11A-12F.

[0126] In some embodiments, a hand shape that satisfies a direct readiness criterion (e.g., by hand 709) is the same as a hand shape that satisfies an indirect readiness criterion (e.g., by hand 711). For example, both a pointing hand shape and a pre-pinch hand shape satisfy both the direct and indirect readiness criteria. In some embodiments, a hand shape that satisfies a direct readiness criterion (e.g., by hand 709) is different from a hand shape that satisfies an indirect readiness criterion (e.g., by hand 711). For example, a pointing hand shape is required for direct readiness, but a pre-pinch hand shape is required for indirect readiness.

[0127] In some embodiments, electronic device 101a (and / or 101b) communicates with one or more input devices, such as a stylus or trackpad. In some embodiments, the criteria for entering a ready state with an input device are different from the criteria for entering a ready state without one of these input devices. For example, the ready state criteria for these input devices do not require detecting the hand shape described above for the direct and indirect ready states without a stylus or trackpad. For example, the ready state criteria when a user is providing input to device 101a and / or 101b using a stylus require the user to be holding the stylus, and the ready state criteria when a user is providing input to device 101a and / or 101b using a trackpad require the user's hand to be resting on the trackpad.

[0128] In some embodiments, each hand of a user (e.g., left and right hands) has an independently associated ready state (e.g., each hand must independently satisfy its ready state criteria before devices 101a and / or 101b respond to input provided by each individual hand). In some embodiments, the ready state criteria for each hand are different from one another (e.g., a different hand shape is required for each hand, allowing only indirect or direct ready states for one or both hands). In some embodiments, the visual indication of the ready state for each hand is different. For example, when scrollable user interface element 705 changes color to indicate a ready state is detected by device 101a and / or 101b, the color of scrollable user interface element 705 may be a first color (e.g., blue) for the right-hand ready state and a second color (e.g., green) for the left-hand ready state.

[0129] In some embodiments, in response to detecting the user's ready state, electronic device 101a prepares to detect input provided by the user (e.g., by the user's hand(s)) and updates the display of scrollable user interface element 705 to indicate that further input is directed to scrollable user interface element 705. For example, as shown in FIG. 7B , scrollable user interface element 705 is updated on electronic device 101a by increasing the thickness of the line around the border of scrollable user interface element 705. In some embodiments, electronic device 101a updates the appearance of scrollable user interface element 705 in different or additional ways, such as by changing the color of the background of scrollable user interface element 705, displaying a highlight around scrollable user interface element 705, updating the size of scrollable user interface element 705, updating the position of scrollable user interface element 705 within the three-dimensional environment (e.g., displaying scrollable user interface element 705 closer to the user's point of view within the three-dimensional environment), etc. In some embodiments, the second electronic device 101b does not update the appearance of the scrollable user interface element 705 to indicate the readiness state of the user of the first electronic device 101a.

[0130] In some embodiments, the manner in which electronic device 101a updates scrollable user interface element 705 in response to detecting a ready state is the same regardless of whether the ready state is a direct ready state (e.g., by hand 709) or an indirect ready state (e.g., by hand 711). In some embodiments, the manner in which electronic device 101a updates scrollable user interface element 705 in response to detecting a ready state is different depending on whether the ready state is a direct ready state (e.g., by hand 709) or an indirect ready state (e.g., by hand 711). For example, if electronic device 101a updates the color of scrollable user interface element 705 in response to detecting a ready state, electronic device 101a uses a first color (e.g., blue) in response to a direct ready state (e.g., by hand 709) and a second color (e.g., green) in response to an indirect ready state (e.g., by hand 711).

[0131] In some embodiments, after detecting a ready state for the scrollable user interface element 705, the electronic device 101a updates the ready state target based on an indication of the user's focus. For example, the electronic device 101a directs an indirect ready state (e.g., by the hand 711) to the selectable option 707 (e.g., subsequently removes the ready state from the scrollable user interface element 705) in response to detecting that the location of the gaze 701b has moved from the scrollable user interface element 705 to the selectable option 707. As another example, the electronic device 101a directs a direct ready state (e.g., by the hand 709) to the selectable option 707 (e.g., subsequently removes the ready state from the scrollable user interface element 705) in response to detecting that the hand 709 has moved from within a threshold distance (e.g., 0.5, 1, 2, 3, 4, 5, 10, 15, 30 centimeters, etc.) of the scrollable user interface element 705 to within a threshold distance of the selectable option 707.

[0132] 7B , device 101b detects that a user of second electronic device 101b directs their gaze 701c toward selectable option 707 while the user's hand 715 is in an inactive state (e.g., hand state A). Because electronic device 101b does not detect the user's ready state, electronic device 101b refrains from updating selectable option 707 to indicate the user's ready state. In some embodiments, as described above, electronic device 101b updates the appearance of selectable option 707 to indicate that the user's gaze 701c is directed toward selectable option 707 in a manner different from how electronic device 101b updates user interface elements to indicate the ready state.

[0133] In some embodiments, the electronic devices 101a and 101b perform an action in response to an input only when a ready state is detected before the electronic devices 101a and 101b detect the input. FIG. 7C shows users of the electronic devices 101a and 101b providing input to the electronic devices 101a and 101b, respectively. In FIG. 7B, the first electronic device 101a detected the user's ready state, but the second electronic device 101b, as described above, did not detect the ready state. Thus, in FIG. 7C, the first electronic device 101a performs an action in response to detecting the user input, but the second electronic device 101b refrains from performing an action in response to detecting the user input.

[0134] 7C, first electronic device 101a detects scrolling input directed at scrollable user interface element 705. FIG. 7C illustrates direct scrolling input provided by hand 709 and / or indirect scrolling input provided by hand 711. Direct scrolling input includes detecting hand 709 within a direct input threshold (e.g., 0.05, 0.1, 0.2, 0.3, 0.5, 1 centimeter, etc.) or touching scrollable user interface element 705 while hand 709 is in a pointing hand shape (e.g., hand state E) while hand 709 is moving in a direction in which scrollable user interface element 705 is scrollable (e.g., vertical movement or horizontal movement). Indirect scrolling input includes detecting hand 711 further than a direct input ready state threshold (e.g., 0.5, 1, 2, 3, 4, 5, 10, 15, 30 centimeters, etc.) and / or further than a direct input threshold (e.g., 0.05, 0.1, 0.2, 0.3, 0.5, 1 centimeter, etc.) from scrollable user interface element 705, detecting hand 711 in a pinch hand shape (e.g., a hand shape with the thumb touching another finger on hand 711, hand state C) while detecting the user's gaze 701b on scrollable user interface element 705, and movement of hand 711 in a direction in which scrollable user interface element 705 is scrollable (e.g., vertical or horizontal movement).

[0135] In some embodiments, electronic device 101a requires that scrollable user interface element 705 be within a user's zone of attention 703 for scrolling input to be detected. In some embodiments, electronic device 101a does not require that scrollable user interface element 705 be within a user's zone of attention 703 for scrolling input to be detected. In some embodiments, electronic device 101a requires that a user's gaze 701b be directed at a scrollable user interface element 705 for scrolling input to be detected. In some embodiments, electronic device 101a does not require that a user's gaze 701b be directed at a scrollable user interface element 705 for scrolling input to be detected. In some embodiments, electronic device 101a requires that a user's gaze 701b be directed at a scrollable user interface element 705 for indirect scrolling input but not direct scrolling input.

[0136] In response to detecting the scroll input, the first electronic device 101a scrolls the content within the scrollable user interface element 705 according to the movement of the hand 709 or hand 711, as shown in FIG. 7C. In some embodiments, the first electronic device 101a sends an indication of the scrolling to the second electronic device 101b (e.g., via a server), and in response, the second electronic device 101b scrolls the scrollable user interface element 705 in the same manner as the first electronic device 101a scrolls the scrollable user interface element 705. For example, the scrollable user interface element 705 within the three-dimensional environment is now scrolled, and thus, electronic devices (including, for example, electronic devices other than the electronic device that detected the input to scroll the scrollable user interface element 705) that display perspectives of the three-dimensional environment that include the scrollable user interface element 705 reflect the scrolled state of the user interface element. In some embodiments, if the user readiness state shown in FIG. 7B is not detected before detecting the input shown in FIG. 7C, the electronic devices 101a and 101b refrain from scrolling the scrollable user interface element 705 in response to the input shown in FIG. 7C.

[0137] Thus, in some embodiments, the results of the user input are synchronized between the first electronic device 101a and the second electronic device 101b. For example, if the second electronic device 101b detects a selection of the selectable option 707, both the first electronic device 101a and the second electronic device 101b update the appearance (e.g., color, style, size, position, etc.) of the selectable option 707 while the selection input is detected and perform actions according to the selection.

[0138] 7B before detecting the input of FIG. 7C, electronic device 101a scrolls scrollable user interface 705 in response to the input. In some embodiments, electronic devices 101a and 101b forgo performing an action in response to the detected input without first detecting the ready state.

[0139] For example, in FIG. 7C , a user of second electronic device 101b provides an indirect selection input directed toward selectable option 707 with hand 715. In some embodiments, detecting the selection input includes detecting that the user's hand 715 is making a pinch gesture (e.g., hand state C) while the user's gaze 701c is directed toward selectable option 707. Because second electronic device 101b did not detect the ready state (e.g., of FIG. 7B ) before detecting the input of FIG. 7C , second electronic device 101b refrains from selecting option 707 and refrains from performing an action in accordance with the selection of option 707. In some embodiments, second electronic device 101b detects the same input (e.g., an indirect input) as first electronic device 101a in FIG. 7C , but because the ready state was not detected before the input was detected, second electronic device 101b does not perform an action in accordance with the input. In some embodiments, if the second electronic device 101b detects a direct input without first detecting a ready state, the second electronic device 101b still refrains from performing an action in response to the direct input because the ready state was not detected before the input was detected.

[0140] 8A-8K are flowcharts illustrating a method 800 for performing or not performing an action in response to a user input, depending on whether detecting a user's ready state precedes the user input, according to some embodiments. In some embodiments, method 800 is performed on a computer system (e.g., computer system 101 of FIG. 1 , such as a tablet, smartphone, wearable computer, or head-mounted device) that includes a display generation component (e.g., display generation component 120 of FIGS. 1, 3, and 4 ) (e.g., a head-up display, a display, a touchscreen, a projector, etc.) and one or more cameras (e.g., a camera pointing downward in a user's hand (e.g., a color sensor, an infrared sensor, and other depth-sensing camera) or a camera pointing forward from the user's head). In some embodiments, method 800 is governed by instructions stored on a non-transitory computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A ). Some operations of method 800 are optionally combined and / or the order of some operations is optionally changed.

[0141] In some embodiments, method 800 is performed on electronic device 101 a or 101 b, which is in communication with a display generation component and one or more input devices (e.g., a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer). In some embodiments, the display generation component is a display integrated with the electronic device (optionally a touchscreen display), an external display such as a monitor, projector, television, or a hardware component (optionally built-in or external) for projecting a user interface and making the user interface visible to one or more users, etc. In some embodiments, the one or more input devices include electronic devices or components capable of receiving user input (e.g., capturing user input, detecting user input, etc.) and transmitting information related to the user input to the electronic device. Examples of input devices include a touchscreen, a mouse (e.g., external), a trackpad (optionally integrated or external), a touchpad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the electronic device), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, an eye tracking device, and / or a motion sensor (e.g., hand tracking device, hand motion sensor), etc. In some embodiments, the electronic device is in communication with a hand tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touchscreen, trackpad). In some embodiments, the hand tracking device is a wearable device such as a smart glove. In some embodiments, the hand tracking device is a handheld input device such as a remote control or a stylus.

[0142] In some embodiments, such as in FIG. 7A , the electronic device 101a displays (802a) a user interface that includes a user interface element (e.g., 705) via a display generation component. In some embodiments, the user interface element is an interactive user interface element, and in response to detecting input directed at the user interface element, the electronic device performs an action associated with the user interface element. For example, the user interface element is a selectable option that, when selected, causes the electronic device to perform an action such as displaying a separate user interface, changing a setting on the electronic device, or starting playback of content. As another example, the user interface element is a container (e.g., a window) in which a user interface / content is displayed, and in response to detecting selection of the user interface element followed by movement input, the electronic device updates the position of the user interface element according to the movement input. In some embodiments, the user interface and / or user interface elements are displayed in (e.g., the user interface is and / or is displayed within) a three-dimensional environment that is generated, displayed, or otherwise made viewable by a device (e.g., a computer-generated reality (CGR) environment such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment).

[0143] In some embodiments, such as FIG. 7C , while displaying a user interface element (e.g., 705), the electronic device 101a detects (802b) an input from a predetermined portion (e.g., 709) of a user of the electronic device 101a (e.g., a hand, arm, head, eye, etc.) via one or more input devices. In some embodiments, detecting the input includes detecting, via a hand tracking device, that the user optionally performs a predetermined gesture with their hand while the user's gaze is directed at the user interface element. In some embodiments, the predetermined gesture is a pinch gesture that includes touching a thumb to another finger (e.g., index finger, middle finger, ring finger, pinky finger) of the same hand while looking at the user interface element. In some embodiments, the input is a direct or indirect interaction with the user interface element, such as described with reference to methods 1000, 1200, 1400, 1600, 1800, and / or 2000.

[0144] In some embodiments, in response to detecting an input from a predetermined portion of the user of the electronic device (802c), and in accordance with a determination that the pose (e.g., position, orientation, hand shape) of the predetermined portion of the user (e.g., 709) prior to detecting the input satisfies one or more criteria, the electronic device performs a distinct action (802d) according to the input from the predetermined portion of the user (e.g., 709) of the electronic device 101a, such as in FIG. 7C . In some embodiments, the pose of the user's physical characteristic is the orientation and / or shape of the user's hand. For example, the pose satisfies one or more criteria if the electronic device detects that the user's hand is oriented with the user's palm facing away from the user's torso while in a pre-pinch hand shape in which the user's thumb is within a threshold distance (e.g., 0.5, 1, 2 centimeters, etc.) of the other fingers (e.g., index finger, middle finger, ring finger, little finger) of the thumb. As another example, one or more criteria are met when the hand is in a pointing hand shape with one or more fingers extended and one or more other fingers curled toward the user's palm. Input by the user's hand following detection of the posture is optionally recognized as directed at a user interface element, and the device optionally performs a respective action according to the subsequent input by the hand. In some embodiments, the respective action includes scrolling a user interface, selecting an option, activating a setting, or navigating to a new user interface. In some embodiments, in response to detecting an input including a movement of a user's part following a selection after detecting a predetermined posture, the electronic device scrolls the user interface. For example, the electronic device first detects the user's gaze directed at the user interface while detecting a pointing hand shape, then detects a movement of the user's hand from the user's torso in a direction in which the user interface is scrollable, and scrolls the user interface in response to the sequence of inputs.As another example, in response to detecting a user's gaze toward an option for activating a setting on the electronic device while detecting a pre-pinch hand shape followed by a pinch hand shape, the electronic device activates a setting on the electronic device.

[0145] In some embodiments, such as in FIG. 7C , in response to detecting (802c) an input from a predetermined portion (e.g., 715) of the user of the electronic device 101b, and in response to determining that the posture of the predetermined portion (e.g., 715) of the user prior to detecting the input does not satisfy one or more criteria, such as in FIG. 7B , the electronic device 101b refrains from performing the individual action in accordance with the input from the predetermined portion (e.g., 715) of the user of the electronic device 101b, such as in FIG. 7C , in response to detecting that the user's gaze was not directed toward a user interface element while the posture and input were detected. In some embodiments, in response to determining that the user's gaze was directed toward a user interface element while the posture and input were detected, the electronic device performs the individual action in accordance with the input.

[0146] The above-described method of performing or not performing a first action depending on whether the posture of a predetermined part of the user before detecting the input meets one or more criteria provides an efficient way of reducing accidental user input, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, makes the user device interface more efficient, and further reduces power usage and improves the battery life of the electronic device by allowing the user to use the electronic device more quickly and efficiently while reducing usage errors, and further reduces the likelihood that the electronic device will perform an unintended action and then be canceled.

[0147] 7A , while a pose of a predefined portion of the user (e.g., 709) does not satisfy one or more criteria (e.g., before detecting input from the predefined portion of the user), the electronic device 101a displays a user interface element (e.g., 705) with visual characteristics (e.g., size, color, position, translucency) having a first value and displays a second user interface element (e.g., 707) included in the user interface with visual characteristics (e.g., size, color, position, translucency) having a second value (804a). In some embodiments, displaying the user interface element with visual characteristics having the first value and displaying the second user interface element with visual characteristics having the second value indicates that input focus is not directed toward either the user interface element or the second user interface element and / or that the electronic device will not direct input from the predefined portion of the user toward either the user interface element or the second user interface element.

[0148] 7B , while the posture of the predetermined portion of the user (e.g., 709) satisfies one or more criteria, the electronic device 101a updates (804b) visual characteristics of the user interface element (e.g., 705) to which the input focus is directed, which includes (e.g., before detecting input from the predetermined portion of the user), updating (804c) the user interface element (e.g., 705) to be displayed with visual characteristics (e.g., size, color, translucency) having a third value (e.g., different from the first value while maintaining the display of a second user interface element with visual characteristics having a second value). In some embodiments, the input focus is directed to the user interface element in accordance with a determination that the user's gaze is directed to the user interface element, optionally including a disambiguation technique according to method 1200. In some embodiments, the input focus is directed to the user interface element pursuant to a determination that the predetermined portion of the user is within a threshold distance (e.g., 0.5, 1, 2, 3, 4, 5, 10, 30, 50 centimeters, etc.) of the user interface element (e.g., a threshold distance for direct input). For example, before the predetermined portion of the user meets one or more criteria, the electronic device displays the user interface element in a first color, and in response to detecting that the predetermined portion of the user meets the one or more criteria and the input focus is directed to the user interface element, the electronic device displays the user interface element in a second color different from the first color to indicate that input from the predetermined portion of the user is directed to the user interface element.

[0149] In some embodiments, while the posture of the predefined portion of the user (e.g., 705) satisfies one or more criteria, such as in FIG. 7B , the electronic device 101a updates (804b) visual characteristics of the user interface element to which the input focus is directed (e.g., in the manner that the electronic device 101a updates user interface element 705 of FIG. 7B ), which includes, pursuant to a determination (e.g., before detecting input from the predefined portion of the user) that the input focus is directed to a second user interface element, the electronic device 101a updates (804d) the second user interface element to be displayed with visual characteristics having a fourth value (e.g., different from the second value, while maintaining the display of the user interface element with the visual characteristics having the first value) (e.g., updating the appearance of user interface element 707 of FIG. 7B if user interface element 707 has the input focus instead of user interface element 705 having the input focus as in FIG. 7B ). In some embodiments, the input focus is directed to the second user interface element in accordance with a determination that the user's gaze is directed to the second user interface element, optionally including a disambiguation technique according to method 1200. In some embodiments, the input focus is directed to the second user interface element in accordance with a determination that a predetermined portion of the user is within a threshold distance (e.g., 0.5, 1, 2, 3, 4, 5, 10, 50 centimeters, etc.) of the second user interface element (e.g., a threshold distance for direct input). For example, before the predetermined portion of the user meets one or more criteria, the electronic device displays the second user interface element in a first color, and in response to detecting that the predetermined portion of the user meets the one or more criteria and the input focus is directed to the second user interface element, the electronic device displays the second user interface element in a second color different from the first color to indicate that input is directed to the user interface element.

[0150] The above-described method of updating the visual characteristics of a user interface element to which input focus is directed in response to detecting that a predetermined portion of the user satisfies one or more criteria provides an efficient way of indicating to a user which user interface element input is directed, which simplifies the interaction between the user and the electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving battery life of the electronic device while reducing errors in use.

[0151] 7B , the input focus is directed 806a to a user interface element (e.g., 705) in accordance with a determination that a predetermined portion of the user (e.g., 709) is within a threshold distance (e.g., 0.5, 1, 2, 3, 4, 5, 10, 50 centimeters, etc.) of a location corresponding to the user interface element (e.g., 705) (e.g., not within a threshold distance of a second user interface element). In some embodiments, the threshold distance is associated with direct input, as described with reference to methods 800, 1000, 1200, 1400, 1600, 1800, and / or 2000. For example, the input focus is directed to a user interface element in response to detecting a finger of the user's hand in the shape of a pointing hand within a threshold distance of the user interface element.

[0152] In some embodiments, the input focus is directed (806b) to the second user interface element (e.g., 707) of FIG. 7B pursuant to a determination that a predetermined portion of the user (e.g., 709) is within a threshold distance (e.g., 0.5, 1, 2, 3, 4, 5, 10, 50 centimeters, etc.) of the second user interface element (e.g., not within the threshold distance of the user interface element, such as if the user's hand 709 were within the threshold distance of user interface element 707 instead of user interface element 705 of FIG. 7B ). In some embodiments, the threshold distance is associated with direct input, as described with reference to methods 800, 1000, 1200, 1400, 1600, 1800, and / or 2000. For example, the input focus is directed to the second user interface element in response to detecting the fingers of the user's hand in the shape of a pointing hand within the threshold distance of the second user interface element.

[0153] The above-described method of directing input focus based on which predetermined part of the user is within a threshold distance of which user interface element provides an efficient way of directing user input when using a predetermined part of the user to provide input, which simplifies the interaction between the user and the electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving the battery life of the electronic device while reducing errors in use.

[0154] In some embodiments, such as in FIG. 7B , the input focus is directed (808a) to a user interface element (e.g., 705) in accordance with a determination that a user's gaze (e.g., 701b) is directed toward the user interface element (e.g., 705) (e.g., a predetermined portion of the user is not within a threshold distance of the user interface element and / or any interactive user interface elements). In some embodiments, determining that the user's gaze is directed toward a user interface element includes one or more disambiguation techniques according to method 1200. For example, the electronic device directs the input focus to a user interface element for indirect input in response to detecting the user's gaze directed toward the user interface element.

[0155] In some embodiments, the input focus is directed (808b) to the second user interface element (e.g., 707) of FIG. 7B in accordance with a determination that the user's gaze is directed toward the second user interface element (e.g., 707) (e.g., a predetermined portion of the user is not within a threshold distance of the second user interface element and / or any interactable user interface elements). For example, if the user's gaze is directed toward user interface element 707 of FIG. 7B rather than toward user interface element 705, the input focus is directed toward user interface element 707. In some embodiments, determining that the user's gaze is directed toward the second user interface element includes one or more disambiguation techniques according to method 1200. For example, the electronic device directs the input focus to the second user interface element for indirect input in response to detecting the user's gaze directed toward the second user interface element.

[0156] The above-described method of directing input focus to a user interface at which a user is looking provides an efficient way of directing user input without the user of additional input devices (e.g., other than eye-tracking and hand-tracking devices), which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving the battery life of the electronic device while reducing errors in use.

[0157] In some embodiments, such as in FIG. 7B , updating the visual characteristics of the user interface element (e.g., 705) at which the input focus is directed includes updating (810b) the visual characteristics of the user interface element (e.g., 705) at which the input focus is directed in accordance with a determination that the pose of the predefined portion of the user (e.g., 709) satisfies a first set of one or more criteria, such as in FIG. 7B , in accordance with a determination that the pose of the predefined portion of the user (e.g., 709) is less than a threshold distance (e.g., 1, 2, 3, 4, 5, 10, 15, 30 centimeters, etc.) from a location corresponding to the user interface element (e.g., 705) (and, optionally, the visual characteristics of the user interface element at which the input focus is directed are not updated in accordance with a determination that the pose of the predefined portion of the user does not satisfy the first set of one or more criteria) (e.g., associated with direct input as described with reference to methods 800, 1000, 1200, 1400, 1600, 1800 and / or 2000). For example, a first set of one or more criteria may include detecting a pointing hand shape (e.g., a shape in which fingers extend outward from an otherwise closed hand) while the user's hand is within a direct input threshold distance of a user interface element.

[0158] In some embodiments, such as in FIG. 7B , updating the visual characteristics of the user interface element (e.g., 705) at which the input focus is directed includes updating (810c) the visual characteristics of the user interface element (e.g., 705) at which the input focus is directed in accordance with a determination that the pose of the pre-defined portion of the user (e.g., 711) exceeds a threshold distance (e.g., 1, 2, 3, 4, 5, 10, 15, 30 centimeters, etc.) from a location corresponding to the user interface element (e.g., 705) and in accordance with a determination that the pose of the pre-defined portion of the user (e.g., 711) satisfies a second set of one or more criteria (e.g., associated with indirect input as described with reference to methods 800, 1000, 1200, 1400, 1600, 1800, and / or 2000) that differ from the first set of one or more criteria, such as in FIG. 7B (and, optionally, the visual characteristics of the user interface element at which the input focus is directed are not updated in accordance with a determination that the pose of the pre-defined portion of the user does not satisfy the second set of one or more criteria). For example, the second set of one or more criteria includes detecting a pre-pinch hand shape instead of detecting a pointing hand shape while the user's hand exceeds a direct input threshold from a user interface element. In some embodiments, a hand shape that satisfies the one or more first criteria is different from a hand shape that satisfies the one or more second criteria. In some embodiments, the one or more criteria are not met if a predetermined portion of the user exceeds a threshold distance from a location corresponding to a user interface element and a posture of the predetermined portion of the user satisfies the first set of one or more criteria without satisfying the second set of one or more criteria. In some embodiments, the one or more criteria are not met if a predetermined portion of the user is less than a threshold distance from a location corresponding to a user interface element and a posture of the predetermined portion of the user satisfies the second set of one or more criteria without satisfying the first set of one or more criteria.

[0159] The above-described method of using different criteria to evaluate a predetermined portion of a user depending on whether the predetermined portion of the user is within a threshold distance of a location corresponding to a user interface element provides an efficient and intuitive way of interacting with user interface elements that is tailored to whether the input is direct or indirect, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving the battery life of the electronic device while reducing errors in use.

[0160] 7B , poses of the predefined portion of the user (e.g., 709) that satisfy one or more criteria include poses of the predefined portion of the user (e.g., 709) that satisfy a first set of one or more criteria (812b) pursuant to a determination that the predefined portion of the user (e.g., 709) is less than a threshold distance (e.g., 1, 2, 3, 4, 5, 10, 15, 30 centimeters, etc.) from a location corresponding to the user interface element (e.g., 705) (e.g., associated with direct input as described with reference to methods 800, 1000, 1200, 1400, 1600, 1800, and / or 2000). For example, the first set of one or more criteria includes detecting a pointing hand shape (e.g., fingers extending outward from an otherwise closed hand) while the user's hand is within a direct input threshold distance of the user interface element.

[0161] In some embodiments, such as FIG. 7B , a pose of the predefined portion of the user (e.g., 711) that satisfies one or more criteria includes satisfying (812a) a second set of one or more criteria (e.g., associated with indirect input as described with reference to methods 800, 1000, 1200, 1400, 1600, 1800, and / or 2000) that differs from (812c) a first set of one or more criteria pursuant to a determination that the predefined portion of the user (e.g., 711) exceeds a threshold distance (e.g., 1, 2, 3, 4, 5, 10, 15, 30 centimeters, etc.) from a location corresponding to the user interface element (e.g., 705). For example, the second set of one or more criteria includes detecting a pre-pinch hand shape while the user's hand exceeds a direct input threshold from the user interface element. In some embodiments, the hand shape that satisfies the one or more first criteria differs from the hand shape that satisfies the one or more second criteria. In some embodiments, the one or more criteria are not met if the predetermined portion of the user is more than a threshold distance from a location corresponding to the user interface element and a posture of the predetermined portion of the user satisfies a first set of one or more criteria without satisfying a second set of one or more criteria. In some embodiments, the one or more criteria are not met if the predetermined portion of the user is less than a threshold distance from a location corresponding to the user interface element and a posture of the predetermined portion of the user satisfies a second set of one or more criteria without satisfying the first set of one or more criteria.

[0162] The above-described method of using different criteria to evaluate a predetermined portion of a user depending on whether the predetermined portion of the user is within a threshold distance of a location corresponding to a user interface element provides an efficient and intuitive way of interacting with user interface elements that is tailored to whether the input is direct or indirect, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving the battery life of the electronic device while reducing errors in use.

[0163] In some embodiments, a pose of a predefined portion of a user that satisfies one or more criteria, such as in FIG. 7B , includes a pose of a predefined portion of a user that satisfies a first set of one or more criteria (814b) pursuant to a determination that the predefined portion of a user is holding (e.g., interacting with, or touching) an input device (e.g., a stylus, a remote control, a trackpad) of one or more input devices (e.g., if the user's hand 709 in FIG. 7B were holding an input device). In some embodiments, the predefined portion of a user is the user's hand. In some embodiments, the first set of one or more criteria is met when the user holds a stylus or controller in a hand within a predefined region of the three-dimensional environment and / or at a predefined orientation relative to a user interface element and / or relative to the user's torso. In some embodiments, a first set of one or more criteria is met when a user holds the remote control within a predetermined area of ​​the three-dimensional environment, at a predetermined orientation relative to the user interface elements and / or relative to the user's torso, and / or while the user's thumb and fingers are resting on individual components of the remote control (e.g., buttons, trackpad, touchpad, etc.). In some embodiments, the first set of one or more criteria is met when a user is holding or interacting with the trackpad and a predetermined part of the user is in contact with the touch-sensitive surface of the trackpad (e.g., without pressing the trackpad as is done to make a selection).

[0164] In some embodiments, such as in FIG. 7B , the pose of the predefined portion of the user (e.g., 709) that satisfies the one or more criteria includes a pose of the predefined portion of the user (e.g., 709) that satisfies a second set of one or more criteria (e.g., different from the first set of one or more criteria) (814a) pursuant to a determination that the predefined portion of the user (e.g., 709) is not holding an input device. In some embodiments, the second set of one or more criteria is satisfied when the user of the electronic device is not holding, touching, or interacting with the input device, and the user's pose is a predefined pose (e.g., a pose including a pinch or pointing hand shape) as described above instead of holding a stylus or controller in the hand. In some embodiments, when the predefined portion of the user is holding an input device and the second set of one or more criteria is satisfied and the first set of one or more criteria is not satisfied, the pose of the predefined portion of the user does not satisfy the one or more criteria. In some embodiments, when the predetermined part of the user is not holding the input device, a first set of one or more criteria is met, and a second set of one or more criteria is not met, the posture of the predetermined part of the user does not satisfy one or more criteria.

[0165] The above-described method of evaluating a predetermined portion of a user according to different criteria depending on whether the user is holding an input device provides an efficient way of switching between accepting input using an input device (e.g., an input device other than an eye tracking device and / or a hand tracking device) and accepting input without an input device, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further reduces power usage and improves the battery life of the electronic device by allowing the user to use the electronic device more quickly and efficiently.

[0166] 7B , poses of the predefined portion of the user (e.g., 709) that satisfy one or more criteria include poses of the predefined portion of the user (e.g., 709) that satisfy a first set of one or more criteria (816b) pursuant to a determination that the predefined portion of the user (e.g., 709) is less than a threshold distance (e.g., 0.5, 1, 2, 3, 4, 5, 10, 15, 30, 50 centimeters, etc., corresponding to direct input) from a location corresponding to a user interface element (e.g., 705). For example, the first set of one or more criteria includes detecting a pointing hand shape and / or a pre-pinch hand shape while the user's hand is within the direct input threshold distance of the user interface element.

[0167] In some embodiments, such as FIG. 7B , the poses of the predefined portion of the user (e.g., 711) that satisfy the one or more criteria include a pose of the predefined portion of the user (e.g., 711) that satisfies a first set of one or more criteria (816c) pursuant to a determination that the predefined portion of the user (e.g., 711) exceeds a threshold distance (e.g., 0.5, 1, 2, 3, 4, 5, 10, 15, 30, 50 centimeters, etc., corresponding to indirect input) from a location corresponding to a user interface element (e.g., 705). For example, while the user's hand exceeds the direct input threshold from the user interface element, the second set of one or more criteria includes detecting a pre-pinch hand shape and / or a pointing hand shape that is the same as the hand shape used to satisfy the one or more criteria. In some embodiments, the hand shape that satisfies the one or more first criteria is the same regardless of whether the predefined portion of the hand is above or below the threshold distance from a location corresponding to the user interface element.

[0168] The above-described method of evaluating the posture of a predetermined part of a user against a first set of one or more criteria, regardless of the distance between the predetermined part of the user and a location corresponding to a user interface element, provides an efficient and consistent method of detecting user input provided by a predetermined part of the user, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving the battery life of the electronic device while reducing errors in use.

[0169] In some embodiments, such as in FIG. 7C , pursuant to a determination that a predetermined portion of the user in the individual input (e.g., 711) is more than a threshold distance (e.g., 0.5, 1, 2, 3, 4, 5, 10, 15, 30, 50 centimeters, etc., corresponding to indirect input) from a location corresponding to the user interface element (e.g., 705) (e.g., the input is indirect input), the one or more criteria include a criterion that is met when the user's attention is directed to the user interface element (e.g., 705) (818a) (e.g., the criterion is not met if the user's attention is not directed to the user interface element) (e.g., the user's gaze is within a threshold distance of the user interface element and the user interface element is within the user's attention zone, as described with reference to method 1000). In some embodiments, the electronic device determines which user interface element the indirect input is directed to based on the user's attention, such that it is not possible to provide indirect input to an individual user interface element without directing the user's attention to the individual user interface element.

[0170] 7C , pursuant to a determination that a predetermined portion of the user (e.g., 709) in a particular input is less than a threshold distance (e.g., 0.5, 1, 2, 3, 4, 5, 10, 15, 30, 50 centimeters, etc., corresponding to direct input) from a location corresponding to a user interface element (e.g., 705) (e.g., the input is direct input), the one or more criteria do not include a requirement that the user's attention be directed to the user interface element (e.g., 709) for the one or more criteria to be met (818b) (e.g., one or more criteria can be met without the user's attention being directed to the user interface element). In some embodiments, the electronic device determines a target for the direct input based on the location of the predetermined portion of the user relative to a user interface element in the user interface, and directs the input to the user interface element that is closest to the predetermined portion of the user, regardless of whether the user's attention is directed to that user interface element.

[0171] The above-described method of requiring user attention to satisfy one or more criteria while a predetermined portion of the user is beyond a threshold distance from the user interface element and not requiring user attention to satisfy one or more criteria while a predetermined portion of the user is less than the threshold distance from the user interface element provides an efficient method of allowing a user to view other portions of the user interface element while providing direct input, thus saving user time while using the electronic device and reducing user error while providing indirect input, which simplifies the interaction between the user and the electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further reduces power usage and improves battery life of the electronic device by allowing the user to use the electronic device more quickly and efficiently.

[0172] In some embodiments, in response to detecting that a user's gaze (e.g., 701a) is directed toward a first region of the user interface (e.g., 703), the electronic device 101a, via the display generation component, visually obscures (e.g., blurs, dims, darkens, and / or desaturates) (820a) a second region of the user interface relative to the first region of the user interface (e.g., 705), such as in FIG. 7A . In some embodiments, the electronic device alters the display of the second region of the user interface and / or alters the display of the first region of the user interface to achieve the visual obscuration of the second region of the user interface relative to the first region of the user interface.

[0173] 7B , in response to detecting that the user's gaze 701c is directed toward a second region of the user interface (e.g., 702), the electronic device 101b, via the display generation component, visually obscures (e.g., blurs, dims, darkens, and / or desaturates) (820b) the first region of the user interface relative to the second region of the user interface (e.g., 702). In some embodiments, the electronic device alters the display of the first region of the user interface and / or alters the display of the second region of the user interface to achieve the visual obscuration of the first region of the user interface relative to the second region of the user interface. In some embodiments, the first and / or second region of the user interface include one or more virtual objects (e.g., application user interfaces, items of content, representations of other users, files, control elements, etc.) and / or one or more physical objects (e.g., pass-through video including photorealistic representations of real objects, true pass-through where a view of the real objects is visible through transparent portions of the display generation component) that are obscured when the region of the user interface is obscured.

[0174] The above-described method of visually obscuring areas other than the area where the user's gaze is directed provides an efficient way of reducing visual clutter while the user is looking at individual areas of the user interface, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which in turn allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving the battery life of the electronic device.

[0175] In some embodiments, such as in FIG. 7A , the user interface is accessible 822a by the electronic device 101a and the second electronic device 101b (e.g., the electronic device and the second electronic device are in communication (e.g., via a wired or wireless network connection)). In some embodiments, the electronic device and the second electronic device are located remotely from one another. In some embodiments, the electronic device and the second electronic device are co-located (e.g., in the same room, building, etc.). In some embodiments, the electronic device and the second electronic device present the three-dimensional environment in a co-presence session, where representations of users of both devices are associated with unique locations within the three-dimensional environment and where each electronic device displays the three-dimensional environment from the perspective of a separate user's representation.

[0176] 7B , in response to an indication that the gaze 701c of a second user of the second electronic device 101b is directed toward the first region 702 of the user interface, the electronic device 101a, via the display generation component, refrains from visually obscuring (e.g., blurring, dimming, darkening, and / or desaturating) the second region of the user interface relative to the first region of the user interface (822b). In some embodiments, the second electronic device refrains from visually obscuring the second region of the user interface in response to a determination that the gaze of the second user is directed toward the first region of the user interface. In some embodiments, in response to a determination that the gaze of a user of the electronic device is directed toward the first region of the user interface, the second electronic device refrains from visually obscuring the second region of the user interface relative to the first region of the user interface.

[0177] 7B , in response to an indication that the gaze of a second user of the second electronic device 101a is directed toward a second region of the user interface (e.g., 703), the electronic device 101b, via the display generation component, refrains from visually obscuring (e.g., blurring, dimming, darkening, and / or desaturating) the first region of the user interface relative to the second region of the user interface (822c). In some embodiments, the second electronic device refrains from visually obscuring the first region of the user interface in response to a determination that the gaze of a user of the electronic device is directed toward the second region of the user interface. In some embodiments, in response to a determination that the gaze of a user of the electronic device is directed toward the second region of the user interface, the second electronic device refrains from visually obscuring the first region of the user interface relative to the second region of the user interface.

[0178] The above-described method of forgoing visually obscuring areas of a user interface based on the line of sight of a user of a second electronic device provides an efficient way of allowing a user to view different areas of the user interface simultaneously, which simplifies the interaction between the user and the electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which in turn reduces power usage and improves battery life of the electronic device by allowing the user to use the electronic device more quickly and efficiently.

[0179] In some embodiments, such as FIG. 7C , detecting input from a predetermined portion (e.g., 705) of the user of electronic device 101a includes detecting a pinch (e.g., pinch, pinch and hold, pinch and drag, double pinch, pluck, release without velocity, toss with velocity) gesture performed by a predetermined portion (e.g., 709) of the user via a hand tracking device (824a). In some embodiments, detecting a pinch gesture includes detecting the user moving the thumb toward and / or within a predetermined distance of another finger of the thumb hand (e.g., index finger, middle finger, ring finger, pinky finger). In some embodiments, detecting a posture that meets one or more criteria includes detecting that the user is in a ready state, such as a pre-pinch hand shape in which the thumb is within a threshold distance (e.g., 1, 2, 3, 4, 5 centimeters, etc.) of the other fingers.

[0180] The above-described method of detecting inputs including pinch gestures provides an efficient way of accepting user inputs based on hand gestures without requiring the user to physically touch and / or manipulate the input device with their hands, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further reduces power usage and improves the battery life of the electronic device by allowing the user to use the electronic device more quickly and efficiently.

[0181] In some embodiments, such as FIG. 7C , detecting input from a predefined portion (e.g., 709) of the user of electronic device 101a includes detecting a press (e.g., tap, press and hold, press and drag, flick) gesture performed by the predefined portion (e.g., 709) of the user via a hand tracking device (826a). In some embodiments, detecting the press gesture includes detecting a predefined portion of the user pressing a location corresponding to a user interface element (e.g., such as those described with reference to methods 1400, 1600, and / or 2000) displayed in a user interface, such as a user interface element according to method 1800 or a virtual trackpad or other visual indication. In some embodiments, before detecting an input including a press gesture, the electronic device detects a posture of a predefined portion of the user that meets one or more criteria, including detecting a user in a ready state, such as a pointing hand with one or more fingers extended and one or more fingers curled toward the palm. In some embodiments, the press gesture includes moving the user's fingers, hand, or arm while the hand is in the pointing hand configuration.

[0182] The above-described method of detecting input, including press gestures, provides an efficient way of accepting user input based on hand gestures without requiring the user to physically touch and / or manipulate the input device with their hands, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further reduces power usage and improves the battery life of the electronic device by allowing the user to use the electronic device more quickly and efficiently.

[0183] In some embodiments, such as in FIG. 7C , detecting input from a user's predetermined portion (e.g., 709) of electronic device 101a includes detecting lateral movement of the user's predetermined portion (e.g., 709) relative to a location corresponding to a user interface element (e.g., 705) (e.g., as described with reference to method 1800) (828a). In some embodiments, lateral movement includes movement that includes a component perpendicular to a straight-line path between the user's predetermined portion and the location corresponding to the user interface element. For example, if a user interface element is in front of the user's predetermined portion and the user moves the user's predetermined portion left, right, up, or down, the movement is lateral movement. For example, the input is one of a press-and-drag, a pinch-and-drag, or a toss (with velocity) input.

[0184] The above-described method of detecting inputs involving lateral movement of a predetermined portion of a user relative to a user interface element provides an efficient way of providing directional inputs to an electronic device by a predetermined portion of the user, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further reduces power usage and improves the battery life of the electronic device by allowing the user to use the electronic device more quickly and efficiently.

[0185] In some embodiments, such as FIG. 7A , before determining (830a) that the posture of a predetermined portion (e.g., 709) of the user prior to detecting the input satisfies one or more criteria, the electronic device 101a detects (830b) via an eye tracking device that the user's gaze (e.g., 701a) is directed toward a user interface element (e.g., 705) (e.g., according to one or more disambiguation techniques of method 1200).

[0186] In some embodiments, prior to determining (830a) that a posture of a predetermined portion (e.g., 709) of the user prior to detecting the input satisfies one or more criteria, in response to detecting that the user's gaze (e.g., 701a) is directed toward a user interface element (e.g., 705), such as in FIG. 7A , the electronic device 101a displays (830c), via the display generation component, a first indication that the user's gaze (e.g., 701a) is directed toward the user interface element (e.g., 705). In some embodiments, the first indication is a highlight overlaid on or displayed around the user interface element. In some embodiments, the first indication is a change in color or location (e.g., toward the user) of the user interface element. In some embodiments, the first indication is a symbol or icon overlaid on or displayed near the user interface element.

[0187] The above-described method of displaying a first indication that a user's gaze is directed at a user interface element provides an efficient way of communicating to a user that input focus is based on the location where the user is looking, which simplifies the interaction between the user and the electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further reduces power usage and improves battery life of the electronic device by allowing the user to use the electronic device more quickly and efficiently.

[0188] In some embodiments, such as in FIG. 7B , prior to detecting an input from a predetermined portion (e.g., 709) of the user of electronic device 101a, while a posture of the predetermined portion (e.g., 709) of the user prior to detecting the input satisfies one or more criteria (832a) (e.g., while the user's gaze is directed toward a user interface element (e.g., according to one or more disambiguation techniques of method 1200)), electronic device 101a, via the display generation component, displays a second indication (832b) that the posture of the predetermined portion (e.g., 709) of the user prior to detecting the input satisfies one or more criteria, such as in FIG. 7B , where the first indication is different from the second indication. In some embodiments, displaying the second indication includes changing a visual characteristic (e.g., color, size, position, translucency) of a user interface element at which the user is viewing. For example, the second indication is the electronic device moving the user interface element toward the user within the three-dimensional environment. In some embodiments, the second indication is overlaid on or displayed near the user interface element at which the user is looking, hi some embodiments, the second indication is an icon or image displayed in a location within the user interface that is independent of the location at which the user's gaze is directed.

[0189] The above-described method of displaying an indication that a user's posture meets one or more criteria that are different from an indication of the location of the user's gaze provides an efficient way of indicating to a user that the electronic device is ready to accept further input from a predetermined part of the user, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further reduces power usage and improves the battery life of the electronic device by allowing the user to use the electronic device more quickly and efficiently.

[0190] In some embodiments, such as FIG. 7C , while displaying a user interface element (e.g., 705), the electronic device 101a detects (834a) a second input from a second predetermined portion (e.g., 717) (e.g., a second hand) of the user of the electronic device 101a via one or more input devices.

[0191] In some embodiments, in response to detecting a second input from a second predefined portion (e.g., 717) of a user of the electronic device (834b), and in accordance with determining that a posture (e.g., position, orientation, hand shape) of the second predefined portion (e.g., 711) of the user of the electronic device prior to detecting the second input satisfies one or more second criteria, such as in FIG. 7B , the electronic device 101a performs a second distinct action in accordance with the second input from the second predefined portion (e.g., 711) of the user of the electronic device 101a (834c). In some embodiments, the one or more second criteria differ from the one or more criteria in that a different predefined portion of the user performs the posture, but the one or more criteria and the one or more second criteria are otherwise the same. For example, the one or more criteria require the user's right hand to be in a ready state, such as a pre-pinch or pointing hand shape, and the one or more second criteria require the user's left hand to be in a ready state, such as a pre-pinch or pointing hand shape. In some embodiments, the one or more criteria differ from the one or more second criteria. For example, a first subset of postures may meet one or more criteria for a user's right hand, and a second, different subset of postures may meet one or more criteria for a user's left hand.

[0192] In some embodiments, such as in FIG. 7C , in response to detecting a second input from a second predefined portion (e.g., 715) of the user of electronic device 101b (834b), and in response to determining that a posture of the second predefined portion (e.g., 721) of the user prior to detecting the second input does not satisfy one or more second criteria, such as in FIG. 7B , the electronic device refrains from performing a second distinct action in accordance with the second input from the second predefined portion (e.g., 715) of electronic device 101b (834d). In some embodiments, the electronic device can detect inputs from the predefined portion and / or the second predefined portion of the user independently of one another. In some embodiments, to perform an action in accordance with an input provided by the user's left hand, the user's left hand must have a posture that satisfies one or more criteria prior to providing the input, and to perform an action in accordance with an input provided by the user's right hand, the user's right hand must have a posture that satisfies the second one or more criteria. In some embodiments, in response to detecting a posture of a predetermined portion of the user that satisfies the one or more criteria followed by an input provided by the second predetermined portion of the user without the second predetermined portion of the user first satisfying the second one or more criteria, the electronic device refrains from performing an action in accordance with the input of the second predetermined portion of the user. In some embodiments, in response to detecting a posture of a second predetermined portion of the user that satisfies the second one or more criteria followed by an input provided by the second predetermined portion of the user without the second predetermined portion of the user first satisfying the one or more criteria, the electronic device refrains from performing an action in accordance with the input of the second predetermined portion of the user.

[0193] The above method of accepting input from a second predetermined portion of the user independent of the predetermined portion of the user provides an efficient way of increasing the speed at which a user can provide input to an electronic device, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further reduces power usage and improves the battery life of the electronic device by allowing the user to use the electronic device more quickly and efficiently.

[0194] In some embodiments, such as in FIGS. 7A-7C , the user interface is accessible (836a) by electronic device 101a and second electronic device 101b (e.g., the electronic device and second electronic device are in communication (e.g., via a wired or wireless network connection)). In some embodiments, the electronic device and second electronic device are located remotely from one another. In some embodiments, the electronic device and second electronic device are co-located (e.g., in the same room, building, etc.). In some embodiments, the electronic device and second electronic device present the three-dimensional environment in a co-presence session, where representations of users of both devices are associated with unique locations within the three-dimensional environment and where each electronic device displays the three-dimensional environment from the perspective of a separate user's representation.

[0195] In some embodiments, such as FIG. 7A , before detecting that a posture of a predetermined portion (e.g., 709) of the user prior to detecting the input satisfies one or more criteria, the electronic device 101a displays (836b) a user interface element (e.g., 705) with visual characteristics (e.g., size, color, translucency, position) having a first value.

[0196] In some embodiments, such as in FIG. 7B , while a posture of the predefined portion of the user (e.g., 709) prior to detecting the input satisfies one or more criteria, the electronic device 101a displays (836c) the user interface element (e.g., 705) with visual characteristics (e.g., size, color, translucency, position) having a second value different from the first value. In some embodiments, the electronic device updates the visual appearance of the user interface element in response to detecting that the posture of the predefined portion of the user satisfies one or more criteria. In some embodiments, the electronic device updates only the appearance of the user interface element to which the user's attention is directed (e.g., according to the user's line of sight or the user's attention zone according to method 1000). In some embodiments, the second electronic device maintains the display of the user interface element with visual characteristics having the first value in response to the predefined portion of the user satisfying one or more criteria.

[0197] In some embodiments, while displaying the user interface element with the visual characteristic having the first value, while a posture of a predetermined portion of the second user of the second electronic device 101b satisfies one or more criteria (optionally in response to an indication to that effect), the electronic device 101a maintains display of the user interface element with the visual characteristic having the first value (836d), similar to the way electronic device 101b maintains display of the user interface element (e.g., 705) while a portion of the user of the first electronic device 101a (e.g., 709) satisfies one or more criteria in Figure 7B. In some embodiments, in response to detecting that a posture of the predetermined portion of the user of the second electronic device satisfies the one or more criteria, the second electronic device updates the user interface element to be displayed with the visual characteristic having the second value, similar to the way electronic devices 101a and 101b both scroll the user interface element (e.g., 705) in response to input detected by electronic device 101a (e.g., via hand 709 or 711) in Figure 7C. In some embodiments, in response to an indication that a posture of a user of the electronic device satisfies one or more criteria while displaying the user interface element with the visual characteristics having the first value, the second electronic device continues displaying the user interface element with the visual characteristics having the first value. In some embodiments, in response to a determination that a posture of a user of the electronic device satisfies one or more criteria and an indication that a posture of a user of the second electronic device satisfies the one or more criteria, the electronic device displays the user interface element with the visual characteristics having a third value.

[0198] The above-described method of not synchronizing updates of visual characteristics of user interface elements across electronic devices provides an efficient way of showing the parts of the user interface that a user is interacting with without causing confusion by also showing the parts of the user interface that other users are interacting with, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further reduces power usage and improves the battery life of the electronic device by allowing the user to use the electronic device more quickly and efficiently.

[0199] In some embodiments, in response to detecting an input from a predetermined portion of a user of the electronic device (e.g., 709 or 711), the electronic device 101a displays (836a) the user interface element (e.g., 705) with a visual characteristic having a third value (e.g., the third value is different from the first value and the second value), such as in Figure 7C. In some embodiments, in response to the input, the electronic device and the second electronic device perform respective actions according to the input.

[0200] In some embodiments, in response to an indication of input from a predetermined portion of a second user of the second electronic device (e.g., after the second electronic device detects that a predetermined portion of the user of the second electronic device satisfies one or more criteria), the electronic device 101a displays the user interface element (e.g., 705) with visual characteristics having a third value (836b), as if the electronic device 101b displayed the user interface element (e.g., 705) in the same manner as the electronic device 101a displayed the user interface element (e.g., 705) in response to the electronic device 101a detecting user input from the hand (e.g., 709 or 711) of the user of the electronic device 101a. In some embodiments, in response to the input from the second electronic device, the electronic device and the second electronic device perform respective operations according to the input. In some embodiments, the electronic device displays an indication that a user of the second electronic device has provided input directed at the user interface element, but does not present an indication of the hover state of the user interface element.

[0201] The above-described method of updating user interface elements in response to input, regardless of the device on which the input is detected, provides an efficient way of indicating the current interaction state of user interface elements displayed by both devices, which simplifies the interaction between the user and the electronic device (e.g., by clearly indicating which parts of the user interface other users are interacting with), improves the usability of the electronic device, and makes the user device interface more efficient, which in turn allows users to use the electronic device more quickly and efficiently, thereby reducing power usage, improving the battery life of the electronic device, and avoiding errors requiring subsequent correction caused by changes in the interaction state of the user interface elements.

[0202] 9A-9C illustrate an example method by which electronic device 101a processes user input based on a zone of attention associated with the user, according to some embodiments.

[0203] FIG. 9A illustrates electronic device 101a displaying a three-dimensional environment via display generation component 120a. It should be understood that in some embodiments, electronic device 101a utilizes one or more of the techniques described with reference to FIGS. 9A-9C in a two-dimensional environment or user interface without departing from the scope of this disclosure. As described above with reference to FIGS. 1-6, electronic device 101a optionally includes display generation component 120a (e.g., a touchscreen) and multiple image sensors 314a. The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor that electronic device 101a can use to capture one or more images of a user or a portion of a user while the user interacts with electronic device 101a. In some embodiments, display generation component 120a is a touchscreen capable of detecting a user's hand gestures and movements. In some embodiments, the user interface described below may also be implemented in a head-mounted display that includes a display generation component that displays the user interface to the user and sensors that detect the physical environment and / or the movement of the user's hands (e.g., external sensors facing outward from the user) and / or the user's line of sight (e.g., internal sensors facing inward toward the user's face).

[0204] 9A illustrates electronic device 101a presenting, via display generation component 120a, a first selectable option 903, a second selectable option 905, and a representation 904 of a table (e.g., table 604 of FIG. 6B ) within the physical environment of electronic device 101a. In some embodiments, representation 904 of the table is a photorealistic image (e.g., pass-through video or digital pass-through) of the table generated by display generation component 120a. In some embodiments, representation 904 of the table is a view of the table through a transparent portion of display generation component 120a (e.g., a true or actual pass-through). In some embodiments, electronic device 101a displays the three-dimensional environment from a perspective associated with a user of the electronic device within the three-dimensional environment.

[0205] In some embodiments, the electronic device 101a defines the user's zone of attention 907 as a cone volume within the three-dimensional environment based on the user's line of sight 901a. For example, the zone of attention 907 is optionally a cone centered on a line defined by the user's line of sight 901a (e.g., a line passing through the location of the user's line of sight in the three-dimensional environment and a viewpoint associated with the electronic device 101a) and includes a volume of the three-dimensional environment within a predetermined angle (e.g., 1, 2, 3, 5, 10, 15 degrees, etc.) from the line defined by the user's line of sight 901a. Thus, in some embodiments, the two-dimensional area of ​​the zone of attention 907 increases as a function of distance from the viewpoint associated with the electronic device 101a. In some embodiments, the electronic device 101a determines whether to respond to an input and / or a user interface element toward which the input is directed based on the user's zone of attention.

[0206] As shown in FIG. 9A , a first selectable option 903 is within a user's zone of attention 907, and a second selectable option 905 is outside the user's zone of attention. As shown in FIG. 9A , the selectable option 903 can be within the zone of attention 907 even when the user's gaze 901 a is not directed at the selectable option 903. In some embodiments, the selectable option 903 can be within the zone of attention 907 while the user's gaze is directed at the selectable option 903. FIG. 9A also shows a user's hand 909 in a direct input ready state (e.g., hand state D). In some embodiments, the direct input ready state is the same as or similar to the direct input ready state(s) described above with reference to FIGS. 7A-8K. Furthermore, in some embodiments, the direct input described herein shares one or more characteristics of the direct input described with reference to methods 800, 1200, 1400, 1600, 1800, and / or 2000. For example, user's hand 909 is in a pointing hand configuration and is within a direct readiness threshold distance (e.g., 0.5, 1, 2, 3, 5, 10, 15, 30 centimeters, etc.) of first selectable option 903. FIG. 9A also shows user's hand 911 in a direct input readiness state. In some embodiments, hand 911 is a substitute for hand 909. In some embodiments, electronic device 101a can detect two of the user's hands at once (e.g., according to one or more of method 1600). For example, user's hand 911 is in a pointing hand configuration and is within a readiness threshold distance of second selectable option 905.

[0207] In some embodiments, electronic device 101a requires that a user interface element be within zone of attention 907 to accept input. For example, because first selectable option 903 is within the user's zone of attention 907, electronic device 101a updates the first selectable option 903 to indicate that further input (e.g., from hand 909) is to be directed to the first selectable option 903. As another example, because second selectable option 905 is outside the user's zone of attention 907, electronic device 101a refrains from updating the second selectable option 905 to indicate that further input (e.g., from hand 911) is to be directed to the second selectable option 905. It should be understood that even though the user's line of sight 901a is not directed to the first selectable option 903, electronic device 101a is still configured to direct input to the first selectable option 903 because the first selectable option 903 is within zone of attention 907, which is optionally wider than the user's line of sight.

[0208] In FIG. 9B , electronic device 101a detects a user's hand 909 making a direct selection of first selectable option 903. In some embodiments, the direct selection includes moving hand 909 while the hand is in a pointing hand configuration to a location touching first selectable option 903 or within a direct selection threshold (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2 centimeters, etc.) thereof. As shown in FIG. 9B , once the input is detected, first selectable option 903 is no longer within user's attention zone 907. In some embodiments, attention zone 907 moves as user's gaze 901b moves. In some embodiments, after electronic device 101a detects the ready state of hand 909 shown in FIG. 9A , attention zone 907 moves to the location shown in FIG. 9B . In some embodiments, the input shown in FIG. 9B is detected before ready state 907 moves to the location shown in FIG. 9B . In some embodiments, the input shown in FIG. 9B is detected after ready state 907 moves to the location shown in FIG. 9B . Although the first selectable option 903 is no longer within the user's zone of attention 907, in some embodiments, the electronic device 101a still updates the color of the first selectable option 903 in response to the input, as the first selectable option 903 was within the zone of attention 907 during the ready state, as shown in FIG. 9A . In some embodiments, in addition to updating the appearance of the first selectable option 903, the electronic device 101a performs an action in accordance with the selection of the first selectable option 903. For example, the electronic device 101a performs an action, such as activating / deactivating a setting associated with the option 903, starting playback of content associated with the option 903, displaying a user interface associated with the option 903, or a different action associated with the option 903.

[0209] In some embodiments, the selection input is detected only in response to detecting that the user's hand 909 has moved from the side of the first selectable option 903 that is visible in Figure 9B to a location that touches or is within the direct selection threshold of the first selectable option 903. For example, if the user instead reaches around the first selectable option 903 and touches the first selectable option 903 from behind the first selectable option 903 that is not visible in Figure 9B, the electronic device 101a optionally refrains from updating the appearance of the first selectable option 903 and / or refrains from performing an action in accordance with the selection.

[0210] In some embodiments, in addition to continuing to accept press inputs (e.g., select inputs) that were initiated while the first selectable option 903 was within the zone of attention 907 and continued while the first selectable option 903 was not within the zone of attention 907, the electronic device 101a accepts other types of inputs that were initiated while the user interface element at which the input was directed was within the zone of attention, even if the user interface element is no longer within the zone of attention when the input continues. For example, the electronic device 101a can continue a drag input, in which the electronic device 101a updates the position of a user interface element in response to the user input, even if the drag input was continued after the user interface element moved outside the zone of attention (e.g., was initiated when the user interface element was within the zone of attention). As another example, the electronic device 101a can continue a scroll input in response to a user input, even if the scroll input was continued after the user interface element moved outside the zone of attention 907 (e.g., was initiated when the user interface element was within the zone of attention). As shown in FIG. 9A, in some embodiments, if a user interface element was within the attention zone when the ready state was detected, the input is accepted even if the user interface element at which the input is directed is outside the attention zone for part of the input.

[0211] Further, in some embodiments, the location of the attention zone 907 remains at a distinct position within the three-dimensional environment for a threshold time (e.g., 0.5, 1, 2, 3, 5 seconds, etc.) after detecting a movement of the user's gaze. For example, while the user's gaze 901a and the attention zone 907 are at the location shown in Figure 9A, the electronic device 101a detects that the user's gaze 901b moves to the location shown in Figure 9B. In this example, the attention zone 907 remains at the location shown in Figure 9A for the threshold time before moving the attention zone 907 to the location shown in Figure 9B in response to the user's gaze 901b moving to the location shown in Figure 9B. Thus, in some embodiments, inputs initiated after the user's gaze has moved and directed to user interface elements within the original zone of attention (e.g., zone of attention 907 in FIG. 9A) are optionally responded to by electronic device 101a so long as these inputs are initiated within a threshold time (e.g., 0.5, 1, 2, 3, 5 seconds, etc.) of the user's gaze moving to the location of FIG. 9B. In some embodiments, electronic device 101a does not respond to such inputs that are initiated after the threshold time of the user's gaze moving to the location of FIG. 9B.

[0212] In some embodiments, the electronic device 101a cancels the user input if the user moves their hand away from the user interface element at which the input is directed or if they do not provide further input for a threshold time (e.g., 1, 2, 3, 5, 10 seconds, etc.) after the ready state is detected. For example, if the user moves their hand 909 to the location shown in Figure 9C after the electronic device 101a detects the ready state as shown in Figure 9A, the electronic device 101a will revert the appearance of the first selectable option 903 to no longer indicating that input is directed at the first selectable option 903 and no longer accepting direct input from the hand 909 directed at option 903 (e.g., unless and until the ready state is again detected).

[0213] 9C , the first selectable option 903 is still within the user's zone of attention 907. The user's hand 909 is, optionally, in a hand shape corresponding to a direct ready state (e.g., a pointing hand shape, hand state D). Because the user's hand 909 has moved a threshold distance (e.g., 1, 2, 3, 5, 10, 15, 20, 30, 50 centimeters, etc.) away from the first selectable option 903 and / or because the user's hand 909 has moved a threshold distance (e.g., 1, 2, 3, 5, 10, 15, 20, 30, 50 centimeters, etc.) away from the first selectable option 903, the electronic device 101 a is no longer configured to direct input from the hand 909 to the first selectable option 903. 9A , the electronic device 101a stops directing further input from the hand to the first user interface element 903. Similarly, in some embodiments, if the user begins providing additional input (e.g., in addition to meeting the ready state criteria, begins providing a press input to the element 903 but has not yet reached the press distance threshold required to complete the press / select input) and then moves the hand a threshold distance away from the first selectable option 903 and / or moves the hand a threshold distance away from the first selectable option 903, the electronic device 101a cancels the input. As described above with reference to FIG. 9B, it should be understood that if an input is initiated while the first selectable option 903 is within the user's zone of attention 907, the electronic device 101a optionally does not cancel the input in response to detecting that the user's line of sight 901b or the user's zone of attention 907 moves away from the first selectable option 903.

[0214] 9A-9C illustrate examples of determining whether to accept direct input directed at a user interface element based on a user's zone of attention 907, it should be understood that electronic device 101a may similarly determine whether to accept indirect input directed at a user interface element based on a user's zone of attention 907. For example, various results shown in and described with reference to FIGS. 9A-9C optionally apply to indirect input as well (e.g., as described with reference to methods 800, 1200, 1400, 1800, etc.). In some embodiments, a zone of attention is not required to accept direct input, but is required for indirect input.

[0215] 10A-10H are flowcharts illustrating a method 1000 for processing user input based on a zone of attention associated with a user, according to some embodiments. In some embodiments, method 1000 is performed on a computer system (e.g., computer system 101 of FIG. 1 , such as a tablet, smartphone, wearable computer, or head-mounted device) that includes a display generation component (e.g., display generation component 120 of FIGS. 1, 3, and 4) (e.g., a head-up display, a display, a touchscreen, a projector, etc.) and one or more cameras (e.g., a camera pointing downward in a user's hand (e.g., a color sensor, an infrared sensor, and other depth-sensing camera) or a camera pointing forward from the user's head). In some embodiments, method 1000 is governed by instructions stored on a non-transitory computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A ). Some operations of method 1000 are optionally combined and / or the order of some operations is optionally changed.

[0216] In some embodiments, method 1000 is performed on an electronic device 101 a in communication with a display generation component and one or more input devices (e.g., a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer). In some embodiments, the display generation component is a display integrated with the electronic device (optionally a touchscreen display), an external display such as a monitor, projector, television, or a hardware component (optionally built-in or external) for projecting a user interface and making the user interface visible to one or more users, etc. In some embodiments, the one or more input devices include electronic devices or components capable of receiving user input (e.g., capturing user input, detecting user input, etc.) and transmitting information related to the user input to the electronic device. Examples of input devices include a touchscreen, a mouse (e.g., external), a trackpad (optionally integrated or external), a touchpad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the electronic device), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, an eye tracking device, and / or a motion sensor (e.g., hand tracking device, hand motion sensor), etc. In some embodiments, the electronic device is in communication with a hand tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touchscreen, trackpad). In some embodiments, the hand tracking device is a wearable device such as a smart glove. In some embodiments, the hand tracking device is a handheld input device such as a remote control or a stylus.

[0217] In some embodiments, such as in FIG. 9A , electronic device 101a displays (1002a) a first user interface element (e.g., 903, 905) via display generation component 120a. In some embodiments, the first user interface element is an interactive user interface element, and in response to detecting input directed at the first user interface element, the electronic device performs an action associated with the first user interface element. For example, the first user interface element is a selectable option that, when selected, causes the electronic device to perform an action, such as displaying a separate user interface, changing a setting on the electronic device, or starting playback of content. As another example, the first user interface element is a container (e.g., a window) in which a user interface / content is displayed, and in response to detecting selection of the first user interface element followed by movement input, the electronic device updates the position of the first user interface element according to the movement input. In some embodiments, the user interface and / or user interface elements are displayed in (e.g., the user interface is and / or is displayed within) a three-dimensional environment that is generated, displayed, or otherwise made viewable by a device (e.g., a computer-generated reality (CGR) environment such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment).

[0218] 9B , while displaying a first user interface element (e.g., 909), the electronic device 101a detects (1002b) a first input directed at the first user interface element (e.g., 909) via one or more input devices. In some embodiments, detecting the first user input includes detecting via a hand tracking device that the user has performed a predetermined gesture (e.g., a pinch gesture in which the user touches the thumb to another finger (e.g., index finger, middle finger, ring finger, pinky finger) of the same hand). In some embodiments, detecting the input includes detecting that the user has performed a pointing gesture in which one or more fingers are extended and one or more fingers are curled toward the user's palm, moving the user's hand a predetermined distance (e.g., 2, 5, 10 centimeters, etc.) from the user's torso in a pressing or pushing motion. In some embodiments, the pointing gesture and the pushing action are detected while the user's hand is within a threshold distance (e.g., 1, 2, 3, 5, 10 centimeters, etc.) of a first user interface element in the three-dimensional environment. In some embodiments, the three-dimensional environment includes a representation of a virtual object and a user. In some embodiments, the three-dimensional environment includes a representation of the user's hand, which may be a photorealistic representation of the hand, a pass-through video of the user's hand, or a view of the user's hand through a transparent portion of a display generation component. In some embodiments, the input is a direct or indirect interaction with a user interface element, such as described with reference to methods 800, 1200, 1400, 1600, 1800, and / or 2000.

[0219] In some embodiments, in response to detecting 1002c a first input directed at a first user interface element (e.g., 903), in accordance with determining (e.g., when the first input is detected) that the first user interface element (e.g., 903) is within a zone of attention (e.g., 907) associated with a user of electronic device 101a, such as in FIG. 9A , electronic device 101a performs 1002d a first action corresponding to the first user interface element (e.g., 903). In some embodiments, the zone of attention includes a region of the three-dimensional environment within a predetermined threshold distance (e.g., 5, 10, 30, 50, 100 centimeters, etc.) and / or threshold angle (e.g., 5, 10, 15, 20, 30, 45 degrees, etc.) of a location in the three-dimensional environment to which the user's gaze is directed. In some embodiments, the zone of attention includes a region of the three-dimensional environment between a location in the three-dimensional environment to which the user's gaze is directed and one or more physical features of the user (e.g., the user's hand, arm, shoulder, torso, etc.). In some embodiments, the attention zone is a three-dimensional region of the three-dimensional environment. For example, the attention zone is cone-shaped, with the tip of the cone corresponding to the user's eye / point of view and the base of the cone corresponding to the area of ​​the three-dimensional environment toward which the user's gaze is directed. In some embodiments, a first user interface element is within the attention zone associated with the user while the user's gaze is directed toward the first user interface element and / or when the first user interface element enters the cone volume of the attention zone. In some embodiments, the first action is one of making a selection, activating a setting on the electronic device, initiating a process to move a virtual object within the three-dimensional environment, displaying a new user interface that is not currently displayed, playing an item of content, saving a file, initiating communication (e.g., phone call, email, message) with another user, and / or scrolling a user interface. In some embodiments, the first input is detected by detecting a posture and / or movement of a predetermined portion of the user.For example, the electronic device detects a user moving a finger to a location within a threshold distance (e.g., 0.1, 0.3, 0.5, 1, 3, 5 centimeters, etc.) of a first user interface element in the three-dimensional environment, with the hand / finger in a position corresponding to the index finger of the hand pointing with the other fingers bent into the hand.

[0220] 9A , in response to detecting 1002c a first input directed at a first user interface element (e.g., 905), and in response to determining that the first user interface element (e.g., 905) is not within a zone of attention associated with the user (e.g., when the first input is detected), the electronic device 101a refrains from performing 1002e a first action. In some embodiments, the first user interface element is not within a zone of attention associated with the user if the user's gaze is directed at a user interface element other than the first user interface element and / or if the first user interface element does not fall within a cone volume of the zone of attention.

[0221] The above method of performing or not performing a first action depending on whether a first user interface element is within a zone of attention associated with a user provides an efficient way of reducing accidental user input, which simplifies the interaction between the user and the electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving battery life of the electronic device while reducing errors in use.

[0222] In some embodiments, the first input directed to the first user interface element (e.g., 903) is an indirect input (1004a) directed to the first user interface element (e.g., 903 in FIG. 9C). In some embodiments, the indirect input is an input provided by a predetermined part of the user (e.g., the user's hand, finger, arm, etc.) while the predetermined part of the user is beyond a threshold distance (e.g., 0.2, 1, 2, 3, 5, 10, 30, 50 centimeters, etc.) from the first user interface element. In some embodiments, the indirect input is similar to the indirect input described with reference to methods 800, 1200, 1400, 1600, 1800, and / or 2000.

[0223] 9B , while displaying the first user interface element (e.g., 905), the electronic device 101a detects (1004b) a second input via one or more input devices, where the second input corresponds to a direct input directed at a respective user interface element (e.g., 903). In some embodiments, the direct input is similar to the direct input described with reference to methods 800, 1200, 1400, 1600, 1800, and / or 2000. In some embodiments, the direct input is provided by a predetermined part of the user (e.g., hand, finger, arm) while the predetermined part of the user is less than a threshold distance (e.g., 0.2, 1, 2, 3, 5, 10, 30, 50 centimeters, etc.) from the first user interface element. In some embodiments, detecting the direct input includes detecting a ready state of the hand (e.g., a pointing hand with one or more fingers extended and one or more fingers curled toward the palm) followed by detecting that the user is performing a predefined gesture with their hand (e.g., a press gesture in which the user moves an extended finger to the location of a respective user interface element and the other fingers are curled toward the palm). In some embodiments, the ready state is detected according to one or more of method 800.

[0224] 9B , in response to detecting the second input, the electronic device 101a performs (1004c) an action associated with the individual user interface element (e.g., 903) regardless of whether the individual user interface element is within a zone of attention (e.g., 907) associated with the user (e.g., because it is a direct input). In some embodiments, the electronic device performs (1004c) an action associated with the first user interface element in response to the indirect input only if the indirect input is detected while the user's gaze is directed toward the first user interface element. In some embodiments, the electronic device performs (1004c) an action associated with the user interface element within the user's zone of attention in response to the direct input regardless of whether the user's gaze is directed toward the user interface element when the direct input is detected.

[0225] The above method of refraining from performing a second action in response to detecting an indirect input while the user's gaze is not directed at a first user interface element provides a way to reduce or prevent the execution of actions that the user does not want, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which in turn allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving the battery life of the electronic device while reducing errors in use.

[0226] In some embodiments, such as in FIG. 9B , a zone of attention (e.g., 907) associated with a user is based (1006a) on the direction (and / or location) of the line of sight (e.g., 901b) of a user of an electronic device. In some embodiments, the zone of attention is defined as a cone-shaped volume (e.g., extending from a point at the user's viewpoint into the three-dimensional environment) that includes the point in the three-dimensional environment at which the user is looking and locations in the three-dimensional environment between the point at which the user is looking and the user within a predetermined threshold angle (e.g., 5, 10, 15, 20, 30, 45 degrees, etc.) of the user's line of sight. In some embodiments, in addition to or instead of being based on the user's line of sight, the zone of attention is based on the orientation of the user's head. For example, the zone of attention is defined as a cone-shaped volume that includes locations in the three-dimensional environment within a predetermined threshold angle (e.g., 5, 10, 15, 20, 30, 45 degrees, etc.) of a line perpendicular to the user's face. As another example, the attention zone is a cone centered at the average of a line extending from the user's gaze and a line perpendicular to the user's face, or the sum of a cone centered at the user's gaze and a cone centered at a line perpendicular to the user's face.

[0227] The above-described method of basing attention zones on the direction of a user's gaze provides an efficient way to direct user input based on gaze without additional input (e.g., to move the input focus, such as moving a cursor), which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving the battery life of the electronic device while reducing errors in use.

[0228] 9A , while a first user interface element (e.g., 903) is within a zone of attention (e.g., 907) associated with a user, the electronic device 101a detects (1008a) that one or more criteria for moving the zone of attention (e.g., 903) to a location where the first user interface element (e.g., 903) is not within the zone of attention. In some embodiments, the zone of attention is based on the user's line of sight, and the one or more criteria are met when the user's line of sight moves to a new location such that the first user interface element is no longer within the zone of attention. For example, the zone of attention includes an area of ​​the user interface within 10 degrees of a line along the user's line of sight, and the user's line of sight moves to a location such that the first user interface element is more than 10 degrees from the user's line of sight.

[0229] 9B , after detecting that one or more criteria have been met (1008b), the electronic device 101a detects (1008c) a second input directed at the first user interface element (e.g., 903). In some embodiments, the second input is a direct input in which the user's hand is within a threshold distance (e.g., 0.2, 1, 2, 3, 5, 10, 30, 50 centimeters, etc.) of the first user interface element.

[0230] 9B , after detecting that one or more criteria have been met (1008b), in response to detecting a second input (1008d) directed at the first user interface element (e.g., 903), the electronic device 101a performs a second action (1008e) corresponding to the first user interface element (e.g., 903) in accordance with a determination that the second input was detected within a respective time threshold (e.g., 0.01, 0.02, 0.05, 0.1, 0.2, 0.3, 0.5, 1 second, etc.) of the one or more criteria being met. In some embodiments, the user's zone of attention remains stationary until the time threshold (e.g., 0.01, 0.02, 0.05, 0.1, 0.2, 0.3, 0.5, 1 second, etc.) has elapsed since the one or more criteria were met.

[0231] 9B , after detecting that one or more criteria have been met (1008b), in response to detecting a second input directed at a first user interface element (e.g., 903) (1008d), the electronic device 101a refrains from performing a second action (1008f) in accordance with a determination that the second input was detected after a respective time threshold (e.g., 0.01, 0.02, 0.05, 0.1, 0.2, 0.3, 0.5, 1 second, etc.) of the one or more criteria being met. In some embodiments, once the time threshold (e.g., 0.01, 0.02, 0.05, 0.1, 0.2, 0.3, 0.5, 1 second, etc.) has elapsed since the one or more criteria for moving the attention zone were met, the electronic device updates the position of the attention zone associated with the user (e.g., based on the user's new gaze location). In some embodiments, the electronic device moves the attention zone gradually over the time threshold, and begins the movement with or without a time delay after detecting the user's gaze movement. In some embodiments, the electronic device refrains from performing the second action in response to an input detected while the first user interface element is not within the user's zone of attention.

[0232] The above-described method of performing a second action in response to a second input being received within a time threshold of one or more criteria for moving the attention zone provides an efficient way of accepting user input without requiring the user to maintain gaze for the duration of the input and avoiding accidental input by preventing activation of user interface elements after the predetermined time threshold has elapsed and the attention zone has moved, which simplifies the interaction between the user and the electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving battery life of the electronic device while reducing errors in use.

[0233] 9A-9B, the first input includes a first portion followed by a second portion (1010a). In some embodiments, detecting the first portion of the input includes detecting a user's predefined portion readiness state, as described with reference to method 800. In some embodiments, in response to the first portion of the input, the electronic device moves the input focus to a respective user interface element. For example, the electronic device updates the appearance of the respective user interface element to indicate that the input focus is directed to the respective user interface element. In some embodiments, the second portion of the input is a selection input. For example, a first portion of the input includes detecting the user's hand within a first threshold distance (e.g., 3, 5, 10, 15 centimeters, etc.) of the individual user interface elements while making a predetermined hand shape (e.g., a pointing hand shape with one or more fingers extended and one or more fingers curled toward the palm), and a second portion of the input includes detecting the user's hand within a second, lower threshold distance (e.g., contact, 0.1, 0.3, 0.5, 1, 2 centimeters, etc.) of the individual user interface elements while maintaining the pointing hand shape.

[0234] In some embodiments, such as FIG. 9A , while detecting a first input (1010b), the electronic device 101a detects a first portion of the first input while a first user interface element (e.g., 903) is within a zone of attention (e.g., 907) (1010c).

[0235] In some embodiments, such as FIG. 9A , while detecting a first input (1010b), in response to detecting a first portion of the first input, the electronic device 101a performs a first portion of a first action corresponding to a first user interface element (e.g., 903) (1010d). In some embodiments, the first portion of the first action includes identifying the first user interface element as having input focus of the electronic device and / or updating the appearance of the first user interface element to indicate that the input focus is directed toward the first user interface element. For example, in response to detecting that the user is making a pre-pinch hand gesture within a threshold distance (e.g., 1, 2, 3, 5, 10 centimeters, etc.) of the first user interface element, the electronic device changes the color of the first user interface element to indicate that the input focus is directed toward the first user interface element (e.g., similar to a cursor “hovering” over the user interface element). In some embodiments, the first portion of the input includes a selection of scrollable content within the user interface and a first portion of the user’s movement of a predetermined portion. In some embodiments, in response to a first portion of the user's movement of the predetermined portion, the electronic device scrolls the scrollable content by a first amount.

[0236] In some embodiments, such as in FIG. 9B , while detecting the first input (1010b), the electronic device 101a detects a second portion of the first input (1010e) while the first user interface element (e.g., 903) is outside the zone of attention. In some embodiments, after detecting the first portion of the first input and before detecting the second portion of the second input, the electronic device detects that the zone of attention no longer includes the first user interface element. For example, the electronic device detects a user's gaze directed toward a portion of the user interface such that the first user interface element is outside a distance or angle threshold of the user's zone of attention. For example, the electronic device detects a user making a pinch hand shape within a threshold distance (e.g., 1, 2, 3, 5, 10 centimeters, etc.) of the first user interface element while the zone of attention does not include the first user interface element. In some embodiments, the second portion of the first input includes a continuation of the user's movement of a predetermined portion. In some embodiments, in response to the user's continued movement of the predetermined portion, the electronic device continues scrolling the scrollable content. In some embodiments, the second portion of the first input is detected after a threshold time has elapsed since detecting the first portion of the input (e.g., the threshold time that an input must be detected after a ready state for the input is detected in order to trigger an action, as described above).

[0237] In some embodiments, such as in FIG. 9B , while detecting the first input (1010b), in response to detecting a second portion of the first input, the electronic device 101a performs (1010f) a second portion of the first operation corresponding to a first user interface element (e.g., 903). In some embodiments, the second portion of the first operation is an operation performed in response to detecting a selection of the first user interface element. For example, if the first user interface element is an option for starting playback of a content item, the electronic device starts playback of the content item in response to detecting the second portion of the first operation. In some embodiments, the electronic device performs the operation in response to detecting the second portion of the first input a threshold time after detecting the first portion of the input (e.g., a threshold time that the input must be detected after a ready state for input is detected in order to trigger an action as described above).

[0238] The above-described method of performing a second portion of a first operation corresponding to a first user interface element in response to detecting a second portion of the input while the first user interface element is outside of a zone of attention provides an efficient way of performing an operation in response to an input initiated while the first user interface element is outside of a zone of attention, even if the zone of attention moves away from the first user interface element before the input is completed, which simplifies the interaction between a user and an electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further allows a user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving battery life of the electronic device while reducing errors in use.

[0239] 9A-9B , the first input corresponds to a press input, a first portion of the first input corresponds to the initiation of the press input, and a second portion of the first input corresponds to a continuation of the press input (1012a). In some embodiments, detecting the press input includes detecting that a user has made a predetermined shape with their hand (e.g., a pointing shape with one or more fingers extended and one or more fingers bent toward the palm). In some embodiments, detecting the initiation of the press input includes detecting that the user is making the predetermined shape with their hand while their hand or a portion of their hand (e.g., the tip of one of the extended fingers) is within a first threshold distance (e.g., 3, 5, 10, 15, 30 centimeters, etc.) of the first user interface element. In some embodiments, detecting the continuation of the press input includes detecting that the user is making the predetermined shape with their hand while their hand or a portion of their hand (e.g., the tip of one of the extended fingers) is within a second threshold distance (e.g., 0.1, 0.5, 1, 2 centimeters, etc.) of the first user interface element. In some embodiments, the electronic device, in response to detecting the initiation of a press input while the first user interface element is within the zone of attention, performs a second action corresponding to the first user interface element, followed by a continuation of the press input (whether or not the first user interface element is within the zone of attention). In some embodiments, in response to a first portion of the press input, the electronic device presses the user interface element away from the user less than the full amount necessary to cause an action in accordance with the press input. In some embodiments, in response to a second portion of the press input, the electronic device continues to press the user interface element the full amount necessary to cause an action, and in response, performs an action in accordance with the press input.

[0240] The above-described method of detecting an imitation of a press input while a first user interface element is in a zone of attention, followed by performing a second action in response to the continued press input, provides an efficient way of detecting user input with a hand tracking device (and optionally an eye tracking device) without an additional input device, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving the battery life of the electronic device while reducing errors in use.

[0241] In some embodiments, the first input corresponds to a drag input, a first portion of the first input corresponds to the initiation of the drag input, and a second portion of the first input corresponds to a continuation of the drag input (1014a). For example, if a user moves hand 909 while selecting user interface element 903 in FIG. 9B , the input is a drag input. In some embodiments, the drag input includes a selection of the user interface element, a movement input, and an end of the drag input (e.g., a release of the selection input, similar to unclicking a mouse or lifting a finger from a touch sensor panel (e.g., a trackpad, touchscreen)). In some embodiments, the initiation of the drag input includes a selection of the user interface element toward which the drag input is directed. For example, the electronic device selects the user interface element in response to detecting that the user makes a pinch hand shape while the hand is within a threshold distance (e.g., 1, 2, 5, 10, 15, 30 centimeters, etc.) of the user interface element. In some embodiments, the continuation of the drag input includes a movement input while the selection is maintained. For example, the electronic device detects that the user maintains a pinch hand shape while moving the hand and moves the user interface element according to the hand movement. In some embodiments, the continuation of the drag input includes an end of the drag input. For example, the electronic device detects that the user has stopped making a pinch hand shape, such as by removing the thumb from the fingers. In some embodiments, the electronic device detects a selection of the first user interface element while the first user interface element is within the attention zone, and performs an action corresponding to the drag input (e.g., moving the first user interface element, scrolling the first user interface element, etc.) in response to detecting an end of the movement input and / or the drag input while the first user interface element is within or not within the attention zone. In some embodiments, the first portion of the input includes a selection of the user interface element and a portion of the user's movement of a predetermined portion.In some embodiments, in response to the first portion of the input, the electronic device moves the user interface element a first amount according to the amount of movement of a predetermined portion of the electronic device during the first portion of the input. In some embodiments, the second portion of the input includes continued movement of the predetermined portion by the user. In some embodiments, in response to the second portion of the input, the electronic device continues to move the user interface element an amount according to the movement of the predetermined portion by the user during the second portion of the user input.

[0242] The above-described method of detecting the initiation of a drag input while a first user interface element is within a zone of attention and performing an action in response to detecting a continuation of the drag input while the first user interface element is not within the zone of attention provides an efficient way of performing an action in response to a drag input initiated while the first user interface element is within a zone of attention, even if the zone of attention moves away from the first user interface element before the drag input is completed, which simplifies the interaction between a user and an electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further allows a user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving battery life of the electronic device while reducing errors in use.

[0243] 9A-9B , the first input corresponds to a selection input, a first portion of the first input corresponds to the beginning of the selection input, and a second portion of the first input corresponds to a continuation of the selection input (1016a). In some embodiments, the selection input includes detecting that the input focus is directed toward the first user interface element, detecting the beginning of a request to select the first user interface element, and detecting the end of the request to select the first user interface element. In some embodiments, the electronic device directs the input focus toward the first user interface element in response to detecting that the user's hand is directed toward the first user interface element in a ready state according to method 800. In some embodiments, the request to direct the input focus toward the first user interface element is similar to a cursor hover. For example, the electronic device detects a user making a pointing hand shape while the hand is within a threshold distance (e.g., 1, 2, 3, 5, 10, 15, 30 centimeters, etc.) of the first user interface element. In some embodiments, initiation of the request to select the first user interface element includes detecting a selection input similar to a mouse click or a touch down on a touch sensor panel. For example, the electronic device detects a user maintaining a pointing hand shape while the hand is within a second threshold distance (e.g., 0.1, 0.2, 0.3, 0.5, 1 centimeter, etc.) of the first user interface element. In some embodiments, termination of the request to select the user interface element is similar to releasing a mouse click or lifting off the touch sensor panel. For example, the electronic device detects that the user has removed their hand from the first user interface element by at least the second threshold distance (e.g., 0.1, 0.2, 0.3, 0.5, 1 centimeter, etc.).In some embodiments, the electronic device detects that input focus is directed toward the first user interface element while the first user interface element is within the attention zone, and performs a selection operation in response to detecting the start and end of a request to select the first user interface element while the first user interface element is within the attention zone or not.

[0244] The above-described method of performing an action in response to detecting an imitation of a selection input while a first user interface element is in a zone of attention, regardless of whether a continuation of the selection input is detected while the first user interface element is in a zone of attention, provides an efficient way of performing an action in response to a selection input initiated while the first user interface element is in a zone of attention, even if the zone of attention moves away from the first user interface element before the selection input is completed, which simplifies the interaction between a user and an electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further reduces power usage and improves battery life of the electronic device by allowing a user to use the electronic device more quickly and efficiently, while reducing errors in use.

[0245] 9A , detecting the first portion of the first input includes detecting a predefined portion of the user (e.g., 909) having a distinct pose (e.g., a hand shape including a pointing hand shape with one or more fingers extended and one or more fingers curled toward a palm, such as the ready state described with reference to method 800) without detecting movement of the predefined portion of the user (e.g., 909) within a distinct distance (e.g., 1, 2, 3, 5, 10, 15, 30 centimeters, etc.) of a location corresponding to the first user interface element (e.g., 903), and detecting the second portion of the first input includes detecting movement of the predefined portion of the user (e.g., 909) (1018a), such as in FIG. 9B . In some embodiments, detecting a predefined portion of the user having a distinct pose and within a distinct distance of the first user interface element includes detecting the ready state according to one or more of method 800. In some embodiments, the movement of the predetermined portion of the user includes movement from a discrete pose to a second pose associated with selecting a user interface element and / or movement from a discrete distance to a second distance associated with selecting a user interface element. For example, forming a pointing hand shape within a discrete distance of a first user interface element is a first portion of a first input, and maintaining the pointing hand shape while moving the hand to a second distance (e.g., within 0.1, 0.2, 0.3, 0.5, 1 centimeter, etc.) from the first user interface element is a second portion of the first input. As another example, forming a pre-pinch hand shape in which the thumb of a hand is within a threshold distance (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, 3 centimeters, etc.) of another finger of the hand is a first portion of a first input, and detecting a movement of the hand from the pre-pinch shape to a pinch shape in which the thumb is touching the other fingers is a second portion of the first input. In some embodiments, the electronic device detects further movement of the hand following the second portion of the input, such as a hand movement corresponding to a request to drag or scroll the first user interface element.In some embodiments, the electronic device performs an action in response to detecting a predetermined portion of the user having a respective pose while within a respective distance of the first user interface element while the first user interface element is in a zone of attention associated with the user, and subsequently detects movement of the predetermined portion of the user while the first user interface element is in the zone of attention or not.

[0246] The above-described method of performing an action in response to detecting a discrete pose of a predetermined portion of a user within a discrete distance of a first user interface element while the first user interface element is within a zone of attention, and subsequently detecting movement of the predetermined portion of the user while the first user interface element is within a zone of attention, provides an efficient way of performing an action in response to an input initiated while the first user interface element is within a zone of attention, even if the zone of attention moves away from the first user interface element before the input is completed, which simplifies the interaction between a user and an electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further allows a user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving battery life of the electronic device while reducing errors in use.

[0247] In some embodiments, such as FIG. 9B , the first input is provided by a predetermined part of the user (e.g., 909) (e.g., the user's finger, hand, arm, or head), and detecting the first input includes detecting the predetermined part of the user (e.g., 909) within a distance threshold (e.g., 1, 2, 3, 5, 10, 15, 30 centimeters) of a location corresponding to the first user interface element (e.g., 903) (1020a).

[0248] As shown in FIG. 9C, in some embodiments, while detecting a first input directed at a first user interface element (e.g., 903) and before performing a first action, the electronic device 101a detects (1020b) movement of a predetermined portion of the user (e.g., 909) via one or more input devices to a distance that exceeds a distance threshold from a location corresponding to the first user interface element (e.g., 903).

[0249] 9C , in response to detecting movement of the predefined portion (e.g., 909) a distance greater than a distance threshold from a location corresponding to the first user interface element (e.g., 903), the electronic device 101a refrains from performing (1020c) the first operation corresponding to the first user interface element (e.g., 903). In some embodiments, in response to detecting that a user begins to provide input directed at the first user interface element and then moves a predefined portion of the user beyond a threshold distance from a location corresponding to the user interface element before completing the input, the electronic device refrains from performing the first operation corresponding to the input directed at the first user interface element. In some embodiments, the electronic device refrains from performing the first operation in response to the user moving a predefined portion of the user at least a distance threshold away from a location corresponding to the first user interface element, even if the user performed one or more portions of the first input without performing the complete first input while the predefined portion of the user is within the distance threshold of the location corresponding to the first user interface element. For example, the selection input includes detecting a user making a pre-pinch hand shape (e.g., a hand shape with the thumb within a threshold distance (e.g., 0.1, 0.2, 0.5, 1, 2, 3 centimeters, etc.)), followed by a pinch hand shape (e.g., the thumb touching the fingers), followed by the end of the pinch hand shape (e.g., the thumb no longer touching the fingers, the thumb at least 0.1, 0.2, 0.5, 1, 2, 3 centimeters, etc. from the fingers). In this example, the electronic device refrains from performing the first action if the end of the pinch gesture is detected while the hand is more than a threshold distance (e.g., 1, 2, 3, 5, 10, 15, 30 centimeters, etc.) from a location corresponding to the first user interface element, even if the hand was within the threshold distance when the pre-pinch hand shape and / or the pinch hand shape were detected.

[0250] The above-described method of forgoing performing a first action in response to detecting movement of a predetermined portion of a user a distance that exceeds a distance threshold provides an efficient way of canceling a first action after a portion of a first input has been provided, which simplifies the interaction between a user and an electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further allows a user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving battery life of the electronic device while reducing errors in use.

[0251] 9A , the first input is provided by a predetermined portion (e.g., 909) of the user (e.g., a finger, hand, arm, or head of the user), and detecting the first input includes detecting 1022a the predetermined portion (e.g., 909) of the user in a particular spatial relationship to a location corresponding to the first user interface element (e.g., 903) (e.g., detecting the predetermined portion of the user within a predetermined threshold distance (e.g., 1, 2, 3, 5, 10, 15, 20, 30 centimeters, etc.) of the first user interface element, at a predetermined orientation or posture relative to the user interface element). In some embodiments, the particular spatial relationship to a location corresponding to the first user interface element is a portion of the user that is in a ready state according to one or more of method 800.

[0252] 9A , during the first input and prior to performing the first action, while a predefined portion of the user (e.g., 909) is in a particular spatial relationship with a location corresponding to the first user interface element (e.g., 903), the electronic device 101a detects (1022b), via one or more input devices, that the predefined portion of the user (e.g., 909) has not engaged with (e.g., provided no additional input toward) the first user interface element (e.g., 903) within a particular time threshold (e.g., 1, 2, 3, 5 seconds, etc.) of entering the particular spatial relationship with the location corresponding to the first user interface element (e.g., 903). In some embodiments, the electronic device detects a ready state of the predefined portion of the user according to one or more of method 800 without detecting any further input within the time threshold. For example, the electronic device detects that the user's hand is in a pre-pinch hand shape (e.g., the thumb is within a threshold distance (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, 3, etc.) of another finger of the thumb hand) while the hand is within a predetermined threshold distance (e.g., 1, 2, 3, 5, 10, 15, 20, 30 centimeters, etc.) of a first user interface element without detecting a pinch hand shape (e.g., thumb and fingers touching) within a predetermined time period.

[0253] In some embodiments, in response to detecting that a predetermined portion of the user (e.g., 909) has not engaged with the first user interface element (e.g., 903) within a respective time threshold within which the predetermined portion is in a respective spatial relationship with respect to a location corresponding to the first user interface element (e.g., 903), the electronic device 101a forgoes execution of the first action corresponding to the first user interface element (e.g., 903) (1022c), such as in Figure 9C. In some embodiments, in response to detecting a predetermined portion of the user engaging with the first user interface element after the respective time threshold has elapsed, the electronic device forgoes execution of the first action corresponding to the first user interface element. For example, in response to detecting the lapse of a predetermined time threshold while detecting the user's hand in a pre-pinch hand shape (e.g., the thumb within a threshold distance (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, 3, etc.) of another finger of the thumb's hand) while the hand is within a predetermined threshold distance (e.g., 1, 2, 3, 5, 10, 15, 20, 30 centimeters, etc.) of a first user interface element before detecting a pinch hand shape (e.g., thumb and fingers touching), the electronic device refrains from performing a first action even though the pinch hand shape is detected after the lapse of the predetermined threshold time. In some embodiments, in response to detecting a predetermined portion of the user in a respective spatial relationship with a location corresponding to the user interface element, the electronic device updates the appearance of the user interface element (e.g., updates the color, size, translucency, position, etc. of the user interface element). In some embodiments, after a respective time threshold without detecting further input from the predetermined portion of the user, the electronic device returns the updated appearance of the user interface element.

[0254] The above-described method of forgoing a first action in response to detecting the lapse of a threshold time without a predetermined portion of the user engaging the first user interface element provides an efficient way of canceling a request to perform a first action, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving the battery life of the electronic device while reducing errors in use.

[0255] In some embodiments, a first portion of the first input is detected while the user's gaze is directed at the first user interface element (e.g., when gaze 901a in FIG. 9A is directed at user interface element 903), and a second portion of the first input following the first portion of the first input is detected (1024a) while the user's gaze (e.g., 901b) is not directed at the first user interface element (e.g., 903), such as in FIG. 9B. In some embodiments, in response to detecting the first portion of the first input while the user's gaze is directed at the first user interface element and subsequently detecting the second portion of the first input while the user's gaze is not directed at the first user interface element, the electronic device performs an action associated with the first user interface element. In some embodiments, in response to detecting the first portion of the first input while the first user interface element is within the attention zone and subsequently detecting the second portion of the first input while the first user interface element is not within the attention zone, the electronic device performs an action associated with the first user interface element.

[0256] The above-described method of performing an action in response to detecting a first portion of a first input while a user's gaze is directed toward a first user interface element, followed by detecting a second portion of the first input while the user's gaze is not directed toward the first user interface element, provides an efficient way of allowing a user to look away from the first user interface element without canceling the first input, which simplifies the interaction between the user and the electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further allows a user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving the battery life of the electronic device while reducing errors in use.

[0257] 9B , the first input is provided by a predetermined part (e.g., 909) (e.g., a finger, hand, arm, etc.) of the user moving to a location corresponding to the first user interface element (e.g., 903) from within a predetermined range of angles relative to the first user interface element (e.g., 903) (1026a) (e.g., the first user interface object is a three-dimensional virtual object accessible from multiple angles). For example, the first user interface object is a virtual video player including a surface on which content is presented, and the first input is provided by moving the user's hand to the first user interface object by touching the surface of the first user interface object on which content is presented before touching any other surface of the first user interface object.

[0258] In some embodiments, the electronic device 101a detects (1026b) a second input directed toward the first user interface element (e.g., 903) via one or more input devices, where the second input includes a predetermined portion (e.g., 909) of the user moving to a location corresponding to the first user interface element (e.g., 903) from outside a predetermined range of angles relative to the first user interface element (e.g., 903), such as when the hand (e.g., 909) in FIG. 9B approaches the user interface element (e.g., 903) from a side of the user interface element (e.g., 903) opposite the side of the user interface element (e.g., 903) visible in FIG. 9B. For example, the electronic device detects that the user's hand touches a surface of the virtual video player other than the surface on which content is presented (e.g., touches the “back” of the virtual video player).

[0259] In some embodiments, in response to detecting the second input, the electronic device 101a refrains (1026c) from interacting with the first user interface element (e.g., 903) in accordance with the second input. For example, if the hand (e.g., 909) in FIG. 9B approaches the user interface element (e.g., 903) from a side of the user interface element (e.g., 903) opposite the side of the user interface element (e.g., 903) visible in FIG. 9B, the electronic device 101a refrains from performing the selection of the user interface element (e.g., 903) shown in FIG. 9B. In some embodiments, the electronic device interacts with the first user interface element if a predetermined portion of the user moves from within a predetermined angle range to a location corresponding to the first user interface element. For example, in response to detecting that a user's hand has touched a surface of the virtual video player on which content is presented by moving the hand through a surface of the virtual video player other than the surface of the virtual video player on which content is presented, the electronic device refrains from performing an action on the surface of the virtual video player on which content is presented that corresponds to the area of ​​the video player touched by the user.

[0260] The above-described method of forgoing interaction with a first user interface element in response to input provided outside a predetermined range of angles provides an efficient way of preventing accidental input caused by a user inadvertently touching the first user interface element from an angle outside the predetermined range of angles, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving the battery life of the electronic device while reducing errors in use.

[0261] In some embodiments, such as FIG. 9A , the first action is performed 1028a in response to detecting a first input without detecting that the user's gaze (e.g., 901a) is directed toward the first user interface element (e.g., 903). In some embodiments, the attention zone includes an additional region of the three-dimensional environment within a predetermined distance or angle of the user's gaze in addition to the region of the three-dimensional environment toward which the user's gaze is directed. In some embodiments, the electronic device performs an action in response to an input directed toward the first user interface element while the first user interface element is within the attention zone (which is wider than the user's gaze), even if the user's gaze is not directed toward the first user interface element and even if the user's gaze was not directed toward the first user interface element during the user input. In some embodiments, indirect input requires the user's gaze to be directed toward the user interface element toward which the input is directed, and direct input does not require the user's gaze to be directed toward the user interface element toward which the input is directed.

[0262] The above-described method of performing an action in response to input directed at a first user interface element while the user's gaze is not directed at the first user interface element provides an efficient way of allowing a user to view areas of the user interface other than the first user interface element while providing input directed at the first user interface element, which simplifies the interaction between the user and the electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving battery life of the electronic device while reducing errors in use.

[0263] 11A-11C show examples of how an electronic device can enhance interaction with user interface elements at different distances and / or angles relative to a user's line of sight in a three-dimensional environment, according to some embodiments.

[0264] FIG. 11A illustrates electronic device 101 displaying a three-dimensional environment 1101 on a user interface via display generation component 120. It should be understood that in some embodiments, electronic device 101 utilizes one or more of the techniques described with reference to FIGS. 11A-11C in a two-dimensional environment or user interface without departing from the scope of this disclosure. As described above with reference to FIGS. 1-6 , electronic device 101 optionally includes display generation component 120 (e.g., a touchscreen) and multiple image sensors 314. The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor that electronic device 101 can use to capture one or more images of a user or a portion of a user while the user interacts with electronic device 101. In some embodiments, display generation component 120 is a touchscreen capable of detecting a user's hand gestures and movements. In some embodiments, the user interfaces described below may also be implemented in a head-mounted display that includes a display generation component that displays the user interface to the user and sensors that detect the physical environment and / or the movement of the user's hands (e.g., external sensors facing outward from the user) and / or the user's line of sight (e.g., internal sensors facing inward toward the user's face).

[0265] 11A , the three-dimensional environment 1101 includes two user interface objects 1103 a and 1103 b disposed within a region of the three-dimensional environment 1101 that is a first distance from a viewpoint of the three-dimensional environment 1101 associated with a user of the electronic device 101, two user interface objects 1105 a and 1105 b disposed within a region of the three-dimensional environment 1101 that is a second distance greater than the first distance from a viewpoint of the three-dimensional environment 1101 associated with a user of the electronic device 101, two user interface objects 1107 a and 1107 b disposed within a region of the three-dimensional environment 1101 that is a third distance greater than the second distance from a viewpoint of the three-dimensional environment 1101 associated with a user of the electronic device 101, and user interface object 1109. In some embodiments, the three-dimensional environment includes a representation 604 of a table within the physical environment of the electronic device 101 (e.g., as described with reference to FIG. 6B ). In some embodiments, representation 604 of the table is a photorealistic video image (e.g., a video or digital pass-through) of the table displayed by display generation component 120. In some embodiments, representation 604 of the table is a view of the table through a transparent portion of display generation component 120 (e.g., a true pass-through or a physical pass-through).

[0266] 11A-11C illustrate simultaneous or alternative inputs provided by a user's hand based on simultaneous or alternative locations of the user's gaze in the three-dimensional environment. In particular, in some embodiments, electronic device 101 directs indirect input from a user's hand of electronic device 101 (e.g., as described with reference to method 800) to different user interface objects depending on the distance of the user interface object from a viewpoint of the three-dimensional environment associated with the user. For example, in some embodiments, when indirect input from the user's hand is directed to a user interface object relatively close to the user's viewpoint in three-dimensional environment 1101, electronic device 101 optionally directs the detected indirect input to the user interface object toward which the user's gaze is directed. Because at a relatively close distance, device 101 can optionally determine relatively accurately which of two (or more) user interface objects the user's gaze is directed, which is optionally used to determine the user interface object toward which the indirect input should be directed.

[0267] In FIG. 11A , user interface objects 1103a and 1103b are relatively close to a user's viewpoint within three-dimensional environment 1101 (e.g., less than a first threshold distance of 1, 2, 5, 10, 20, 50 feet, etc. from the viewpoint) (e.g., objects 1103a and 1103b are located within a region of three-dimensional environment 1101 that is relatively close to the user's viewpoint). Thus, indirect input provided by hand 1113a detected by device 101 is directed toward user interface object 1103a (e.g., rather than user interface object 1103b), as indicated by the check mark in the figure. This is because user's gaze 1111a is directed toward user interface object 1103a when indirect input provided by hand 1113a is detected. In contrast, in FIG. 11B , user's gaze 1111d is directed toward user interface object 1103b when indirect input provided by hand 1113a is detected. Thus, device 101 directs indirect input from hand 1113a to user interface object 1103b (eg, rather than user interface object 1103a), as indicated by the check mark in the figure.

[0268] In some embodiments, if one or more user interface objects are relatively far from a user's viewpoint within three-dimensional environment 1101, device 101 optionally prevents indirect input from being directed at such one or more user interface objects and / or makes such one or more user interface objects visually less prominent, because at relatively far distances, device 101 optionally cannot relatively accurately determine whether the user's gaze is directed at one or more user interface objects. For example, in FIG. 11A , user interface objects 1107a and 1107b are relatively far from a user's viewpoint within three-dimensional environment 1101 (e.g., greater than a first threshold distance, greater than a second threshold distance from the viewpoint, such as 10, 20, 30, 50, 100, 200 feet, etc.) (e.g., objects 1107a and 1107b are located in an area of ​​three-dimensional environment 1101 that is relatively far from the user's viewpoint). Thus, indirect input provided by hand 1113c that is detected by device 101 while user's gaze 1111c is (e.g., ostensibly) directed at user interface object 1107b (or 1107a) is ignored by device 101 and is not directed at user interface object 1107b (or 1107a), as reflected by the absence of a check mark shown in the figure. In some embodiments, device 101 additionally or alternatively visually obscures (e.g., grays out) user interface objects 1107a and 1107b to indicate that user interface objects 1107a and 1107b are not available for indirect interaction.

[0269] In some embodiments, if one or more user interface objects are greater than a threshold angle from the line of sight of a user of electronic device 101, device 101 optionally prevents indirect input from being directed at such one or more user interface objects and / or visually obscures such one or more user interface objects, e.g., to prevent accidental interaction with one or more user interface objects that are out of such angle. For example, in FIG. 11A , user interface object 1109 is optionally greater than a threshold angle (e.g., 10, 20, 30, 45, 90, 120 degrees, etc.) from user's line of sight 1111a, 1111b, and / or 1111c. Thus, device 101 optionally visually obscures (e.g., grays out) user interface object 1109 to indicate that user interface object 1109 is not available for indirect interaction.

[0270] However, in some embodiments, when indirect input from the user's hand is directed toward a user interface object that is reasonably far from the user's viewpoint in three-dimensional environment 1101, electronic device 101 optionally directs the detected indirect input toward the user interface object based on criteria other than the user's line of sight. This is because, at a reasonable distance, device 101 can optionally accurately determine that the user's line of sight is directed toward a set of two or more user interface objects, but optionally cannot relatively accurately determine which of the set of two or more user interface objects the gaze is directed toward. In some embodiments, when the user's line of sight is directed toward a reasonably distant user interface object that is not co-located with other user interface objects (e.g., beyond a threshold distance of 1, 2, 5, 10, 20 feet, etc. from any other interactable user interface object), device 101 optionally directs the indirect input toward that user interface object without performing the various disambiguation techniques described herein with reference to method 1200. Additionally, in some embodiments, electronic device 101 performs the various disambiguation techniques described herein with reference to method 1200 on user interface objects located within a region (e.g., a volume and / or surface or plane) within the three-dimensional environment defined by the user's line of sight (e.g., the user's line of sight defines the center of that volume and / or surface or plane), but not on user interface objects not located within the region (e.g., regardless of distance from the user's viewpoint). In some embodiments, the size of the region varies based on the distance of the region and / or the user interface objects it contains from the user's viewpoint within the three-dimensional environment (e.g., within a reasonably distant region of the three-dimensional environment). For example, in some embodiments, the size of the region decreases as the region becomes farther from the viewpoint (and increases as the region becomes closer to the viewpoint), and in some embodiments, the size of the region increases as the region becomes farther from the viewpoint (and decreases as the region becomes closer to the viewpoint).

[0271] 11A , user interface objects 1105a and 1105b are moderately far from a user's viewpoint within three-dimensional environment 1101 (e.g., greater than a first threshold distance and less than a second threshold distance from the viewpoint) (e.g., objects 1105a and 1105b are located within a region of three-dimensional environment 1101 that is moderately far from the user's viewpoint). In FIG. 11A , when device 101 detects indirect input from hand 1113b, line of sight 1111b is directed toward (e.g., device 101 detects) user interface object 1105a. Because user interface objects 1105a and 1105b are moderately far from the user's viewpoint, device 101 determines which of user interface objects 1105a and 1105b will accept the input based on characteristics other than user's line of sight 1111b. 11A, device 101 directs input from hand 1113b toward user interface object 1105b, as indicated by the check mark in the figure (and not toward user interface object 1105a, at which, for example, user's gaze 1111b is directed), because user interface object 1105b is closer to the user's viewpoint in three-dimensional environment 1101. In Figure 11B, user's gaze 1111e is directed toward user interface object 1105b (rather than user interface object 1105a in Figure 11A) when input from hand 1113b is detected, and device 101 optionally still directs indirect input from hand 1113b toward user interface object 1105b, as indicated by the check mark in the figure, not because user's gaze 1111e is directed toward user interface object 1105b, but because user interface object 1105b is closer to the user's viewpoint in the three-dimensional environment than user interface object 1105a.

[0272] In some embodiments, criteria in addition to or instead of distance are used to determine which user interface object to direct the indirect input to (e.g., if those user interface objects are reasonably far from the user's viewpoint). For example, in some embodiments, device 101 directs the indirect input to one of the user interface objects based on which of the user interface objects is an application user interface object or a system user interface object. For example, in some embodiments, device 101 prefers system user interface objects and directs the indirect input from hand 1113b in FIG. 11C to user interface object 1105c, as indicated by the check mark, because it is a system user interface object, and user interface object 1105d (to which user's gaze 1111f is directed) is an application user interface object. In some embodiments, device 101 prefers application user interface objects and directs the indirect input from hand 1113b in FIG. 11C to user interface object 1105d. 1105c is not the application user interface object because it is an application user interface object and user interface object 1105c is a system user interface object (eg, not because the user's gaze 1111f is directed at user interface object 1105d).Additionally or alternatively, in some embodiments, the software, application(s) and / or operating system associated with the user interface objects defines a selection priority for the user interface objects, such that if the selection priority gives one user interface object a higher priority than another user interface object, device 101 directs input to that one user interface object (e.g., user interface object 1105c), and if the selection priority gives the other user interface object a higher priority than one user interface object, device 101 directs input to the other user interface object (e.g., user interface object 1105d).

[0273] 12A-12F are flowcharts illustrating a method 1200 for enhancing interaction with user interface elements at different distances and / or angles relative to a user's line of sight in a three-dimensional environment, according to some embodiments. In some embodiments, method 1200 is performed on a computer system (e.g., computer system 101 of FIG. 1 , such as a tablet, smartphone, wearable computer, or head-mounted device) that includes a display generation component (e.g., display generation component 120 of FIGS. 1, 3, and 4 ) (e.g., a head-up display, a display, a touchscreen, a projector, etc.) and one or more cameras (e.g., a camera pointing downward in a user's hand (e.g., color sensors, infrared sensors, and other depth-sensing cameras) or a camera pointing forward from the user's head). In some embodiments, method 1200 is governed by instructions stored on a non-transitory computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A ). Some operations of method 1200 are optionally combined and / or the order of some operations is optionally changed.

[0274] In some embodiments, method 1200 is performed by an electronic device in communication with a display generation component and one or more input devices, including an eye-tracking device. For example, the electronic device may be a mobile device (e.g., a tablet, smartphone, media player, or wearable device) or a computer. In some embodiments, the display generation component may be a display integrated with the electronic device (optionally a touchscreen display), an external display such as a monitor, projector, television, or a hardware component (optionally built-in or external) for projecting a user interface and making the user interface visible to one or more users. In some embodiments, the one or more input devices include electronic devices or components capable of receiving user input (e.g., capturing user input, detecting user input, etc.) and transmitting information related to the user input to the electronic device. Examples of input devices include a touchscreen, a mouse (e.g., external), a trackpad (optionally integrated or external), a touchpad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the electronic device), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, an eye tracking device, and / or a motion sensor (e.g., a hand tracking device, hand motion sensor), etc. In some embodiments, the hand tracking device is a wearable device such as a smart glove. In some embodiments, the hand tracking device is a handheld input device such as a remote control or a stylus.

[0275] In some embodiments, the electronic device, via a display generation component, displays (1202a) a user interface including a first region including a first user interface object and a second user interface object, such as objects 1105a and 1105b in FIG. 11A . In some embodiments, the first and / or second user interface objects are interactive user interface objects, and in response to detecting input directed at the given object, the electronic device performs an action associated with the user interface object. For example, the user interface object is a selectable option that, when selected, causes the electronic device to perform an action such as displaying a separate user interface, changing a setting on the electronic device, or starting playback of content. As another example, the user interface object is a container (e.g., a window) in which a user interface / content is displayed, and in response to detecting selection of the user interface object followed by movement input, the electronic device updates the position of the user interface object according to the movement input. In some embodiments, the first user interface object and the second user interface object are displayed in (e.g., the user interface is and / or is displayed within) a three-dimensional environment generated, displayed, or otherwise viewable by the device (e.g., a computer-generated reality (CGR) environment such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment. In some embodiments, the first region, and therefore the first and second user interface objects, are away from a location corresponding to the location of the user / electronic device in the three-dimensional environment and / or away from the user's viewpoint in the three-dimensional environment (e.g., 2, 5, 10, 15, 20 feet away from a threshold distance, etc.).

[0276] In some embodiments, while displaying a user interface and detecting via an eye tracking device a user's gaze directed toward a first region of the user interface (e.g., the user's gaze intersects with the first region), such as gaze 1111b in FIG. 11A, the first user interface object and / or the second user interface object, or the user's gaze, is within a threshold distance, such as 1, 2, 5, 10 feet, of the intersection with the first region, the first user interface object, and / or the second user interface object. In some embodiments, the first region, the first user interface object, and / or the second user interface object are sufficiently far from the position of the user / electronic device such that the electronic device is unable to determine whether the user's gaze is directed toward the first or second user interface object and / or can only determine that the user's gaze is directed toward the first region of the user interface, and the electronic device detects 1202b a distinct input provided by a predetermined part of the user, such as input from hand 1113b of FIG. 11A via one or more input devices (e.g., a gesture in which a finger, such as the index finger, of the user's hand points toward and / or moves toward the first region, optionally with a threshold movement (e.g., 0.5, 1, 3, 5, 10 cm) and / or a threshold velocity (e.g., 0.5, 1, 3, 5, 10 cm / sec), or a gesture in which the thumb of the hand is pinched together with another finger of the hand). In some embodiments, during the discrete input, the location of the pre-defined portion of the user is away from a location corresponding to the first region of the user interface (e.g., the pre-defined portion of the user remains more than a threshold distance of 2, 5, 10, 15, 20 feet away from the first region, the first user interface object, and / or the second user interface object throughout the discrete input). The discrete input is, optionally, input provided by interaction with the pre-defined portion of the user and / or the user interface object, as described with reference to methods 800, 1000, 1600, 1800, and / or 2000.

[0277] In some embodiments, in response to detecting the distinct input (1202c), in accordance with a determination that one or more first criteria are met (e.g., the first user interface object is closer to the user's viewpoint in the three-dimensional environment than the second user interface object), the first user interface object is a system user interface object (e.g., a user interface object of an operating system of the electronic device as opposed to a user interface object of an application on the electronic device) and the second user interface object is an application user interface object (e.g., a user interface object of an application on the electronic device as opposed to a user interface object of an operating system of the electronic device). In some embodiments, the one or more first criteria are not met based on the user's gaze (e.g., whether the one or more first criteria are met is independent of where the user's gaze is directed in the first region of the user interface), and the electronic device performs an action on the first user interface object based on the distinct input (1202d), such as for user interface object 1105b of FIG. 11A (e.g., without performing an action based on the distinct input on the second user interface object). For example, selecting a first user interface object for further interaction (e.g., without transitioning a second user interface object to a selected state), transitioning a first user interface object to a selected state so that further input interacts with the first user interface object (e.g., without selecting a second user interface object for further interaction), selecting a first user interface object as a button (e.g., without selecting a second user interface object as a button), etc.

[0278] In some embodiments, in accordance with a determination that one or more second criteria different from the first criteria are met (e.g., the second user interface object is closer to the user's viewpoint in the three-dimensional environment than the first user interface object), the second user interface object is a system user interface object (e.g., a user interface object of an operating system of the electronic device as opposed to a user interface object of an application on the electronic device) and the first user interface object is an application user interface object (e.g., a user interface object of an application on the electronic device as opposed to a user interface object of an operating system of the electronic device). In some embodiments, the one or more second criteria are not met based on the user's gaze (e.g., whether the one or more second criteria are met is independent of where the user's gaze is directed in the first region of the user interface), and the electronic device performs an action on the second user interface object based on the separate input (1202e), such as with user interface object 1105c of FIG. 11C (e.g., without performing an action based on the separate input on the first user interface object). For example, selecting a second user interface object for further interaction (e.g., without selecting the first user interface object for further interaction), transitioning the second user interface object to a selected state so that further input interacts with the second user interface object (e.g., without transitioning the first user interface object to a selected state), selecting the second user interface object as a button (e.g., without selecting the first user interface object as a button), etc.The above-described method of clarifying which user interface object a particular input is directed to provides an efficient way of facilitating interaction with user interface objects when there may be uncertainty about which user interface object a given input is directed to, without requiring further user input to designate the given user interface object as the target of the given input, which simplifies the interaction between a user and an electronic device, improves usability of the electronic device, and makes the user device interface more efficient (e.g., by not requiring additional user input for further designation), which further reduces power usage and improves battery life of the electronic device by allowing a user to use the electronic device more quickly and efficiently.

[0279] In some embodiments, the user interface includes a three-dimensional environment (1204a), such as environment 1101 (e.g., the first region is a distinct volume and / or surface located at some x, y, z coordinate within the three-dimensional environment at which a viewpoint of the three-dimensional environment associated with the electronic device is located). In some embodiments, the first and second user interface objects are located within distinct volumes and / or surfaces, and the first region is a distinct distance from a viewpoint associated with the electronic device within the three-dimensional environment (1204b) (e.g., the first region is at a location within the three-dimensional environment that is some distance, angle, position, etc. relative to the location of the viewpoint within the three-dimensional environment). In some embodiments, in accordance with a determination that the distinct distance is a first distance (e.g., 1 foot, 2 feet, 5 feet, 10 feet, 50 feet), the first region has a first size in the three-dimensional environment (1204c), and in accordance with a determination that the distinct distance is a second distance different from the first distance (e.g., 10 feet, 20 feet, 50 feet, 100 feet, 500 feet), the first region has a second size in the three-dimensional environment different from the first size (1204d). For example, the size of a region at which the electronic device initiates actions on first and second user interface objects within the region based on one or more first or second criteria (e.g., not based on the user's gaze being directed at the first or second user interface object) varies based on the distance of the region from a viewpoint associated with the electronic device. In some embodiments, the size of the region decreases as the region of interest becomes farther from the viewpoint; in some embodiments, the size of the region increases as the region of interest becomes farther from the viewpoint.For example, in FIG. 11A, if objects 1105a and 1105 are further away from the user's viewpoint than shown in FIG. 11A, the region that includes objects 1105a and 1105b and over which the criteria-based disambiguation described herein is performed will be different (e.g., larger); if objects 1105a and 1105 are closer to the user's viewpoint than shown in FIG. 11A, the region that includes objects 1105a and 1105b and over which the criteria-based disambiguation described herein is performed will be different (e.g., smaller). The above-described method of operating on regions of different sizes depending on the distance of the regions from a viewpoint associated with the electronic device provides an efficient way of ensuring that the device's operation with respect to potential uncertainties in the input accurately corresponds to potential uncertainties in the input without requiring further user input to manually change the size of the region of interest, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user device interface more efficient, which further allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage and improving the battery life of the electronic device while reducing device malfunctions.

[0280] 11A-11C , the size of the first region in the three-dimensional environment increases (1206a) as the respective distances increase. For example, as the region of interest becomes farther from the viewpoint associated with the electronic device, the size of the region at which the electronic device initiates action on first and second user interface objects within the region based on one or more first or second criteria (e.g., not based on the user's gaze being directed at the first or second user interface objects) increases, optionally corresponding to uncertainty in determining what the user's gaze is directed at because potentially related user interface objects are farther away from the viewpoint associated with the electronic device (e.g., two more user interface objects, from the viewpoint, may make it more difficult to determine whether the user's gaze is directed at the first or second of the two user interface objects, and therefore the electronic device optionally acts based on one or more first or second criteria related to those two user interface objects). The above-described method of operation for regions that increase in size as the regions become farther from a viewpoint associated with the electronic device provides an efficient way of avoiding erroneous responses of the device to gaze-based inputs directed at objects as those objects become farther from a viewpoint associated with the electronic device, which simplifies the interaction between the user and the electronic device, improves the usability of the electronic device, and makes the user-device interface more efficient, which further reduces power usage and improves the battery life of the electronic device by allowing the user to use the electronic device more quickly and efficiently, while reducing malfunctions of the device.

[0281] In some embodiments, the one or more first criteria are satisfied when the first object is closer to a user's viewpoint in the three-dimensional environment than the second object, such as user interface object 1105b in Figure 11A, and the one or more second criteria are satisfied when the second object is closer to a user's viewpoint in the three-dimensional environment than the first object (1208a), such as when user interface object 1105a is closer than user interface object 1105b in Figure 11A. For example, in accordance with a determination that the first user interface object is closer to a viewpoint associated with the electronic device in the three-dimensional environment than the second user interface object, the one or more first criteria are satisfied and the one or more second criteria are not satisfied, and in accordance with a determination that the second user interface object is closer to the viewpoint in the three-dimensional environment than the first user interface object, the one or more second criteria are satisfied and the one or more first criteria are not satisfied. Thus, in some embodiments, whichever user interface object within the first region is closest to the viewpoint is the user interface object to which the device directs input (e.g., regardless of whether the user's gaze is directed at another user interface object within the first region). The above-described method of directing input to user interface objects based on their distance from a viewpoint associated with the electronic device provides an efficient and predictable way of selecting user interface objects for input, which simplifies the interaction between the user and the electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further allows the user to use the electronic device more quickly and efficiently, thereby reducing power usage, improving the battery life of the electronic device, and reducing device malfunctions.

[0282] In some embodiments, one or more first criteria are met or one or more second criteria are met based on the type of the first user interface object (e.g., a user interface object of the operating system of the electronic device, or a user interface object of an application other than the operating system of the electronic device) and the type of the second user interface object (e.g., a user interface object of the operating system of the electronic device, or a user interface object of an application other than the operating system of the electronic device) (1210a). For example, in accordance with a determination that the first user interface object is a system user interface object and the second user interface object is not a system user interface object (e.g., an application user interface object), one or more first criteria are met and one or more second criteria are not met; in accordance with a determination that the second user interface object is a system user interface object and the first user interface object is not a system user interface object (e.g., an application user interface object), one or more second criteria are met and one or more first criteria are not met. Thus, in some embodiments, whichever user interface object in the first region is the system user interface object is the user interface object to which the device directs input (e.g., regardless of whether the user's gaze is directed at another user interface object in the first region).11A, if user interface object 1105b was a system user interface object and user interface object 1105a was an application user interface object, device 101 could direct the input of FIG. 11A to object 1105b instead of object 1105a (e.g., even though object 1105b is farther from the user's viewpoint than object 1105a). The above-described method of directing input to user interface objects based on their type provides an efficient and predictable way of selecting user interface objects for input, which simplifies the interaction between a user and an electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further allows a user to use the electronic device more quickly and efficiently, thereby reducing power usage, improving the battery life of the electronic device, and reducing device malfunctions.

[0283] In some embodiments, the one or more first criteria are met or the one or more second criteria are met based on respective priorities defined for the first and second user interface objects by the electronic device (1212a) (e.g., by software on the electronic device, such as an application or an operating system of the electronic device). For example, in some embodiments, the application(s) and / or operating system associated with the first and second user interface objects define selection priorities for the first and second user interface objects such that if the selection priority gives the first user interface object a higher priority than the second user interface object, the device directs input to the first user interface object (e.g., regardless of whether the user's gaze is directed at another user interface object in the first region), and if the selection priority gives the second user interface object a higher priority than the first user interface object, the device directs input to the second user interface object (e.g., regardless of whether the user's gaze is directed at another user interface object in the first region). For example, in FIG. 11A, if user interface object 1105b is assigned a higher selection priority (e.g., by software in device 101) and user interface object 1105a is assigned a lower selection priority, device 101 may direct the input in FIG. 11A to object 1105b instead of object 1105a (e.g., even though object 1105b is farther from the user's viewpoint than object 1105a).In some embodiments, the relative selection priority of the first and second user interface objects changes over time based on what the respective user interface objects are currently displaying (e.g., a user interface object currently displaying video / playing content has a higher selection priority than the same user interface object displaying paused video content or other content other than the video / playing content). The above-described method of directing input to user interface objects based on operating system and / or application priority provides a flexible way of selecting user interface objects for input, which simplifies the interaction between a user and an electronic device, improves usability of the electronic device, and makes the user device interface more efficient, which further reduces power usage and improves battery life of the electronic device by allowing a user to use the electronic device more quickly and efficiently.

[0284] In some embodiments, in response to detecting the distinct input (1214a), in accordance with a determination that one or more third criteria are met, including criteria that are met when the first region is more than a threshold distance (e.g., 5, 10, 15, 20, 30, 40, 50, 100, 150 feet) in the three-dimensional environment from a viewpoint associated with the electronic device, the electronic device refrains from performing (1214b) the action on the first user interface object and refrains from performing the action on the second user interface object, as described with reference to user interface objects 1107a and 1107b in FIG. 11A . For example, the electronic device optionally disables interaction with user interface objects that are within a region that ...

Claims

1. 1. A method comprising: An electronic device in communication with a display generation component and one or more input devices, displaying, via the display generation component, a user interface including a user interface element; detecting input from a predetermined portion of a user of the electronic device via the one or more input devices while displaying the user interface element; in response to detecting the input from the predetermined portion of the user of the electronic device; performing a discrete action in accordance with the input from the predetermined portion of the user of the electronic device in accordance with a determination that a posture of the predetermined portion of the user prior to detecting the input satisfies one or more criteria; and refraining from performing the individual action in accordance with the input from the predetermined portion of the user of the electronic device in accordance with a determination that the posture of the predetermined portion of the user prior to detecting the input does not satisfy the one or more criteria.

2. displaying the user interface element with visual characteristics having a first value and displaying a second user interface element included in the user interface with the visual characteristics having a second value while the posture of the predetermined portion of the user does not satisfy the one or more criteria; updating the visual characteristics of a user interface element to which input focus is directed while the pose of the predetermined portion of the user satisfies the one or more criteria, wherein the updating comprises: updating the user interface element to be displayed with the visual characteristic having a third value in accordance with a determination that input focus is directed to the user interface element; and updating the second user interface element to be displayed with the visual characteristic having a fourth value in accordance with determining that the input focus is directed toward the second user interface element.

3. the input focus is directed to the user interface element in accordance with a determination that the predetermined portion of the user is within a threshold distance of a location corresponding to the user interface element; The method of claim 2 , wherein the input focus is directed to the second user interface element in accordance with a determination that the predetermined portion of the user is within the threshold distance of the second user interface element.

4. the input focus is directed to the user interface element in accordance with a determination that the user's gaze is directed to the user interface element; The method of claim 2 or 3, wherein the input focus is directed to the second user interface element in accordance with a determination that the user's gaze is directed to the second user interface element.

5. Updating the visual characteristics of a user interface element to which input focus is directed includes: updating the visual characteristics of the user interface element to which the input focus is directed in accordance with a determination that the pose of the predetermined portion of the user satisfies a first set of one or more criteria in accordance with a determination that the predetermined portion of the user is less than a threshold distance from a location corresponding to the user interface element; 4. The method of claim 2 or 3, further comprising: updating the visual characteristics of the user interface element to which the input focus is directed in accordance with a determination that the pose of the predetermined portion of the user satisfies a second set of one or more criteria that is different from the first set of one or more criteria in accordance with a determination that the predetermined portion of the user exceeds the threshold distance from the location corresponding to the user interface element.

6. The posture of the predetermined portion of the user satisfying the one or more criteria may be In accordance with a determination that the predetermined portion of the user is less than a threshold distance from a location corresponding to the user interface element, the pose of the predetermined portion of the user satisfies a first set of one or more criteria; and wherein, in accordance with a determination that the predetermined portion of the user exceeds the threshold distance from the location corresponding to the user interface element, the pose of the predetermined portion of the user satisfies a second set of one or more criteria that is different from the first set of one or more criteria.

7. The posture of the predetermined portion of the user satisfying the one or more criteria may be In accordance with a determination that the predetermined part of the user is holding an input device of the one or more input devices, the pose of the predetermined part of the user satisfies a first set of one or more criteria; and wherein, in accordance with a determination that the predetermined part of the user is not holding the input device, the posture of the predetermined part of the user satisfies a second set of one or more criteria.

8. The posture of the predetermined portion of the user satisfying the one or more criteria may be In accordance with a determination that the predetermined portion of the user is less than a threshold distance from a location corresponding to the user interface element, the pose of the predetermined portion of the user satisfies a first set of one or more criteria; and wherein the pose of the predetermined portion of the user satisfies the first set of one or more criteria in accordance with a determination that the predetermined portion of the user exceeds the threshold distance from the location corresponding to the user interface element.

9. pursuant to a determination that the predetermined portion of the user is more than a threshold distance away from a location corresponding to the user interface element during the individual input, the one or more criteria including a criterion that is satisfied when the user's attention is directed toward the user interface element; 9. The method of claim 1, wherein, pursuant to a determination that the predetermined portion of the user is less than the threshold distance from the location corresponding to the user interface element during the individual input, the one or more criteria do not include a requirement that the attention of the user be directed toward the user interface element in order to satisfy the one or more criteria.

10. In response to detecting that the user's gaze is directed toward a first region of the user interface, visually obscuring, via the display generation component, a second region of the user interface relative to the first region of the user interface; 10. The method of claim 1, further comprising: in response to detecting that the user's gaze is directed toward the second region of the user interface, visually obscuring the first region of the user interface relative to the second region of the user interface via the display generation component.

11. The user interface is accessible by the electronic device and a second electronic device, and the method includes: and refraining, via the display generation component, from visually obscuring the second region of the user interface relative to the first region of the user interface in accordance with an indication that a gaze of a second user of the second electronic device is directed toward the first region of the user interface; 11. The method of claim 10, further comprising: in accordance with an indication that the gaze of the second user of the second electronic device is directed toward the second region of the user interface, via the display generation component, forgoing visual obscuring of the first region of the user interface relative to the second region of the user interface.

12. 12. The method of claim 1, wherein detecting the input from the predetermined portion of the user of the electronic device comprises detecting a pinch gesture performed by the predetermined portion of the user via a hand tracking device.

13. 13. The method of claim 1, wherein detecting the input from the predetermined portion of the user of the electronic device comprises detecting a press gesture performed by the predetermined portion of the user via a hand tracking device.

14. 14. The method of claim 1, wherein detecting the input from the predetermined portion of the user of the electronic device comprises detecting lateral movement of the predetermined portion of the user relative to a location corresponding to the user interface element.

15. before determining that the posture of the predetermined portion of the user prior to detecting the input satisfies the one or more criteria; Detecting via an eye-tracking device that the user's gaze is directed toward the user interface element; 15. The method of claim 1, further comprising: in response to detecting that the user's gaze is directed toward the user interface element, displaying, via the display generation component, a first indication that the user's gaze is directed toward the user interface element.

16. prior to detecting the input from the predetermined portion of the user of the electronic device, while the posture of the predetermined portion of the user prior to detecting the input satisfies the one or more criteria.

16. The method of claim 15, further comprising displaying, via the indication generation component, a second indication that the posture of the predetermined portion of the user prior to detecting the input satisfies the one or more criteria, wherein the first indication is different from the second indication.

17. detecting a second input from a second predetermined portion of the user of the electronic device via the one or more input devices while displaying the user interface element; in response to detecting the second input from the second predetermined portion of the user of the electronic device; performing a second distinct action in accordance with the second input from the second predetermined portion of the user of the electronic device in accordance with a determination that a posture of the second predetermined portion of the user prior to detecting the second input satisfies one or more second criteria; 17. The method of claim 1, further comprising: refraining from performing the second individual action in accordance with the second input from the second predetermined portion of the user of the electronic device in accordance with a determination that the posture of the second predetermined portion of the user prior to detecting the second input does not satisfy the one or more second criteria.

18. The user interface is accessible by the electronic device and a second electronic device, and the method includes: displaying the user interface element with visual characteristics having a first value before detecting that the pose of the predetermined portion of the user before detecting the input satisfies the one or more criteria; displaying the user interface element with the visual characteristic having a second value different from the first value while the pose of the predetermined portion of the user prior to detecting the input satisfies the one or more criteria; 18. The method of claim 1, further comprising: while displaying the user interface element with the visual characteristic having the first value, maintaining the display of the user interface element with the visual characteristic having the first value while a posture of a predetermined portion of a second user of the second electronic device satisfies the one or more criteria.

19. displaying the user interface element with the visual characteristic having a third value in response to detecting the input from the predetermined portion of the user of the electronic device; 20. The method of claim 18, further comprising: displaying the user interface element with the visual characteristic having the third value in response to an indication of input from the predetermined portion of the second user of the second electronic device.

20. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising: displaying, via a display generation component, a user interface including the user interface element; detecting input from a predetermined portion of a user of the electronic device via one or more input devices while displaying the user interface element; in response to detecting the input from the predetermined portion of the user of the electronic device; performing a distinct action in accordance with the input from the predetermined portion of the user of the electronic device in accordance with a determination that a posture of the predetermined portion of the user prior to detecting the input satisfies one or more criteria; an electronic device including instructions for refraining from performing the individual action in accordance with the input from the predetermined portion of the user of the electronic device in accordance with a determination that the posture of the predetermined portion of the user prior to detecting the input does not satisfy the one or more criteria.

21. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the electronic device, cause the electronic device to: Displaying, via a display generation component, a user interface including the user interface element; detecting input from a predetermined portion of a user of the electronic device via one or more input devices while displaying the user interface element; in response to detecting the input from the predetermined portion of the user of the electronic device; performing a discrete action in accordance with the input from the predetermined portion of the user of the electronic device in accordance with a determination that a posture of the predetermined portion of the user prior to detecting the input satisfies one or more criteria; and refraining from performing the individual action in accordance with the input from the predetermined portion of the user of the electronic device in accordance with a determination that the posture of the predetermined portion of the user prior to detecting the input does not satisfy the one or more criteria.

22. 1. An electronic device comprising: one or more processors; Memory and means for displaying, via a display generation component, a user interface including the user interface element; means for detecting input from a predetermined portion of a user of the electronic device via one or more input devices while displaying the user interface elements; in response to detecting the input from the predetermined portion of the user of the electronic device; performing a distinct action in accordance with the input from the predetermined portion of the user of the electronic device in accordance with a determination that a posture of the predetermined portion of the user prior to detecting the input satisfies one or more criteria; means for forgoing performing the individual action in accordance with the input from the predetermined part of the user of the electronic device in accordance with a determination that the posture of the predetermined part of the user prior to detecting the input does not satisfy the one or more criteria.

23. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: means for displaying, via the display generation component, a user interface including user interface elements; means for detecting input from a predetermined portion of a user of the electronic device via the one or more input devices while displaying the user interface elements; in response to detecting the input from the predetermined portion of the user of the electronic device; performing a distinct action in accordance with the input from the predetermined portion of the user of the electronic device in accordance with a determination that a posture of the predetermined portion of the user prior to detecting the input satisfies one or more criteria; and means for refraining from performing the individual action in accordance with the input from the predetermined part of the user of the electronic device in accordance with a determination that the posture of the predetermined part of the user before detecting the input does not satisfy the one or more criteria.

24. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions to perform the method of any one of claims 1 to 19.

25. 20. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 1 to 19.

26. 1. An electronic device comprising: one or more processors; Memory and and means for performing the method of any one of claims 1 to 19.

27. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: and means for executing the method according to any one of claims 1 to 19.

28. 1. A method comprising: An electronic device in communication with a display generation component and one or more input devices, displaying, via the display generation component, a first user interface element; detecting a first input directed at the first user interface element via the one or more input devices while displaying the first user interface element; in response to detecting the first input directed at the first user interface element; performing a first action corresponding to the first user interface element in accordance with determining that the first user interface element is within a zone of attention associated with a user of the electronic device; and forgoing performing the first action in accordance with a determination that the first user interface element is not within the zone of attention associated with the user.

29. the first input directed to the first user interface element is an indirect input directed to the first user interface element, and the method further comprises: detecting, while displaying the first user interface element, a second input via the one or more input devices, the second input corresponding to a direct input directed at a respective user interface element; 29. The method of claim 28, further comprising: in response to detecting the second input, performing an action associated with the individual user interface element regardless of whether the individual user interface element is within the zone of attention associated with the user.

30. 30. The method of claim 28 or 29, wherein the zone of attention associated with the user is based on a direction of gaze of the user of the electronic device.

31. detecting, while the first user interface element is within the zone of attention associated with the user, that one or more criteria for moving the zone of attention to a location where the first user interface element is not within the zone of attention are met; After detecting that the one or more criteria are met, Detecting a second input directed at the first user interface element; in response to detecting the second input directed at the first user interface element; performing a second action corresponding to the first user interface element in accordance with a determination that the second input is detected within a respective time threshold during which the one or more criteria are satisfied; 31. The method of claim 28, further comprising: forgoing performing the second action in accordance with a determination that the second input is detected after the respective time threshold of the one or more criteria is satisfied.

32. The first input includes a first portion followed by a second portion, and the method comprises: While detecting the first input, detecting the first portion of the first input while the first user interface element is within the zone of attention; In response to detecting the first portion of the first input, performing a first portion of the first operation corresponding to the first user interface element; Detecting the second portion of the first input while the first user interface element is outside the zone of attention; 32. The method of claim 28, further comprising: in response to detecting the second portion of the first input, performing a second portion of the first operation corresponding to the first user interface element.

33. 33. The method of claim 32, wherein the first input corresponds to a press input, the first portion of the first input corresponds to a start of the press input, and the second portion of the first input corresponds to a continuation of the press input.

34. 33. The method of claim 32, wherein the first input corresponds to a drag input, the first portion of the first input corresponds to a start of the drag input, and the second portion of the first input corresponds to a continuation of the drag input.

35. 33. The method of claim 32, wherein the first input corresponds to a selection input, the first portion of the first input corresponds to a start of the selection input, and the second portion of the first input corresponds to a continuation of the selection input.

36. 36. The method of claim 32, wherein detecting the first portion of the first input comprises detecting a predetermined portion of the user having a distinct pose within a distinct distance of a location corresponding to the first user interface element without detecting movement of the predetermined portion of the user, and wherein detecting the second portion of the first input comprises detecting the movement of the predetermined portion of the user.

37. the first input is provided by a predetermined portion of the user, and detecting the first input includes detecting the predetermined portion of the user within a distance threshold of a location corresponding to the first user interface element, the method comprising: detecting, while detecting the first input directed at the first user interface element and before performing the first action, movement of the predetermined portion of the user via the one or more input devices to a distance that exceeds the distance threshold from the location corresponding to the first user interface element; 37. The method of claim 28, further comprising: in response to detecting the movement of the predetermined portion from the location corresponding to the first user interface element to the distance that exceeds the distance threshold, refraining from performing the first action corresponding to the first user interface element.

38. the first input is provided by a predetermined portion of the user, and detecting the first input includes detecting the predetermined portion of the user in a distinct spatial relationship to a location corresponding to the first user interface element, the method comprising: detecting, during the first input and prior to performing the first action, that the predetermined portion of the user did not engage with the first user interface element via the one or more input devices within a respective time threshold of entering the respective spatial relationship with the location corresponding to the first user interface element while the predetermined portion of the user was in the respective spatial relationship with the location corresponding to the first user interface element; 38. The method of claim 28, further comprising: refraining from performing the first action corresponding to the first user interface element in response to detecting that the predetermined portion of the user has not engaged with the first user interface element within the respective time threshold of entering the respective spatial relationship with the location corresponding to the first user interface element.

39. 39. The method of any one of claims 28 to 38, wherein a first portion of the first input is detected while the user's gaze is directed toward the first user interface element, and a second portion of the first input subsequent to the first portion of the first input is detected while the user's gaze is not directed toward the first user interface element.

40. The first input is provided by a predetermined portion of the user moving from within a predetermined range of angles relative to the first user interface element to a location corresponding to the first user interface element, and the method further comprises: detecting a second input directed at the first user interface element via the one or more input devices, the second input comprising the predetermined portion of the user moving from outside the predetermined angular range relative to the first user interface element to the location corresponding to the first user interface element; 40. The method of claim 28, further comprising: in response to detecting the second input, forgoing interacting with the first user interface element in accordance with the second input.

41. 41. The method of any one of claims 28 to 40, wherein the first action is performed in response to detecting the first input without detecting that the user's gaze is directed toward the first user interface element.

42. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising: displaying, via a display generation component, the first user interface element; detecting a first input directed at the first user interface element via one or more input devices while displaying the first user interface element; in response to detecting the first input directed at the first user interface element; performing a first action corresponding to the first user interface element in accordance with determining that the first user interface element is within a zone of attention associated with a user of the electronic device; an electronic device comprising instructions for forgoing performing the first action in accordance with a determination that the first user interface element is not within the zone of attention associated with the user;

43. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the electronic device, cause the electronic device to: displaying, via a display generation component, a first user interface element; detecting, while displaying the first user interface element, a first input directed at the first user interface element via one or more input devices; in response to detecting the first input directed at the first user interface element; performing a first action corresponding to the first user interface element in accordance with determining that the first user interface element is within a zone of attention associated with a user of the electronic device; and forgoing performing the first action in accordance with a determination that the first user interface element is not within the zone of attention associated with the user.

44. 1. An electronic device comprising: one or more processors; Memory and means for displaying a first user interface element via a display generation component; means for detecting a first input directed at the first user interface element via one or more input devices while displaying the first user interface element; in response to detecting the first input directed at the first user interface element; performing a first action corresponding to the first user interface element in accordance with determining that the first user interface element is within a zone of attention associated with a user of the electronic device; means for forgoing performing the first action in accordance with a determination that the first user interface element is not within the zone of attention associated with the user.

45. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: means for displaying a first user interface element via the display generation component; means for detecting a first input directed at the first user interface element via the one or more input devices while displaying the first user interface element; in response to detecting the first input directed at the first user interface element; performing a first action corresponding to the first user interface element in accordance with determining that the first user interface element is within a zone of attention associated with a user of the electronic device; means for forgoing performing the first action in accordance with a determination that the first user interface element is not within the zone of attention associated with the user.

46. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions to perform the method of any one of claims 28 to 41.

47. 42. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 28 to 41.

48. 1. An electronic device comprising: one or more processors; Memory and and means for performing the method of any one of claims 28 to 41.

49. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: and means for executing the method according to any one of claims 28 to 41.

50. 1. A method comprising:

1. An electronic device in communication with a display generation component and one or more input devices including an eye tracking device, comprising: displaying, via the display generation component, a user interface including a first region including a first user interface object and a second user interface object; detecting, while displaying the user interface, a discrete input provided by a predetermined portion of the user via the one or more input devices while detecting, via the eye tracking device, a gaze of the user directed toward the first region of the user interface, wherein during the discrete input, a location of the predetermined portion of the user is away from a location corresponding to the first region of the user interface; In response to detecting the individual input, performing an action on the first user interface object based on the individual input in accordance with a determination that one or more first criteria are met; and performing an action on the second user interface object based on the individual input in accordance with a determination that one or more second criteria different from the first criteria are satisfied.

51. the user interface includes a three-dimensional environment; the first region being a discrete distance from a viewpoint associated with the electronic device in the three-dimensional environment; In response to a determination that the respective distance is a first distance, the first region has a first size in the three-dimensional environment; 51. The method of claim 50, wherein, in accordance with a determination that the distinct distance is a second distance different from the first distance, the first region has a second size in the three-dimensional environment different from the first size.

52. 52. The method of claim 51, wherein a size of the first region in the three-dimensional environment increases as the individual distances increase.

53. 53. The method of any one of claims 50 to 52, wherein the one or more first criteria are met when the first object is closer to the user's viewpoint in the three-dimensional environment than the second object, and the one or more second criteria are met when the second object is closer to the user's viewpoint in the three-dimensional environment than the first object.

54. 54. The method of any one of claims 50 to 53, wherein the one or more first criteria are met or the one or more second criteria are met based on the type of the first user interface object and the type of the second user interface object.

55. 55. The method of any one of claims 50 to 54, wherein the one or more first criteria are met or the one or more second criteria are met based on individual priorities defined for the first user interface object and the second user interface object by the electronic device.

56. In response to detecting the individual input, 56. The method of any one of claims 50 to 55, further comprising refraining from performing the action on the first user interface object and refraining from performing the action on the second user interface object in accordance with a determination that one or more third criteria are satisfied, the third criteria including criteria that are satisfied when the first region exceeds a threshold distance from a viewpoint associated with the electronic device in a three-dimensional environment.

57. visually obscuring the first user interface object and the second user interface object relative to an area of ​​the user interface outside the first area in accordance with a determination that the first area exceeds the threshold distance from the viewpoint associated with the electronic device within the three-dimensional environment; 57. The method of claim 56, further comprising: refraining from visually obscuring the first user interface object and the second user interface object relative to the region of the user interface outside the first region in accordance with a determination that the first region is less than the threshold distance from the viewpoint associated with the electronic device in the three-dimensional environment.

58. detecting a second separate input provided by the predetermined portion of the user via the one or more input devices while displaying the user interface; In response to detecting the second discrete input, 58. The method of claim 50, further comprising: refraining from performing the respective action on the first user interface object and refraining from performing the respective action on the second user interface object in accordance with a determination that one or more third criteria are met, the third criteria including a criterion that is met when the first region is greater than a threshold angle from the line of sight of the user in a three-dimensional environment.

59. visually obscuring the first user interface object and the second user interface object relative to a region of the user interface outside the first region in accordance with a determination that the first region exceeds the threshold angle from the viewpoint associated with the electronic device within the three-dimensional environment; 59. The method of claim 58, further comprising: refraining from visually obscuring the first user interface object and the second user interface object relative to the region of the user interface outside the first region in accordance with a determination that the first region is less than the threshold angle from the viewpoint associated with the electronic device in the three-dimensional environment.

60. wherein the one or more first criteria and the one or more second criteria include respective criteria that are met when the first region is more than a threshold distance from a viewpoint associated with the electronic device in a three-dimensional environment and are not met when the first region is less than the threshold distance from the viewpoint associated with the electronic device in the three-dimensional environment, the method comprising: in response to detecting the discrete input and in accordance with a determination that the first region is less than the threshold distance from the viewpoint associated with the electronic device within the three-dimensional environment; performing the action on the first user interface object based on the individual input in accordance with a determination that the user's gaze is directed toward the first user interface object; 60. The method of claim 50, further comprising: performing the action on the second user interface object based on the individual input in accordance with a determination that the user's gaze is directed toward the second user interface object.

61. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising: displaying, via a display generation component, a user interface including a first region including a first user interface object and a second user interface object; detecting, while displaying the user interface, a gaze of the user directed toward the first region of the user interface via an eye tracking device, a discrete input provided by a predetermined portion of the user via the one or more input devices, wherein during the discrete input, a location of the predetermined portion of the user is away from a location corresponding to the first region of the user interface; In response to detecting the individual input, pursuant to a determination that one or more first criteria are met, performing an action on the first user interface object based on the individual input; pursuant to a determination that one or more second criteria different from the first criteria are satisfied, performing an action on the second user interface object based on the individual input.

62. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the electronic device, cause the electronic device to: displaying, via a display generation component, a user interface including a first region including a first user interface object and a second user interface object; detecting, while displaying the user interface, a gaze of the user directed toward the first region of the user interface via an eye tracking device, a discrete input provided by a predetermined portion of the user via the one or more input devices, wherein during the discrete input, a location of the predetermined portion of the user is away from a location corresponding to the first region of the user interface; In response to detecting the individual input, performing an action on the first user interface object based on the individual input in accordance with a determination that one or more first criteria are met; and performing an action on the second user interface object based on the individual input in accordance with a determination that one or more second criteria different from the first criteria are satisfied.

63. 1. An electronic device comprising: one or more processors; Memory and means for displaying, via a display generation component, a user interface including a first region including a first user interface object and a second user interface object; means for detecting, while displaying the user interface, a discrete input provided by a predetermined portion of the user via the one or more input devices while detecting, via an eye tracking device, a gaze of the user directed toward the first region of the user interface, wherein during the discrete input, a location of the predetermined portion of the user is away from a location corresponding to the first region of the user interface; In response to detecting the individual input, pursuant to a determination that one or more first criteria are met, performing an action on the first user interface object based on the individual input; means for performing an action on the second user interface object based on the individual input in accordance with a determination that one or more second criteria different from the first criteria are satisfied.

64. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: means for displaying, via a display generation component, a user interface including a first region including a first user interface object and a second user interface object; means for detecting, while displaying the user interface, a discrete input provided by a predetermined portion of the user via the one or more input devices while detecting, via an eye tracking device, a gaze of the user directed toward the first region of the user interface, wherein during the discrete input, a location of the predetermined portion of the user is away from a location corresponding to the first region of the user interface; In response to detecting the individual input, pursuant to a determination that one or more first criteria are met, performing an action on the first user interface object based on the individual input; means for performing an action on the second user interface object based on the individual input in accordance with a determination that one or more second criteria different from the first criteria are satisfied.

65. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions to perform the method of any one of claims 50 to 60.

66. 61. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 50 to 60.

67. 1. An electronic device comprising: one or more processors; Memory and and means for performing the method of any one of claims 50 to 60.

68. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: and means for executing the method according to any one of claims 50 to 60.

69. 1. A method comprising:

1. An electronic device in communication with a display generation component and one or more input devices including an eye tracking device, comprising: displaying, via the display generation component, a user interface including a plurality of user interface objects of a distinct type, including a first user interface object in a first state and a second user interface object in the first state; displaying, via the display generation component, the first user interface object in a second state different from the first state while maintaining the display of the second user interface object in the first state in accordance with a determination that one or more criteria are satisfied, including a criterion that is satisfied if a first predetermined portion of the user of the electronic device is farther than a threshold distance from a location corresponding to one of the plurality of user interface objects in the user interface while a gaze of the user of the electronic device is directed toward the first user interface object; While the user's gaze is directed at the first user interface object, detecting a movement of the first predetermined portion of the user via the one or more input devices while displaying the first user interface object in the second state; In response to detecting the movement of the first predetermined portion of the user, and displaying the second user interface object in the second state via the display generation component in accordance with a determination that the first predetermined portion of the user has moved within the threshold distance of a location corresponding to the second user interface object.

70. In response to detecting the movement of the first predetermined portion of the user, 70. The method of claim 69, further comprising: displaying, via the display generation component, the first user interface object in the first state in accordance with the determination that the first predetermined portion of the user has moved within the threshold distance of the location corresponding to the second user interface object.

71. In response to detecting the movement of the first predetermined portion of the user, 71. The method of claim 69 or 70, further comprising maintaining the display of the first user interface object in the second state in accordance with a determination that the first predetermined portion of the user has moved within the threshold distance of a location corresponding to the first user interface object.

72. In response to detecting the movement of the first predetermined portion of the user, 72. The method of claim 69, further comprising: displaying, via the display generation component, the third user interface object in the second state in accordance with a determination that the first predetermined portion of the user has moved within the threshold distance of a location corresponding to a third user interface object of the plurality of user interface objects.

73. In response to detecting the movement of the first predetermined portion of the user, in response to a determination that the first predetermined portion of the user has moved within the threshold distance of a location corresponding to the first user interface object and the location corresponding to the second user interface object; displaying, via the display generation component, the first user interface object in the second state in accordance with a determination that the first predetermined portion is closer to the location corresponding to the first user interface object than to the location corresponding to the second user interface object; 73. The method of claim 69, further comprising: displaying, via the display generation component, the second user interface object in the second state in accordance with a determination that the first predetermined portion is closer to the location corresponding to the second user interface object than to the location corresponding to the first user interface object.

74. 74. The method of any one of claims 69 to 73, wherein the one or more criteria include a criterion that is met when the first predetermined portion of the user is in a predetermined posture.

75. In response to detecting the movement of the first predetermined portion of the user, maintaining the display of the first user interface object in the second state in accordance with a determination that the first predetermined portion of the user has moved within the threshold distance of a location corresponding to the first user interface object; the first user interface object in the second state has a first visual appearance when the first predetermined portion of the user is greater than the threshold distance of the location corresponding to the first user interface object; 75. The method of any one of claims 69 to 74, wherein when the first predetermined portion of the user is within the threshold distance of the location corresponding to the first user interface object, the first user interface object in the second state has a second visual appearance that is different from the first visual appearance.

76. displaying, via the display generation component, the first user interface object in the second state in accordance with a determination that one or more second criteria are satisfied, the second criteria including a criterion satisfied when, while the gaze of the user is directed at the first user interface object, a second predetermined portion of the user, different from the first predetermined portion, is farther than the threshold distance from the location corresponding to any of the plurality of user interface objects in the user interface; In accordance with the determination that the one or more criteria are satisfied, the first user interface object in the second state has a first visual appearance; 76. The method of any one of claims 69 to 75, wherein, in accordance with the determination that the one or more second criteria are satisfied, the first user interface object in the second state has a second visual appearance that is different from the first visual appearance.

77. 77. The method of any one of claims 69 to 76, wherein displaying the second user interface object in the second state occurs while the user's gaze remains directed at the first user interface object.

78. 78. The method of any one of claims 69 to 77, wherein displaying the second user interface object in the second state is further subject to a determination that the second user interface object is within a zone of attention associated with the user of the electronic device.

79. 79. The method of any one of claims 69 to 78, wherein the one or more criteria include a criterion that is met when at least one predetermined portion of the user, including the first predetermined portion of the user, is in a predetermined posture.

80. detecting, via the one or more input devices, a first movement of a zone of attention associated with the user while displaying the first user interface object in the second state; In response to detecting the first movement of the zone of attention associated with the user, 80. The method of claim 69, further comprising: displaying, via the display generation component, the third user interface object in the second state in accordance with a determination that the zone of attention includes a third user interface object of the respective type and that the first predetermined portion of the user is within the threshold distance of a location corresponding to the third user interface object.

81. after detecting the first movement of the zone of attention, while displaying the third user interface object in the second state, detecting a second movement of the zone of attention via the one or more input devices, wherein the third user interface object is no longer within the zone of attention as a result of the second movement of the zone of attention; In response to detecting the second movement of the attention zone, 81. The method of claim 80, further comprising: maintaining the display of the third user interface object in the second state in accordance with a determination that the first predetermined portion of the user is within the threshold distance of the third user interface object.

82. in response to detecting the second movement of the zone of attention and in accordance with a determination that the first predetermined portion of the user is not engaging with the third user interface object; displaying the first user interface object in the second state in accordance with a determination that the first user interface object is within the attention zone, the one or more criteria are satisfied, and the user's gaze is directed toward the first user interface object; 82. The method of claim 81, further comprising: displaying the second user interface object in the second state in accordance with a determination that the second user interface object is within the zone of attention, the one or more criteria are satisfied, and the user's gaze is directed toward the second user interface object.

83. While said one or more criteria are met, detecting, via the eye tracking device, a movement of the user's gaze to the second user interface object prior to detecting the movement of the first predetermined portion of the user and while displaying the first user interface object in the second state; 83. The method of any one of claims 69 to 82, further comprising: in response to detecting the movement of the user's gaze to the second user interface object, displaying the second user interface object in the second state via the display generation component.

84. detecting, via the eye tracking device, a movement of the user's gaze toward the first user interface object after detecting the movement of the first predetermined portion of the user and while displaying the second user interface object in the second state in accordance with the determination that the first predetermined portion of the user is within the threshold distance of the location corresponding to the second user interface object; 84. The method of claim 83, further comprising: in response to detecting the movement of the user's gaze to the first user interface object, maintaining the display of the second user interface object in the second state.

85. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising: displaying, via a display generation component, a user interface including a plurality of user interface objects of a distinct type, including a first user interface object in a first state and a second user interface object in the first state; pursuant to a determination that one or more criteria are satisfied, including a criterion that is satisfied if a first predetermined portion of the user of the electronic device is farther than a threshold distance from a location corresponding to one of the plurality of user interface objects in the user interface while a gaze of the user of the electronic device is directed toward the first user interface object, displaying the first user interface object in a second state different from the first state, while maintaining the display of the second user interface object in the first state, via the display generation component; While the user's gaze is directed at the first user interface object, detecting a movement of the first predetermined portion by the user via the one or more input devices while displaying the first user interface object in the second state; In response to detecting the movement of the first predetermined portion of the user, and instructions for displaying, via the display generation component, the second user interface object in the second state in accordance with a determination that the first predetermined portion of the user has moved within the threshold distance of a location corresponding to the second user interface object.

86. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the electronic device, cause the electronic device to: displaying, via a display generation component, a user interface including a plurality of user interface objects of a distinct type, including a first user interface object in a first state and a second user interface object in the first state; displaying, via the display generation component, the first user interface object in a second state different from the first state while maintaining the display of the second user interface object in the first state in accordance with a determination that one or more criteria are satisfied, including a criterion that is satisfied if a first predetermined portion of the user of the electronic device is farther than a threshold distance from a location corresponding to one of the plurality of user interface objects in the user interface while a gaze of the user of the electronic device is directed toward the first user interface object; While the user's gaze is directed at the first user interface object, detecting a movement of the first predetermined portion of the user via the one or more input devices while displaying the first user interface object in the second state; In response to detecting the movement of the first predetermined portion of the user, and displaying, via the display generation component, the second user interface object in the second state in accordance with a determination that the first predetermined portion of the user has moved within the threshold distance of a location corresponding to the second user interface object.

87. 1. An electronic device comprising: one or more processors; Memory and means for displaying, via a display generation component, a user interface including a plurality of user interface objects of a distinct type, including a first user interface object in a first state and a second user interface object in the first state; means for displaying, via the display generation component, the first user interface object in a second state different from the first state while maintaining the display of the second user interface object in the first state in accordance with a determination that one or more criteria are satisfied, including a criterion that is satisfied if a first predetermined portion of the user of the electronic device is farther than a threshold distance from a location corresponding to one of the plurality of user interface objects in the user interface while a gaze of the user of the electronic device is directed toward the first user interface object; While the user's gaze is directed at the first user interface object, detecting a movement of the first predetermined portion by the user via the one or more input devices while displaying the first user interface object in the second state; In response to detecting the movement of the first predetermined portion of the user, means for displaying, via the display generation component, the second user interface object in the second state in accordance with a determination that the first predetermined portion of the user has moved within the threshold distance of a location corresponding to the second user interface object.

88. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: means for displaying, via a display generation component, a user interface including a plurality of user interface objects of a distinct type, including a first user interface object in a first state and a second user interface object in the first state; means for displaying, via the display generation component, the first user interface object in a second state different from the first state while maintaining the display of the second user interface object in the first state in accordance with a determination that one or more criteria are satisfied, including a criterion that is satisfied if a first predetermined portion of the user of the electronic device is farther than a threshold distance from a location corresponding to one of the plurality of user interface objects in the user interface while a gaze of the user of the electronic device is directed toward the first user interface object; While the user's gaze is directed at the first user interface object, detecting a movement of the first predetermined portion by the user via the one or more input devices while displaying the first user interface object in the second state; In response to detecting the movement of the first predetermined portion of the user, means for displaying, via the display generation component, the second user interface object in the second state in accordance with a determination that the first predetermined portion of the user has moved within the threshold distance of a location corresponding to the second user interface object.

89. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions to perform the method of any one of claims 69 to 84.

90. 85. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 69 to 84.

91. 1. An electronic device comprising: one or more processors; Memory and and means for performing the method of any one of claims 69 to 84.

92. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: and means for performing the method of any one of claims 69 to 84.

93. 1. A method comprising:

1. An electronic device in communication with a display generation component and one or more input devices including an eye tracking device, comprising: detecting, via the eye tracking device, a movement of the user's gaze from a first user interface element displayed via the display generation component to a second user interface element displayed via the display generation component while the user's gaze is directed at the first user interface element displayed via the display generation component; in response to detecting the movement of the user's gaze from the first user interface element to the second user interface element displayed via the display generation component; modifying a visual appearance of the second user interface element in accordance with a determination that a second predetermined portion of the user different from the first predetermined portion is available for engagement with the second user interface element; and forgoing altering the visual appearance of the second user interface element in accordance with a determination that the second predetermined portion of the user is not available for engagement with the second user interface element.

94. while one or more criteria are satisfied, including criteria that are satisfied when the first predetermined portion of the user and the second predetermined portion of the user are not engaging with any user interface element; pursuant to determining that the user's gaze is directed toward the first user interface element, displaying the first user interface element with a visual characteristic indicating that engagement with the first user interface element is available, the second user interface element being displayed without the visual characteristic; and displaying the second user interface element with the visual characteristic indicating that engagement with the second user interface element is available in accordance with determining that the user's gaze is directed toward the second user interface element, the first user interface element being displayed without the visual characteristic; detecting an input from the first predetermined portion or the second predetermined portion of the user via the one or more input devices while the one or more criteria are met; In response to detecting the input, performing an action corresponding to the first user interface element in accordance with the determination that the user's gaze was directed toward the first user interface element when the input was received; and 94. The method of claim 93, further comprising: performing an action corresponding to the second user interface element in accordance with the determination that the user's gaze was directed toward the second user interface element when the input was accepted.

95. 95. The method of claim 94, wherein the one or more criteria include a criterion that is met when at least one of the first predetermined portion or the second predetermined portion of the user is available for engagement with a user interface element.

96. in response to detecting the movement of the user's gaze from the first user interface element to the second user interface element displayed via the display generation component; 96. The method of any one of claims 93 to 95, further comprising: refraining from altering the visual appearance of the second user interface element in accordance with a determination that the first predetermined portion and the second predetermined portion of the user are not available for engagement with a user interface element.

97. after changing the visual appearance of the second user interface element to the changed appearance of the second user interface element while the second predetermined portion of the user is available for engagement with the second user interface element, detecting via the eye tracking device that the second predetermined portion of the user is no longer available for engagement with the second user interface element; 97. The method of any one of claims 93 to 96, further comprising: ceasing to display the modified appearance of the second user interface element in response to detecting that the second predetermined portion of the user is no longer available for engagement with the second user interface element.

98. after determining that the second predetermined portion of the user is not available for engagement with the second user interface element, detecting, via the one or more input devices while the user's gaze is directed toward the second user interface element, that the second predetermined portion of the user is currently available for engagement with the second user interface element; 98. The method of any one of claims 93 to 97, further comprising: modifying the visual appearance of the second user interface element in response to detecting that the second predetermined portion of the user is currently available for engagement with the second user interface element.

99. in response to detecting the movement of the user's gaze from the first user interface element to the second user interface element displayed via the display generation component; 99. The method of any one of claims 93 to 98, further comprising refraining from altering the visual appearance of the second user interface element in accordance with a determination that the first predetermined portion and the second predetermined portion of the user are already engaged with a separate user interface element other than the second user interface element.

100. 100. The method of any one of claims 93 to 99, wherein the determination that the second predetermined portion of the user is not available for engagement with the second user interface element is based on a determination that the second predetermined portion of the user is engaged with a third user interface element that is different from the second user interface element.

101. 101. The method of any one of claims 93 to 100, wherein the determination that the second predetermined portion of the user is not available for engagement with the second user interface element is based on a determination that the second predetermined portion of the user is not in a predetermined posture required for engagement with the second user interface element.

102. 102. The method of any one of claims 93 to 101, wherein the determination that the second predetermined portion of the user is not available for engagement with the second user interface element is based on a determination that the second predetermined portion of the user is not detected by the one or more input devices in communication with the electronic device.

103. while displaying the first user interface element and the second user interface element via the display generation component; in response to determining that the first predetermined portion of the user is within a threshold distance of a location corresponding to the first user interface element and that the second predetermined portion of the user is within the threshold distance of a location corresponding to the second user interface element; displaying the first user interface element with visual characteristics that indicate the first predetermined portion of the user is available for direct engagement with the first user interface element; 103. The method of any one of claims 93 to 102, further comprising: displaying the second user interface element with the visual characteristics indicating that the second user interface element is available for the user's direct engagement with the second predetermined portion.

104. while displaying the first user interface element and the second user interface element via the display generation component; in response to determining that the first predetermined portion of the user is within a threshold distance of a location corresponding to the first user interface element and that the second predetermined portion of the user is further than the threshold distance of a location corresponding to the second user interface element but is available for engagement with the second user interface element; displaying the first user interface element with visual characteristics that indicate the first predetermined portion of the user is available for direct engagement with the first user interface element; displaying the second user interface element with visual characteristics indicating that the second predetermined portion of the user is available for indirect engagement with the second user interface element in accordance with determining that the gaze of the user is directed toward the second user interface element; 104. The method of any one of claims 93 to 103, further comprising: in accordance with a determination that the user's gaze is not directed toward the second user interface element, displaying the second user interface element without the visual characteristic indicating that the second predetermined portion of the user is available for indirect engagement with the second user interface element.

105. while displaying the first user interface element and the second user interface element via the display generation component; in response to determining that the second predetermined portion of the user is within a threshold distance of a location corresponding to the second user interface element and that the first predetermined portion of the user is further than the threshold distance of a location corresponding to the first user interface element but is available for engagement with the first user interface element; displaying the second user interface element with a visual characteristic that indicates the second user interface element is available for the user's direct engagement with the second predetermined portion; displaying the first user interface element with visual characteristics indicating that the first predetermined portion of the user is available for indirect engagement with the first user interface element in accordance with determining that the gaze of the user is directed toward the first user interface element; 105. The method of any one of claims 93 to 104, further comprising: in accordance with a determination that the user's gaze is not directed toward the first user interface element, displaying the first user interface element without the visual characteristic indicating that the first predetermined portion of the user is available for indirect engagement with the first user interface element.

106. detecting, after detecting the movement of the user's gaze from the first user interface element to the second user interface element, the second predetermined portion of the user directly engaging with the first user interface element via the one or more input devices while displaying the second user interface element with the changed visual appearance; 106. The method of any one of claims 93 to 105, further comprising: in response to detecting that the second predetermined portion of the user is directly engaging with the first user interface element, refraining from displaying the second user interface element in the altered visual appearance.

107. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising: detecting, via an eye tracking device, a movement of a gaze of a user of the electronic device from a first user interface element displayed via the display generation component to a second user interface element displayed via the display generation component while the gaze of the user of the electronic device is directed at the first user interface element; in response to detecting the movement of the user's gaze from the first user interface element to the second user interface element displayed via the display generation component; modifying a visual appearance of the second user interface element in accordance with a determination that a second predetermined portion of the user different from the first predetermined portion is available for engagement with the second user interface element; and instructions for forgoing altering the visual appearance of the second user interface element in accordance with a determination that the second predetermined portion of the user is not available for engagement with the second user interface element.

108. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the electronic device, cause the electronic device to: detecting, via the eye tracking device, a movement of the user's gaze from a first user interface element displayed via the display generation component to a second user interface element displayed via the display generation component while the user's gaze is directed at the first user interface element displayed via the display generation component; in response to detecting the movement of the user's gaze from the first user interface element to the second user interface element displayed via the display generation component; modifying a visual appearance of the second user interface element in accordance with a determination that a second predetermined portion of the user different from the first predetermined portion is available for engagement with the second user interface element; and forgoing altering the visual appearance of the second user interface element in accordance with a determination that the second predetermined portion of the user is not available for engagement with the second user interface element.

109. 1. An electronic device comprising: one or more processors; Memory and means for detecting, via an eye tracking device, a movement of a user's gaze from a first user interface element displayed via the display generation component to a second user interface element displayed via the display generation component while the user's gaze is directed at the first user interface element; in response to detecting the movement of the user's gaze from the first user interface element to the second user interface element displayed via the display generation component; modifying a visual appearance of the second user interface element in accordance with a determination that a second predetermined portion of the user different from the first predetermined portion is available for engagement with the second user interface element; means for forgoing altering the visual appearance of the second user interface element in accordance with a determination that the second predetermined portion of the user is not available for engagement with the second user interface element.

110. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: means for detecting, via an eye tracking device, a movement of a user's gaze from a first user interface element displayed via the display generation component to a second user interface element displayed via the display generation component while the user's gaze is directed at the first user interface element; in response to detecting the movement of the user's gaze from the first user interface element to the second user interface element displayed via the display generation component; modifying a visual appearance of the second user interface element in accordance with a determination that a second predetermined portion of the user different from the first predetermined portion is available for engagement with the second user interface element; means for forgoing altering the visual appearance of the second user interface element in accordance with a determination that the second predetermined portion of the user is not available for engagement with the second user interface element.

111. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions to perform the method of any one of claims 93 to 106.

112. 107. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 93 to 106.

113. 1. An electronic device comprising: one or more processors; Memory and and means for performing the method of any one of claims 93 to 106.

114. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: and means for executing the method of any one of claims 93 to 106.

115. 1. A method comprising: An electronic device in communication with a display generation component and one or more input devices, displaying a user interface object within a three-dimensional environment via the display generation component; detecting, while displaying the user interface object, a discrete input via the one or more input devices, the discrete input comprising movement of a predetermined portion of a user of the electronic device, wherein during the discrete input, a location of the predetermined portion of the user is away from a location corresponding to the user interface object; While detecting said discrete input, displaying, via the display generation component, a visual indication at a first location within the three-dimensional environment corresponding to the first position of the predetermined portion of the user in accordance with a determination that a first portion of the movement of the predetermined portion of the user satisfies one or more criteria and the predetermined portion of the user is at a first position; and displaying, via the display generation component, a visual indication at a second location within the three-dimensional environment corresponding to the second position of the predetermined portion of the user, the second location being different from the first location, in accordance with a determination that the first portion of the movement of the predetermined portion of the user satisfies the one or more criteria and the predetermined portion of the user is at a second position.

116. While detecting said discrete input, performing a selection action on the user interface object in accordance with the individual input in accordance with the determination that the first portion of the movement of the predetermined portion of the user satisfies the one or more criteria and one or more second criteria are satisfied, including a criterion that is satisfied if the first portion of the movement of the predetermined portion of the user is followed by a second portion of the movement of the predetermined portion of the user; 116. The method of claim 115, further comprising: refraining from performing the selection action on the user interface object in accordance with the determination that the first portion of the movement of the predetermined portion of the user does not satisfy the one or more criteria and the one or more second criteria are satisfied.

117. 117. The method of claim 115 or 116, further comprising displaying, via the display generation component, a representation of the predetermined portion of the user that moves in accordance with the movement of the predetermined portion of the user while detecting the discrete input.

118. 118. The method of any one of claims 115 to 117, wherein the predetermined portion of the user is visible through the display generation component of the three-dimensional environment.

119. 119. The method of any one of claims 115 to 118, further comprising, while detecting the individual input, modifying a display of the user interface object in accordance with the individual input in accordance with the determination that the first portion of the movement of the predetermined portion of the user satisfies the one or more criteria.

120. Modifying the display of the user interface object comprises:

120. The method of claim 119, further comprising: moving the user interface object backward in the three-dimensional environment in accordance with the movement of the predetermined portion of the user toward the location corresponding to the user interface object in accordance with a determination that the predetermined portion of the user moves toward a location corresponding to the user interface object after the first portion of the movement of the predetermined portion of the user satisfies the one or more criteria.

121. the user interface objects are displayed in a respective user interface via the display generation component; In response to determining that the respective input is a scroll input, the electronic device moves the respective user interface and the user interface object backward in accordance with the movement of the predetermined portion of the user toward the location corresponding to the user interface object; 121. The method of claim 120, wherein, in accordance with a determination that the respective input is an input other than a scroll input, the electronic device moves the user interface object relative to the respective user interface without moving the respective user interface.

122. While detecting said discrete input, detecting movement of the predetermined portion of the user away from the location corresponding to the user interface object after detecting the movement of the predetermined portion of the user toward the user interface object and after moving the user interface object backward in the three-dimensional environment; 122. The method of claim 120 or 121, further comprising: in response to detecting the movement of the predetermined portion of the user away from the location corresponding to the user interface object, moving the user interface object forward in the three-dimensional environment in accordance with the movement of the predetermined portion of the user away from the location corresponding to the user interface object.

123. the visual indication at the first location within the three-dimensional environment corresponding to the first position of the predetermined portion of the user is displayed near a representation of the predetermined portion of the user visible within the three-dimensional environment at a first discrete location within the three-dimensional environment; 123. The method of any one of claims 115 to 122, wherein the visual indication at the second location within the three-dimensional environment corresponding to the second position of the predetermined portion of the user is displayed near the representation of the predetermined portion of the user visible in the three-dimensional environment at a second distinct location within the three-dimensional environment.

124. detecting, while displaying the user interface object, a second discrete input via the one or more input devices, the second discrete input including movement of the predetermined portion of the user, wherein during the second discrete input, the location of the predetermined portion of the user is at the location corresponding to the user interface object; While detecting the second discrete input, 124. The method of any one of claims 119 to 123, further comprising: via the display generation component, modifying a display of the user interface object in accordance with the second individual input without displaying the visual indication at the location corresponding to the predetermined portion of the user.

125. The electronic device performs a distinct action in response to the distinct input, and the method further comprises: detecting, while displaying the user interface object, a third discrete input via the one or more input devices, the third discrete input including a movement of the predetermined portion of the user including a same type of movement of the predetermined portion of the user as the movement of the predetermined portion of the user in the discrete input, wherein during the third discrete input, the location of the predetermined portion of the user is at the location corresponding to the user interface object; 125. The method of any one of claims 119 to 124, further comprising: performing the discrete action in response to detecting the third discrete input.

126. before detecting the discrete input; displaying the user interface object with a respective visual characteristic having a first value in accordance with a determination that the user's gaze is directed toward the user interface object; 126. The method of any one of claims 115 to 125, further comprising: displaying the user interface object with the individual visual characteristic having a second value different from the first value in accordance with a determination that the user's gaze is not directed toward the user interface object.

127. While detecting said discrete input, after the first portion of the movement of the predetermined portion of the user satisfies the one or more criteria; 127. The method of claim 115, further comprising: performing a tap action on the user interface object in accordance with a determination that a second portion of the movement of the predetermined portion of the user has been detected that satisfies one or more second criteria, including a criterion that is met when a second portion of the movement of the predetermined portion of the user includes movement toward the location corresponding to the user interface object that is greater than a movement threshold, followed by one or more third criteria, including a criterion that is met when a third portion of the movement is away from the location corresponding to the user interface object and is detected within a time threshold of the second portion of the movement.

128. While detecting said discrete input, after the first portion of the movement of the predetermined portion of the user satisfies the one or more criteria; 128. The method of claim 115, further comprising: performing a scrolling operation on the user interface object in accordance with a determination that the second portion of the movement of the predetermined portion of the user satisfies one or more second criteria, including a criterion that is met when the second portion of the movement of the predetermined portion of the user includes movement greater than a movement threshold toward the location corresponding to the user interface object, followed by one or more third criteria, including a criterion that is met when the third portion of the movement is lateral movement relative to the location corresponding to the user interface object.

129. While detecting said discrete input, detecting, via the one or more input devices, a second portion of the movement of the predetermined portion of the user away from the location corresponding to the user interface object after the first portion of the movement of the predetermined portion of the user satisfies the one or more criteria; 129. The method of any one of claims 115 to 128, further comprising: in response to detecting the second portion of the movement, updating an appearance of the visual indication according to the second portion of the movement.

130. Updating the appearance of the visual indication includes ceasing display of the visual indication, the method comprising: detecting, after ceasing to display the visual indication, a second discrete input via the one or more input devices, the second discrete input including a second movement of the predetermined portion of the user, wherein during the second discrete input, the location of the predetermined portion of the user is away from the location corresponding to the user interface object; While detecting the second discrete input, 130. The method of claim 129, further comprising: in accordance with a determination that a first portion of the second movement satisfies the one or more criteria, displaying, via the display generation component, a second visual indication at a location within the three-dimensional environment corresponding to the predetermined portion of the user in the second individual input.

131. The discrete input corresponds to a scrolling input directed at the user interface object, and the method further comprises:

131. The method of any one of claims 115 to 130, further comprising scrolling the user interface object in accordance with the discrete input while maintaining display of the visual indication.

132. While detecting said discrete input, after the first portion of the movement of the predetermined portion of the user satisfies the one or more criteria, detecting, via the one or more input devices, a second portion of the movement of the predetermined portion of the user that satisfies one or more second criteria, including a criterion that is satisfied when the second portion of the movement corresponds to a distance between a location corresponding to the visual indication and the predetermined portion of the user; 132. The method of any one of claims 115 to 131, further comprising: in response to detecting the second portion of the movement of the predetermined portion of the user, generating audio feedback indicating that the one or more second criteria are met.

133. detecting that one or more second criteria are satisfied while displaying the user interface object, the second criteria including a criterion that is satisfied when the predetermined portion of the user has a distinct pose while the location of the predetermined portion of the user is away from the location corresponding to the user interface object; 133. The method of any one of claims 115 to 132, further comprising: in response to detecting that the one or more second criteria are met, displaying, via the display generation component, a virtual surface near a location corresponding to the predetermined portion of the user and away from the user interface object.

134. detecting, while displaying the virtual surface, a discrete movement of the predetermined portion of the user via the one or more input devices toward a location corresponding to the virtual surface; 134. The method of claim 133, further comprising: in response to detecting the individual movements, modifying the visual appearance of the virtual surface in accordance with the individual movements.

135. detecting, while displaying the virtual surface, a discrete movement of the predetermined portion of the user via the one or more input devices toward a location corresponding to the virtual surface; 134. The method of claim 132 or 133, further comprising, in response to detecting the individual movements, modifying a visual appearance of the user interface object in accordance with the individual movements.

136. 136. The method of any one of claims 133 to 135, wherein displaying the virtual surface near a location corresponding to the predetermined portion of the user includes displaying the virtual surface at discrete distances from the location corresponding to the predetermined portion of the user, the discrete distances corresponding to amounts of movement of the predetermined portion of the user toward the location corresponding to the virtual surface required to perform an action on the user interface object.

137. 137. The method of any one of claims 133 to 136, further comprising, while displaying the virtual surface, displaying a visual indication on the virtual surface of the distance between the predetermined portion of the user and a location corresponding to the virtual surface.

138. detecting, while displaying the virtual surface, movement of the predetermined portion of the user via the one or more input devices to a discrete location that is greater than a threshold distance from a location corresponding to the virtual surface; 138. The method of any one of claims 133 to 137, further comprising ceasing display of the virtual surface in the three-dimensional environment in response to detecting the movement of the predetermined portion of the user to the individual location.

139. Displaying the virtual surface near the predetermined portion of the user includes: displaying the virtual surface at a third location within the three-dimensional environment corresponding to the first discrete position of the predetermined portion of the user in accordance with a determination that the predetermined portion of the user is in a first discrete position when the one or more second criteria are met; 139. The method of any one of claims 133 to 138, comprising: when the one or more second criteria are satisfied in accordance with a determination that the predetermined portion of the user is in a second discrete position different from the first discrete position, displaying the virtual surface at a fourth location different from the third location in the three-dimensional environment corresponding to the second discrete position of the predetermined portion of the user.

140. while displaying the visual indication corresponding to the predetermined portion of the user; detecting a second discrete input via the one or more input devices, the second discrete input comprising a movement of a second predetermined portion of the user, wherein during the second discrete input, a location of the second predetermined portion of the user is away from the location corresponding to the user interface object; While detecting the second discrete input, In response to a determination that the first portion of the movement of the second predetermined portion of the user satisfies the one or more criteria, via the display generation component: the visual indication corresponding to the predetermined portion of the user; and and simultaneously displaying a visual indication at a location corresponding to the second predetermined portion of the user within the three-dimensional environment.

141. 141. The method of any one of claims 115 to 140, further comprising, while detecting the discrete input, displaying a discrete visual indication on the user interface object that indicates a discrete distance the predetermined portion of the user needs to move toward the location corresponding to the user interface object in order to engage the user interface object.

142. detecting, while displaying the user interface object, that the user's gaze is directed toward the user interface object; 142. The method of any one of claims 115 to 141, further comprising: in response to detecting that the user's gaze is directed toward the user interface object, displaying the user interface object with a respective visual characteristic having a first value.

143. the three-dimensional environment includes representations of individual objects within a physical environment of the electronic device, and the method further comprises: Detecting that one or more second criteria are met, including a criterion that is met when the user's gaze is directed toward the representation of the individual object and a criterion that is met when the predetermined portion of the user is in a individual pose; and 143. The method of any one of claims 115 to 142, further comprising: in response to detecting that the one or more second criteria are met, displaying, via the display generation component, one or more selectable options near the representation of the individual object, the one or more selectable options being selectable to perform a particular action associated with the individual object.

144. after the first portion of the movement of the predetermined portion of the user satisfies the one or more criteria, detecting, via the one or more input devices, a second portion of the movement of the predetermined portion of the user that satisfies one or more second criteria while displaying the visual indication corresponding to the predetermined portion of the user; In response to detecting the second portion of the movement of the predetermined portion of the user, In response to a determination that the user's gaze is directed toward the user interface object and that the user interface object is interactive, displaying, via the display generation component, a visual indication that the second portion of the movement of the predetermined portion of the user satisfies the one or more second criteria; performing an action corresponding to the user interface object in accordance with the individual input; in response to a determination that the gaze of the user is not directed toward an interactive user interface object, 144. The method of any one of claims 115 to 143, further comprising: displaying, via the display generation component, the visual indication that the second portion of the movement of the predetermined portion of the user satisfies the one or more second criteria without performing an action in accordance with the individual input.

145. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising: displaying user interface objects within a three-dimensional environment via said display generation component; detecting, while displaying the user interface object, a discrete input comprising movement of a predetermined portion of a user of the electronic device via the one or more input devices, wherein during the discrete input, a location of the predetermined portion of the user is away from a location corresponding to the user interface object; While detecting said discrete input, displaying, via the display generation component, a visual indication at a first location within the three-dimensional environment corresponding to the first position of the predetermined portion of the user in accordance with determining that a first portion of the movement of the predetermined portion of the user satisfies one or more criteria and the predetermined portion of the user is at a first position; In accordance with a determination that the first portion of the movement of the predetermined portion of the user satisfies the one or more criteria and the predetermined portion of the user is at a second position, the electronic device includes instructions for displaying, via the display generation component, a visual indication at a second location different from the first location within the three-dimensional environment corresponding to the second position of the predetermined portion of the user.

146. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the electronic device, cause the electronic device to: displaying a user interface object within a three-dimensional environment via the display generation component; detecting, while displaying the user interface object, a discrete input via the one or more input devices, the discrete input comprising movement of a predetermined portion of a user of the electronic device, wherein during the discrete input, a location of the predetermined portion of the user is away from a location corresponding to the user interface object; While detecting said discrete input, displaying, via the display generation component, a visual indication at a first location within the three-dimensional environment corresponding to the first position of the predetermined portion of the user in accordance with a determination that a first portion of the movement of the predetermined portion of the user satisfies one or more criteria and the predetermined portion of the user is at a first position; and in accordance with a determination that the first portion of the movement of the predetermined portion of the user satisfies the one or more criteria and the predetermined portion of the user is at a second position, displaying, via the display generation component, a visual indication at a second location within the three-dimensional environment that corresponds to the second position of the predetermined portion of the user, the second location being different from the first location.

147. 1. An electronic device comprising: one or more processors; Memory and means for displaying user interface objects within a three-dimensional environment via said display generation component; means for detecting, while displaying the user interface object, a discrete input comprising movement of a predetermined portion of a user of the electronic device via the one or more input devices, wherein during the discrete input, a location of the predetermined portion of the user is away from a location corresponding to the user interface object; While detecting said discrete input, displaying, via the display generation component, a visual indication at a first location within the three-dimensional environment corresponding to the first position of the predetermined portion of the user in accordance with determining that a first portion of the movement of the predetermined portion of the user satisfies one or more criteria and the predetermined portion of the user is at a first position; and means for displaying, via the display generation component, a visual indication at a second location within the three-dimensional environment corresponding to the second position of the predetermined portion of the user, the second location being different from the first location, in accordance with a determination that the first portion of the movement of the predetermined portion of the user satisfies the one or more criteria and the predetermined portion of the user is at a second position.

148. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: means for displaying user interface objects within a three-dimensional environment via said display generation component; means for detecting, while displaying the user interface object, a discrete input comprising movement of a predetermined portion of a user of the electronic device via the one or more input devices, wherein during the discrete input, a location of the predetermined portion of the user is away from a location corresponding to the user interface object; While detecting said discrete input, displaying, via the display generation component, a visual indication at a first location within the three-dimensional environment corresponding to the first position of the predetermined portion of the user in accordance with determining that a first portion of the movement of the predetermined portion of the user satisfies one or more criteria and the predetermined portion of the user is at a first position; and means for displaying, via the display generation component, a visual indication at a second location within the three-dimensional environment corresponding to the second position of the predetermined portion of the user, the second location being different from the first location, in accordance with a determination that the first portion of the movement of the predetermined portion of the user satisfies the one or more criteria and the predetermined portion of the user is at a second position.

149. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions to perform the method of any one of claims 115 to 144.

150. 145. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 115 to 144.

151. 1. An electronic device comprising: one or more processors; Memory and and means for performing the method of any one of claims 115 to 144.

152. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: and means for performing the method of any one of claims 115 to 144.

153. 1. A method comprising: An electronic device in communication with a display generation component and one or more input devices, displaying a user interface object via the display generation component; detecting, while displaying the user interface object, an input directed at the user interface object by a first predetermined portion of a user of the electronic device via the one or more input devices; and displaying, via the display generation component, a simulated shadow displayed on the user interface object while detecting the input directed at the user interface object, wherein the simulated shadow has an appearance based on a position of an element relative to the user interface object that is indicative of an interaction with the user interface object.

154. 154. The method of claim 153, wherein the element comprises a cursor displayed at a location corresponding to a location away from the first predetermined portion of the user and controlled by movement of the first predetermined portion of the user.

155. while displaying the user interface object and a second user interface object and before detecting the input directed to the user interface object by the first predetermined portion of the user; displaying, via the display generation component, the cursor at a predetermined distance from the user interface object in accordance with a determination that one or more first criteria are met, the first criteria including criteria that are met when the user's gaze is directed toward the user interface object; 155. The method of claim 154, further comprising: displaying, via the display generation component, the cursor at the predetermined distance from the second user interface object in accordance with a determination that one or more second criteria are met, the second criteria including criteria that are met when the user's gaze is directed toward the second user interface object.

156. 154. The method of claim 153, wherein the simulated shadow comprises a simulated shadow of a virtual representation of the first predetermined portion of the user.

157. 154. The method of claim 153, wherein the simulated shadow comprises a simulated shadow of the first predetermined physical part of the user.

158. while detecting the input directed at the user interface object and while displaying the simulated shadow displayed on the user interface object. detecting a progression of the input directed to the user interface object by the first predetermined portion of the user via the one or more input devices; 158. The method of any one of claims 153 to 157, further comprising: in response to detecting the progression of the input directed toward the user interface object, modifying a visual appearance of the simulated shadow displayed on the user interface object in accordance with the progression of the input directed toward the user interface object by the first predetermined portion of the user.

159. 159. The method of claim 158, wherein altering the visual appearance of the simulated shadow comprises altering the brightness at which the simulated shadow is displayed.

160. 160. The method of claim 158 or 159, wherein modifying the visual appearance of the simulated shadow comprises modifying the blurriness with which the simulated shadow is displayed.

161. 161. The method of any one of claims 158 to 160, wherein altering the visual appearance of the simulated shadow comprises altering a size of the simulated shadow.

162. while detecting the input directed at the user interface object and while displaying the simulated shadow displayed on the user interface object. detecting, via the one or more input devices, a first portion of the input corresponding to laterally moving the element relative to the user interface object; displaying the simulated shadow at a first location on the user interface object with a first visual appearance in response to detecting the first portion of the input; detecting, via the one or more input devices, a second portion of the input corresponding to laterally moving the element relative to the user interface object; 162. The method of any one of claims 153 to 161, further comprising: in response to detecting the second portion of the input, displaying the simulated shadow at a second location on the user interface object that is different from the first location and with a second visual appearance that is different from the first visual appearance.

163. 163. The method of any one of claims 153 to 162, wherein the user interface object is a virtual surface, and the input detected at a location near the virtual surface provides input to a second user interface object remote from the virtual surface.

164. 164. The method of any one of claims 153 to 163, wherein the first predetermined portion of the user is directly interacting with the user interface object and the simulated shadow is displayed on the user interface object.

165. the simulated shadow corresponds to the first predetermined portion of the user in accordance with a determination that the first predetermined portion of the user is within a threshold distance of a location corresponding to the user interface object; 165. The method of any one of claims 153 to 164, wherein, in accordance with a determination that the first predetermined portion of the user is farther than the threshold distance from the location corresponding to the user interface object, the simulated shadow corresponds to a cursor controlled by the first predetermined portion of the user.

166. detecting a second input directed at the user interface object by a second predetermined portion of the user while detecting the input directed at the user interface object by the first predetermined portion of the user; on the user interface object while simultaneously detecting the input and the second input directed to the user interface object. the simulated shadow indicating an interaction of the first predetermined portion of the user with the user interface object; 166. The method of any one of claims 153 to 165, further comprising simultaneously displaying the second predetermined portion of the user relative to the user interface object and a second simulated shadow that is indicative of an interaction with the user interface object.

167. 167. The method of any one of claims 153 to 166, wherein the simulated shadow indicates how much movement of the first predetermined part of the user is required to engage the user interface object.

168. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising: displaying a user interface object via said display generation component; detecting, while displaying the user interface object, an input directed at the user interface object by a first predetermined portion of a user of the electronic device via the one or more input devices; 12. An electronic device comprising: instructions for displaying, via the display generation component, a simulated shadow displayed on the user interface object while detecting the input directed at the user interface object, the simulated shadow having an appearance based on a position of an element relative to the user interface object that indicates an interaction with the user interface object.

169. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the electronic device, cause the electronic device to: displaying a user interface object via the display generation component; detecting, while displaying the user interface object, an input directed at the user interface object by a first predetermined portion of a user of the electronic device via the one or more input devices; and displaying, via the display generation component, a simulated shadow displayed on the user interface object while detecting the input directed at the user interface object, the simulated shadow having an appearance based on a position of an element relative to the user interface object that is indicative of an interaction with the user interface object.

170. 1. An electronic device comprising: one or more processors; Memory and means for displaying user interface objects via said display generation component; detecting, while displaying the user interface object, an input directed at the user interface object by a first predetermined portion of a user of the electronic device via the one or more input devices; and displaying, via the display generation component, a simulated shadow displayed on the user interface object while detecting the input directed at the user interface object, wherein the simulated shadow has an appearance based on a position of an element relative to the user interface object that indicates an interaction with the user interface object.

171. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: means for displaying user interface objects via said display generation component; detecting, while displaying the user interface object, an input directed at the user interface object by a first predetermined portion of a user of the electronic device via the one or more input devices; and displaying, via the display generation component, a simulated shadow displayed on the user interface object while detecting the input directed at the user interface object, wherein the simulated shadow has an appearance based on a position of an element relative to the user interface object that indicates an interaction with the user interface object.

172. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions to perform the method of any one of claims 153 to 167.

173. 168. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 153 to 167.

174. 1. An electronic device comprising: one or more processors; Memory and and means for performing the method of any one of claims 153 to 167.

175. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: and means for performing the method of any one of claims 153 to 167.

176. 1. A method comprising: An electronic device in communication with a display generation component and one or more input devices, displaying, via the display generation component, a user interface including separate regions including a first user interface element and a second user interface element; detecting, while displaying the user interface, a first input directed at the first user interface element within the discrete region via the one or more input devices; In response to detecting the first input directed at the first user interface element, modifying an appearance of the first user interface element to indicate that a further input directed at the first user interface element will cause selection of the first user interface element; detecting a second input via the one or more input devices while displaying the first user interface element in the altered appearance; In response to detecting the second input, In response to determining that the second input includes a movement corresponding to a movement away from the first user interface element, in accordance with determining that the movement corresponds to movement within the discrete region of the user interface, forgoing selection of the first user interface element and altering an appearance of the second user interface element to indicate that further input directed at the second user interface element will cause selection of the second user interface element; and refraining from selecting the first user interface element without changing the appearance of the second user interface element in accordance with a determination that the movement corresponds to movement in a first direction outside the individual region of the user interface.

177. in response to detecting the second input and in accordance with a determination that the movement corresponds to movement outside the discrete region of the user interface in a second direction; forgoing selection of the first user interface element in accordance with a determination that the first input includes input provided by a predetermined portion of a user and that the predetermined portion of the user is farther than a threshold distance from a location corresponding to the first user interface element; 177. The method of claim 176, further comprising: selecting the first user interface element in accordance with the second input in accordance with a determination that the first input includes an input provided by the predetermined portion of the user while the predetermined portion of the user is closer than the threshold distance from the location corresponding to the first user interface element.

178. 178. The method of claim 176 or 177, wherein the first input includes an input provided by a predetermined portion of a user, and selection of the first user interface element is withheld in accordance with the determination that the movement of the second input corresponds to movement outside the discrete area of ​​the user interface in the first direction, regardless of whether the predetermined portion of the user is closer or farther than a threshold distance from a location corresponding to the first user interface element in the first input.

179. detecting, while displaying the user interface, a third input via the one or more input devices directed to a third user interface element in the discrete region, the third user interface element being a slider element, the third input including a moving portion for controlling the slider element; in response to detecting the third input directed at the third user interface element, modifying an appearance of the third user interface element to indicate that a further input directed at the third user interface element will cause further control of the third user interface element, and updating the third user interface element according to the moving portion of the third input; detecting a fourth input while displaying the third user interface element with the changed appearance and while updating the third user interface element according to the moved portion of the third input; In response to detecting the fourth input, In response to determining that the fourth input includes a movement corresponding to a movement away from the third user interface element, maintaining the altered appearance of the third user interface element to indicate that further input directed at the third user interface element will cause further control of the third user interface element; and 179. The method of any one of claims 176 to 178, further comprising: updating the third user interface element according to the movement of the fourth input, regardless of whether the movement of the fourth input corresponds to movement outside the separate region of the user interface.

180. The moving portion of the third input includes an input provided by a predetermined portion of a user having a discrete size, and updating the third user interface element according to the moving portion of the third input includes: updating the third user interface element by a first amount determined based on the first velocity of the predetermined portion of the user and the respective magnitude of the movement portion of the third input in accordance with a determination that the predetermined portion of the user moved at a first velocity during the movement portion of the third input; 180. The method of claim 179, comprising: in accordance with a determination that the predetermined portion of the user moved at a second speed faster than the first speed during the movement portion of the third input, updating the third user interface element by a second amount greater than the first amount, determined based on the second speed of the predetermined portion of the user and the individual magnitudes of the movement portion of the third input, wherein for the individual magnitudes of the movement portion of the third input, the second amount of movement of the third user interface element is greater than the first amount of movement of the third user interface element.

181. the movement of the second input is provided by a discrete movement of a predetermined part of a user; In accordance with a determination that the discrete region of the user interface has a first size, the movement of the second input corresponds to movement outside the discrete region of the user interface in accordance with a determination that the discrete movement of the predetermined portion of the user has a first magnitude; 181. A method according to any one of claims 176 to 180, wherein in accordance with a determination that the individual area of ​​the user interface has a second size different from the first size, the movement of the second input corresponds to movement outside the individual area of ​​the user interface in accordance with a determination that the individual movement of the predetermined portion of the user has the first size.

182. detecting the first input includes detecting that a gaze of a user of the electronic device is directed toward the first user interface element; detecting the second input includes detecting the movement corresponding to movement away from the first user interface element and detecting that the user's gaze is no longer directed toward the first user interface element; 182. The method of any one of claims 176 to 181, wherein forgoing the selection of the first user interface element and changing the appearance of the second user interface element to indicate that further input directed at the second user interface element will cause the second user interface element to be selected is performed while the user's gaze is not directed at the first user interface element.

183. Detecting the first input includes detecting a gaze of a user of the electronic device directed toward the discrete region of the user interface, the method further comprising: detecting, while displaying the first user interface element in the altered appearance and prior to detecting the second input, that the user's gaze is directed via the one or more input devices to a second region of the user interface that is different from the distinct region; In response to detecting that the line of sight of the user is directed toward the second area, 183. The method of any one of claims 176 to 182, further comprising: in accordance with a determination that the second region includes a third user interface element, modifying an appearance of the third user interface element to indicate that further input directed at the third user interface element will cause an interaction with the third user interface element.

184. 184. The method of any one of claims 176 to 183, wherein the first input comprises movement of a predetermined portion of a user of the electronic device in space within an environment of the electronic device, the predetermined portion of the user not in contact with a physical input device.

185. 185. The method of any one of claims 176 to 184, wherein the first input comprises a pinch gesture performed by a hand of a user of the electronic device.

186. 186. The method of any one of claims 176 to 185, wherein the first input comprises movement of a finger of a hand of a user of the electronic device through space within an environment of the electronic device.

187. In response to detecting the second input, In response to the determination that the second input includes a movement corresponding to a movement away from the first user interface element, 187. The method of any one of claims 176 to 186, further comprising, in accordance with the determination that the movement corresponds to movement within the individual region of the user interface, changing the appearance of the first user interface element to indicate that further input is no longer directed to the first user interface element.

188. In response to determining that the second input is provided by a predetermined portion of a user of the electronic device and that the predetermined portion of the user is farther than a threshold distance from a location corresponding to the individual region, the movement of the second input corresponds to movement within the discrete region of the user interface if the second input satisfies one or more first criteria, and the movement of the second input corresponds to movement outside the discrete region of the user interface if the second input does not satisfy the one or more first criteria; in response to determining that the second input is provided by the predetermined portion of the user of the electronic device and that the predetermined portion of the user is closer than the threshold distance from the location corresponding to the individual region; 188. The method of any one of claims 176 to 187, wherein the movement of the second input corresponds to movement within the discrete region of the user interface if the second input satisfies one or more second criteria different from the first criteria, and the movement of the second input corresponds to movement outside the discrete region of the user interface if the second input does not satisfy the one or more second criteria.

189. modifying the appearance of the first user interface element to indicate that further input directed at the first user interface element will cause selection of the first user interface element includes moving the first user interface element away from the user's viewpoint within the three-dimensional environment; 189. The method of any one of claims 176 to 188, wherein altering the appearance of the second user interface element to indicate that further input directed at the second user interface element will cause selection of the second user interface element comprises moving the second user interface element away from the user's viewpoint within the three-dimensional environment.

190. detecting a third input via the one or more input devices while displaying the second user interface element in the altered appearance to indicate that a further input directed at the second user interface element will cause selection of the second user interface element; In response to detecting the third input, 190. The method of any one of claims 176 to 189, further comprising: in accordance with a determination that the third input corresponds to a further input directed at the second user interface element, selecting the second user interface element in accordance with the third input.

191. prior to detecting the first input, selecting the first user interface element requires an input associated with a first magnitude; the first input includes an input of a second magnitude less than the first magnitude; prior to detecting the second input, selecting the second user interface element requires an input associated with a third magnitude; 191. The method of any one of claims 176 to 190, wherein in response to detecting the second input, selection of the second user interface element requires a further input associated with the third magnitude that is smaller than the second magnitude of the first input.

192. the first input includes a selection initiation portion followed by a second portion, and the appearance of the first user interface element is changed to indicate that a further input directed at the first user interface element will cause selection of the first user interface element in accordance with the first input including the selection initiation portion; the appearance of the second user interface element is changed to indicate that a further input directed at the second user interface element will cause selection of the second user interface element without the electronic device detecting another selection initiation portion after the selection initiation portion included in the first input, the method comprising: detecting a third input directed at the second user interface element via the one or more input devices while displaying the second user interface element without the altered appearance; In response to detecting the third input, In accordance with determining that the third input includes the selection initiation portion, modifying the appearance of the second user interface element to indicate that a further input directed at the second user interface element will cause selection of the second user interface element; 192. The method of any one of claims 176 to 191, further comprising: refraining from altering the appearance of the second user interface element in accordance with a determination that the third input does not include the selection start portion.

193. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising: displaying, via a display generation component, a user interface including separate regions including the first user interface element and the second user interface element; detecting, while displaying the user interface, a first input directed at the first user interface element within the discrete region via one or more input devices; In response to detecting the first input directed at the first user interface element, modifying an appearance of the first user interface element to indicate that a further input directed at the first user interface element will cause selection of the first user interface element; detecting a second input via the one or more input devices while displaying the first user interface element in the modified appearance; In response to detecting the second input, In response to determining that the second input includes a movement corresponding to a movement away from the first user interface element, pursuant to determining that the movement corresponds to movement within the discrete region of the user interface, forgoing selection of the first user interface element and altering an appearance of the second user interface element to indicate that further input directed at the second user interface element will cause selection of the second user interface element; 1. An electronic device comprising: instructions for forgoing selection of the first user interface element without changing the appearance of the second user interface element in accordance with a determination that the movement corresponds to movement in a first direction outside the discrete region of the user interface.

194. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the electronic device, cause the electronic device to: displaying, via a display generation component, a user interface including separate regions including a first user interface element and a second user interface element; detecting, while displaying the user interface, a first input directed at the first user interface element within the discrete region via one or more input devices; In response to detecting the first input directed at the first user interface element, modifying an appearance of the first user interface element to indicate that a further input directed at the first user interface element will cause selection of the first user interface element; detecting a second input via the one or more input devices while displaying the first user interface element in the altered appearance; In response to detecting the second input, In response to determining that the second input includes a movement corresponding to a movement away from the first user interface element, in accordance with determining that the movement corresponds to movement within the discrete region of the user interface, forgoing selection of the first user interface element and altering an appearance of the second user interface element to indicate that further input directed at the second user interface element will cause selection of the second user interface element; and forgoing selection of the first user interface element without changing the appearance of the second user interface element in accordance with a determination that the movement corresponds to movement in a first direction outside the discrete region of the user interface.

195. 1. An electronic device comprising: one or more processors; Memory and means for displaying, via a display generation component, a user interface including separate regions including a first user interface element and a second user interface element; means for detecting, while displaying the user interface, a first input directed to the first user interface element within the discrete region via one or more input devices; means for, in response to detecting the first input directed at the first user interface element, altering an appearance of the first user interface element to indicate that a further input directed at the first user interface element will cause selection of the first user interface element; means for detecting a second input via the one or more input devices while displaying the first user interface element in the modified appearance; In response to detecting the second input, In response to determining that the second input includes a movement corresponding to a movement away from the first user interface element, pursuant to determining that the movement corresponds to movement within the discrete region of the user interface, forgoing selection of the first user interface element and altering an appearance of the second user interface element to indicate that further input directed at the second user interface element will cause selection of the second user interface element; means for forgoing selection of the first user interface element without changing the appearance of the second user interface element in accordance with a determination that the movement corresponds to movement in a first direction outside the discrete region of the user interface.

196. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: means for displaying, via a display generation component, a user interface including separate regions including a first user interface element and a second user interface element; means for detecting, while displaying the user interface, a first input directed to the first user interface element within the discrete region via one or more input devices; means for, in response to detecting the first input directed at the first user interface element, altering an appearance of the first user interface element to indicate that a further input directed at the first user interface element will cause selection of the first user interface element; means for detecting a second input via the one or more input devices while displaying the first user interface element in the modified appearance; In response to detecting the second input, In response to determining that the second input includes a movement corresponding to a movement away from the first user interface element, pursuant to determining that the movement corresponds to movement within the discrete region of the user interface, forgoing selection of the first user interface element and altering an appearance of the second user interface element to indicate that further input directed at the second user interface element will cause selection of the second user interface element; means for refraining from selection of the first user interface element without changing the appearance of the second user interface element in accordance with a determination that the movement corresponds to movement in a first direction outside the individual region of the user interface.

197. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions to perform the method of any one of claims 176 to 192.

198. 193. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 176 to 192.

199. 1. An electronic device comprising: one or more processors; Memory and and means for performing the method of any one of claims 176 to 192.

200. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: and means for performing the method of any one of claims 176 to 192.

Citation Information

Patent Citations

  • Handheld equipment, and control method and program

    JP2015170213A

  • Input device

    JP2019105967A

  • Display control device, display control method, and program

    WO2018074054A1

Cited By

  • Game machine

    JP2025158133A