Method for interacting with user interface based on attention

By displaying the gaze virtual objects and selecting operations based on user attention, optimizing the user interface of virtual reality and augmented reality systems, the inefficiency problem in existing systems is solved, achieving more efficient user interaction and energy savings.

CN120266083APending Publication Date: 2025-07-04APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380081314.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-04
Filing Date
2023-09-22
Publication Date
2025-07-04

Smart Images

  • Figure CN120266083A_ABST
    Figure CN120266083A_ABST
Patent Text Reader

Abstract

A gaze virtual object is displayed, the gaze virtual object being selectable based on attention directed to the gaze virtual object to perform an operation associated with the selectable virtual object. An indication of the user's attention is displayed. A magnified view of the region of the user interface is displayed. The value of the slider element is adjusted based on the attention of the user. The user interface elements are moved at respective rates based on the attention of the user. Text is entered into the text entry field in response to the verbal input. The value of the value selection user interface object is updated based on the attention of the user. Movement of a virtual object is facilitated based on direct touch interaction. User input is facilitated to display a selected refined user interface object. When the criteria are met, a visual indicator is displayed that indicates a progress towards the selection of the virtual object.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 377,024, filed on September 24, 2022; U.S. Provisional Application No. 63 / 503,138, filed on May 18, 2023; U.S. Provisional Application No. 63 / 506,080, filed on June 3, 2023; and U.S. Provisional Application No. 63 / 506,124, filed on June 4, 2023, the entire contents of which are incorporated herein by reference for all purposes. Technical Field

[0003] The present invention generally relates to computer systems that provide computer-generated experiences, including but not limited to electronic devices that provide virtual reality and mixed reality experiences via a display. Background Art

[0004] In recent years, the development of computer systems for augmented reality has increased significantly. Example augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays) for computer systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Example virtual elements include virtual objects such as digital images, videos, text, icons, and control elements (such as buttons and other graphics). Summary of the Invention

[0005] Some methods and interfaces for interacting with an environment that includes at least some virtual elements (e.g., an application, an augmented reality environment, a mixed reality environment, and a virtual reality environment) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems in which virtual object manipulation is complex, cumbersome, and error-prone impose a significant cognitive burden on the user and detract from the experience of the virtual / augmented reality environment. In addition, these methods take longer than necessary, thereby wasting the energy of the computer system. This latter consideration is particularly important in battery-powered devices.

[0006] Accordingly, there is a need for computer systems with improved methods and interfaces to provide computer-generated experiences to users, such that the interaction between the user and the computer system is more effective and intuitive for the user. Such methods and interfaces optionally supplement or replace conventional methods for providing extended reality experiences. Such methods and interfaces form a more effective human-machine interface by helping the user understand the connection between the inputs provided and the device's response to those inputs, thereby reducing the quantity, degree, and / or nature of the inputs from the user.

[0007] The above-mentioned deficiencies and other problems associated with the user interface of a computer system are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a laptop computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a watch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touch screen" or "touch screen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, in addition to the display generation component, the computer system also has one or more output devices, which include one or more haptic output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, a memory, and one or more modules, a program or set of instructions stored in the memory for performing multiple functions. In some embodiments, the user interacts with the GUI through contact and gestures of a stylus and / or finger on a touch-sensitive surface, movement of the user's eyes and hands in space relative to the GUI (and / or the computer system) or the user's body (as captured by cameras and other motion sensors), and / or voice input (as captured by one or more audio input devices). In some embodiments, the functions performed through the interaction optionally include image editing, drawing, presentation, word processing, spreadsheet creation, playing games, making and receiving phone calls, video conferencing, sending and receiving emails, instant messaging, test support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. The executable instructions for performing these functions are optionally included in a transient and / or non-transitory computer-readable storage medium or other computer program product configured to be executed by one or more processors.

[0008] There is a need for electronic devices having improved methods and interfaces for interacting with content in a three-dimensional environment. Such methods and interfaces can supplement or replace conventional methods for interacting with content in a three-dimensional environment. Such methods and interfaces reduce the amount, degree, and / or nature of input from the user and result in a more efficient human-machine interface. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charges.

[0009] In some embodiments, a computer system displays a gaze virtual object that is selectable based on attention directed to the gaze virtual object to perform an operation associated with an optional virtual object. In some embodiments, the computer system displays an indication of the user's attention. In some embodiments, the computer system displays a magnified view of a region of the user interface. In some embodiments, the computer system adjusts the value of a slider element based on the user's attention. In some embodiments, the computer system moves a user interface element (e.g., a thumb) in the user interface (e.g., a slider element) at a corresponding rate based on the user's attention. In some embodiments, the computer system enters text into a text entry field in response to speech input. In some embodiments, the computer system updates the value of a value selection user interface object based on the user's attention. In some embodiments, the computer system facilitates direct touch interaction in a three-dimensional environment. In some embodiments, the computer system facilitates direct touch interaction with content in a three-dimensional environment. In some embodiments, the computer system promotes movement of a virtual object relative to the user's viewpoint based on direct touch interaction in a three-dimensional environment. In some embodiments, the computer system facilitates user input for displaying a selection refinement user interface object in a three-dimensional environment. In some embodiments, the computer system displays a visual indicator indicating progress towards selecting an optional virtual object when certain criteria are met.

[0010] Note that the various embodiments described above can be combined with any other embodiment described herein. The features and advantages described in this specification are not exhaustive, and in particular, many additional features and advantages will be apparent to those of ordinary skill in the art from the drawings, the specification, and the claims. Additionally, it should be noted that the language used in this specification has been selected for readability and guidance purposes in principle, and may not have been selected to depict or define the subject matter of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To better understand the various described embodiments, reference should be made to the following detailed description in conjunction with the following drawings, in which like reference numerals indicate corresponding parts in all the drawings.

[0012] Figure 1A is a block diagram showing an operating environment of a computer system for providing an XR experience according to some embodiments.

[0013] Figures 1B to 1P is for providing an XR experience in the Figure 1A operating environment of a computer system example.

[0014] Figure 2 is a block diagram showing a controller of a computer system configured to manage and coordinate a user's XR experience according to some embodiments.

[0015] Figure 3 is a block diagram of a display generation component of a computer system configured to provide a visual component of an XR experience to a user according to some embodiments.

[0016] Figure 4 is a block diagram of a hand tracking unit of a computer system configured to capture a user's gesture input according to some embodiments.

[0017] Figure 5 is a block diagram of an eye tracking unit of a computer system configured to capture a user's gaze input according to some embodiments.

[0018] Figure 6 is a flowchart of a flash-assisted gaze tracking pipeline according to some embodiments.

[0019] Figures 7A to 7O illustrates an example of a computer system that displays a gaze virtual object according to some embodiments, where the gaze virtual object is selectable based on the attention directed to it to perform an operation associated with an optional virtual object.

[0020] Figures 8A to 8I is a flowchart of an exemplary method of displaying a gaze virtual object according to some embodiments, where the gaze virtual object is selectable based on the attention directed to it to perform an operation associated with an optional virtual object.

[0021] Figures 9A to 9J is a flowchart of an exemplary method of displaying a gaze virtual object according to some embodiments, where the gaze virtual object is selectable based on the attention directed to it to perform an operation associated with an optional virtual object.

[0022] Figures 10A to 10G is a flowchart of an exemplary method of displaying a gaze virtual object according to some embodiments, where the gaze virtual object is selectable based on the attention directed to it to perform an operation associated with an optional virtual object.

[0023] Figures 11A to 11C illustrates an example of a computer system that displays an indication of a user's attention according to some embodiments.

[0024] Figures 12A to 12L is a flowchart of a method of displaying an indication of a user's attention according to some embodiments.

[0025] Figures 13A to 13C illustrates an example of a first computer system that displays a magnified view of a region of a user interface according to some embodiments.

[0026] Figures 14A to 14E A flowchart of a method for showing an enlarged view of a region of a display user interface according to some embodiments.

[0027] Figures 15A to 15F An example of a computer system for adjusting the value of a slider element based on a user's attention according to some embodiments is shown.

[0028] Figures 16A to 16H A flowchart of a method for adjusting the value of a slider element based on a user's attention according to some embodiments.

[0029] Figures 17A to 17G An example of a computer system for moving a user interface element (e.g., a thumb) in a user interface (e.g., a slider element) at a corresponding rate based on a user's attention according to some embodiments is shown.

[0030] Figures 18A to 18F A flowchart of a method for moving a user interface element (e.g., a thumb) in a user interface (e.g., a slider element) at a corresponding rate based on a user's attention according to some embodiments.

[0031] Figures 19A to 19J An example of a computer system for showing an indication of a user's attention according to some embodiments is shown.

[0032] Figures 20A to 20E A flowchart of a method for showing an indication of a user's attention according to some embodiments.

[0033] Figures 21A to 21I An example of a computer system for entering text into a text entry field in response to speech input according to some embodiments is shown.

[0034] Figures 22A to 22H A flowchart of a method for entering text into a text entry field in response to speech input according to some embodiments.

[0035] Figures 23A to 23M An example of a computer system for updating the value of a value selection user interface object based on a user's attention according to some embodiments is shown.

[0036] Figures 24A to 24H A flowchart of a method for showing a value selection user interface object according to some embodiments, the value selection user interface object being selectable based on attention directed to the value selection user interface object to navigate among options for the value selection user interface object and select a value from these options.

[0037] Figures 25A to 25JIllustrates an example of a computer system that facilitates touch interactions with one or more virtual objects in a three-dimensional environment according to some embodiments.

[0038] Figures 26A to 26J Is a flowchart showing a method for facilitating direct touch interactions with content in a three-dimensional environment according to some embodiments.

[0039] Figures 27A to 27H Is a flowchart showing a method for facilitating direct touch interactions with content in a three-dimensional environment according to some embodiments.

[0040] Figures 28A to 28H Is a flowchart showing a method for displaying an indication of a user's attention according to some embodiments.

[0041] Figures 29A to 29G Illustrates an example of a computer system that facilitates the movement of a virtual object relative to a user's viewpoint based on a direct touch interaction in a three-dimensional environment according to some embodiments.

[0042] Figures 30A to 30I Is a flowchart showing a method for facilitating the movement of a virtual object relative to a user's viewpoint based on a direct touch interaction in a three-dimensional environment according to some embodiments.

[0043] Figures 31A to 31J Illustrates an example of a computer system that facilitates user input for displaying a selection refinement user interface object in a three-dimensional environment according to some embodiments.

[0044] Figures 32A to 32H Is a flowchart showing a method for facilitating user input for displaying a selection refinement user interface object in a three-dimensional environment according to some embodiments.

[0045] Figures 33A to 33H Illustrates an example of a computer system that displays a visual indicator indicating progress towards selecting an optional virtual object when certain criteria are met according to some embodiments.

[0046] Figure 34A Is a flowchart showing a method for displaying a visual indicator indicating progress towards selecting an optional virtual object when certain criteria are met according to some embodiments. Detailed Description

[0047] According to some embodiments, the present disclosure relates to a user interface for providing an extended reality (XR) experience to a user.

[0048] The systems, methods, and GUIs described herein improve user interface interactions with virtual / augmented reality environments in a variety of ways.

[0049] In some embodiments, the computer system displays a virtual object for watching, which is selectively used to perform operations associated with an optional virtual object based on the attention directed to the virtual object for watching. In some embodiments, the computer system displays an indication of the user's attention. In some embodiments, the computer system displays an enlarged view of the area of ​​the user interface. In some embodiments, the computer system adjusts the value of the slider element based on the user's attention. In some embodiments, the computer system moves the user interface element (e.g., thumb) in the user interface (e.g., slider element) at a corresponding rate based on the user's attention. In some embodiments, the computer system enters text into the text entry field in response to speech input. In some embodiments, the computer system updates the value of the user interface object based on the user's attention to select the value of the user interface object. In some embodiments, the computer system promotes the movement of the virtual object relative to the user's viewpoint according to the direct touch interaction in the three-dimensional environment. In some embodiments, the computer system promotes the movement of the virtual object relative to the user's viewpoint according to the direct touch interaction in the three-dimensional environment. In some embodiments, the computer system promotes the user input for displaying a selection to refine the user interface object in a three-dimensional environment. In some embodiments, the computer system displays a visual indicator indicating the progress towards selecting an optional virtual object when certain criteria are met.

[0050] Figures 1A to 6 A description of an example computer system for providing an XR experience to a user is provided (such as described below with respect to methods 800 , 900 , 1000 , 1200 , 1400 , 1600 , 1800 , 2000 , 2200 , 2400 , 2600 , 2700 , 2800 , 3000 , 3200 , and / or 3400 ). Figures 7A to 7O An example of a computer system displaying a gaze virtual object that is selectable based on attention directed toward the gaze virtual object to perform an operation associated with the selectable virtual object is shown according to some embodiments. Figures 8A to 8I , Figures 9A to 9J as well as Figures 10A to 10G is a flowchart illustrating an exemplary method of displaying a gaze virtual object according to some embodiments, wherein the gaze virtual object is selectable based on attention directed to the gaze virtual object to perform an operation associated with the selectable virtual object. Figures 7A to 7O The user interface in Figures 8A to 8I , Figures 9A to 9J as well as Figures 10A to 10G process. Figures 11A to 11C Example techniques are shown for displaying an indication of a user's attention according to some embodiments. Figures 12A to 12L is a flowchart of a method of displaying an indication of a user's attention according to various embodiments. Figures 28A to 28HFlowchart of a method for indicating a user's attention according to various embodiments. Figures 11A to 11C The user interface in Figures 12A to 12L is used to illustrate the Figures 28A to 28H processes in Figures 13A to 13C An example technique showing a magnified view of a region of a display user interface according to some embodiments. Figures 14A to 14E Flowchart of a method for magnifying a view of a region of a display user interface according to various embodiments. Figures 13A to 13C The user interface in Figures 14A to 14E is used to illustrate the Figures 15A to 15F processes in Figures 16A to 16H An example technique showing adjusting the value of a slider element based on a user's attention according to some embodiments. Figures 15A to 15F The user interface in Figures 16A to 16H is used to illustrate the Figures 17A to 17G processes in Figures 18A to 18F An example technique showing moving a user interface element (e.g., a thumb) in a user interface (e.g., a slider element) at a corresponding rate based on a user's attention according to some embodiments. Figures 17A to 17G The user interface in Figures 18A to 18F is used to illustrate the Figures 19A to 19J processes in Figures 20A to 20E An example technique showing an indication of a user's attention according to some embodiments. Figures 19A to 19J Flowchart of a method for indicating a user's attention according to various embodiments. Figure 20A in Figure 20E is used to illustrate the Figures 21A to 21I processes in Figures 22A to 22H An example technique showing entering text into a text entry field in response to speech input according to some embodiments. Figures 21A to 21I Flowchart of a method for entering text into a text entry field in response to speech input according to various embodiments. Figures 22A to 22H The user interface in Figures 23A to 23M is used to illustrate the Figures 23A to 23M processes in Figures 24A to 24H An example technique showing updating the value of a value selection user interface object based on a user's attention according to some embodiments. Figures 25A to 25J An example technique showing facilitating touch interaction with one or more virtual objects in a three-dimensional environment according to some embodiments. Figures 26A to 26Jis a flowchart of a method for facilitating direct touch interaction with content in a three-dimensional environment according to some embodiments. Figures 25A to 25J The user interface in Figures 26A to 26J is used to illustrate the Figures 27A to 27H is a flowchart of a method for facilitating direct touch interaction with content in a three-dimensional environment according to some embodiments. Figures 25A to 25J The user interface in Figures 27A to 27H is used to illustrate the Figures 29A to 29G illustrates an example technique for facilitating movement of a virtual object relative to a user's viewpoint based on direct touch interaction in a three-dimensional environment according to some embodiments. Figures 30A to 30I is a flowchart of a method for facilitating movement of a virtual object relative to a user's viewpoint based on direct touch interaction in a three-dimensional environment according to various embodiments. Figures 29A to 29G The user interface in Figures 30A to 30I is used to illustrate the Figures 31A to 31J illustrates an example technique for facilitating user input for displaying a selection refinement user interface object in a three-dimensional environment according to some embodiments. Figures 32A to 32H is a flowchart of a method for facilitating user input for displaying a selection refinement user interface object in a three-dimensional environment according to some embodiments. Figures 31A to 31J The user interface in Figures 32A to 32H is used to illustrate the Figures 33A to 33H illustrates an example technique for displaying a visual indicator indicating progress towards selecting an optional virtual object when certain criteria are met according to some embodiments. Figure 34A is a flowchart of a method for displaying a visual indicator indicating progress towards selecting an optional virtual object when certain criteria are met according to various embodiments. Figures 33A to 33H The user interface in Figure 34A is used to illustrate the

[0051] The processes described below enhance the operability of a device and make the user-device interface more efficient through various techniques (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), including providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional display controls, performing an operation when a set of conditions has been met without further user input, improving privacy and / or security, providing a more diverse, detailed, and / or realistic user experience while saving storage space, and / or additional techniques. These techniques also reduce power usage and extend the battery life of the device by enabling the user to use the device faster and more efficiently. Saving battery power, and thus weight, improves the ergonomics of the device. These techniques also enable real-time communication, allow for the use of fewer and / or less precise sensors, resulting in a more compact, lighter, and cheaper device, and enable the device to be used under various lighting conditions. These techniques reduce energy usage, thereby reducing the heat emitted by the device, which is particularly important for wearable devices, where it can become uncomfortable for the user to wear the device if too much heat is generated within the operating parameters of the device components.

[0052] In addition, in a method where one or more of the steps described herein depend on one or more conditions being met, it should be understood that the method can be repeated in multiple iterations such that, during the repeated process, all of the conditions that determine the steps in the method are met in different iterations of the method. For example, if a method requires performing a first step (if a condition is met) and a second step (if the condition is not met), then one of ordinary skill in the art will know to repeat the stated steps until both the condition is met and the condition is not met (in no particular order). Thus, a method described as having one or more steps that depend on one or more conditions being met can be rewritten as a method that repeats until each condition described in the method has been met. However, this does not require the system or computer-readable medium to state that the system or computer-readable medium includes instructions for performing conditional operations based on the satisfaction of the corresponding one or more conditions and is thus capable of determining whether the possible conditions have been met without explicitly repeating the steps of the method until all of the conditions that determine the steps in the method have been met. One of ordinary skill in the art will also understand that, similar to a method with conditional steps, a system or computer-readable storage medium can repeat the steps of the method as many times as needed to ensure that all conditional steps have been performed.

[0053] In some embodiments, as Figure 1AAs shown, an XR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a haptic sensor, an orientation sensor, a proximity sensor, a temperature sensor, a position sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., household appliances, wearable devices, etc.). In some embodiments, one or more of the input devices 125, output devices 155, sensors 190, and peripheral devices 195 are integrated with the display generation component 120 (e.g., in a head-mounted device or a handheld device).

[0054] When describing an XR experience, various terms are used to distinctively refer to several related but different environments that a user can sense and / or with which the user can interact (e.g., interact using inputs detected by the computer system 101 that generates the XR experience, where these inputs cause the computer system that generates the XR experience to generate audio, visual, and / or haptic feedback corresponding to the various inputs provided to the computer system 101). The following is a subset of these terms:

[0055] Physical environment: The physical environment refers to the physical world that people can sense and / or interact with without the help of an electronic system. Physical environments such as a physical park include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment, such as through vision, touch, hearing, taste, and smell.

[0056] Extended Reality: In contrast, an extended reality (XR) environment is a fully or partially simulated environment in which people sense and / or interact via an electronic system. In XR, a subset of a person's physical movements or their representation is tracked, and in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner consistent with at least one physical law. For example, an XR system can detect a person's head rotation, and in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. In some cases (e.g., for accessibility reasons), the adjustment of the characteristics of virtual objects in the XR environment can be made in response to a representation of a physical movement (e.g., a voice command). A person can use any of their senses to sense and / or interact with XR objects, including vision, hearing, touch, taste, and smell. For example, a person can sense and / or interact with an audio object that creates a 3D or spatial audio environment that provides the perception of point audio sources in 3D space. Additionally, an audio object can enable audio transparency that selectively introduces ambient sounds from the physical environment with or without computer-generated audio. In some XR environments, a person can sense and / or interact only with audio objects.

[0057] Examples of XR include virtual reality and mixed reality.

[0058] Virtual Reality: A virtual reality (VR) environment is a simulated environment that is designed to be fully computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects with which a person can sense and / or interact. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with the virtual objects in the VR environment by way of a simulation of the person's presence within the computer-generated environment and / or by way of a simulation of a subset of the person's physical movements within the computer-generated environment.

[0059] Mixed Reality: Compared with a VR environment that is designed to be based entirely on computer-generated sensory inputs, a mixed reality (MR) environment is an analog environment that is designed to include, in addition to computer-generated sensory inputs (e.g., virtual objects), sensory inputs or their representations from the physical environment. On the virtual continuum, an MR environment is any condition between the fully physical environment at one end and the virtual reality environment at the other end, excluding these two ends. In some MR environments, the computer-generated sensory inputs can respond to changes in the sensory inputs from the physical environment. Additionally, some electronic systems for presenting an MR environment can track the position and / or orientation relative to the physical environment so that virtual objects can interact with real objects (i.e., physical items or their representations from the physical environment). For example, the system can cause movements such that a virtual tree appears stationary relative to the physical ground.

[0060] Examples of mixed reality include augmented reality and augmented virtuality.

[0061] Augmented Reality: An augmented reality (AR) environment is a simulated environment in which one or more virtual objects are superimposed over a physical environment or a representation of a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system can be configured to present virtual objects on the transparent or translucent display such that a person using the system perceives the virtual objects superimposed over the physical environment. Alternatively, the system can have an opaque display and one or more imaging sensors that capture images or video of the physical environment, which are representations of the physical environment. The system combines the images or video with virtual objects and presents the combination on the opaque display. A person uses the system to indirectly view the physical environment via the images or video of the physical environment and perceives the virtual objects superimposed over the physical environment. As used herein, the video of the physical environment displayed on the opaque display is referred to as "passthrough video," meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Further alternatively, the system can have a projection system that projects virtual objects into the physical environment, such as as a hologram or on a physical surface, such that a person using the system perceives the virtual objects superimposed over the physical environment. An augmented reality environment is also a simulated environment in which a representation of the physical environment is transformed by computer-generated sensory information. For example, in providing passthrough video, the system can transform one or more sensor images to impose an alternative perspective (e.g., viewpoint) different from the perspective captured by the imaging sensor. As another example, a representation of the physical environment can be transformed by graphically modifying (e.g., magnifying) portions thereof such that the modified portions can be a representative but not a true version of the originally captured image. As yet another example, a representation of the physical environment can be transformed by graphically removing portions thereof or blurring portions thereof.

[0062] Augmented Virtuality: An augmented virtuality (AV) environment is a simulated environment in which a virtual environment or computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory inputs can be representations of one or more characteristics of the physical environment. For example, an AV park can have virtual trees and virtual buildings, but a person's face is a realistic reproduction from an image of a physical person. As another example, a virtual object can adopt the shape or color of a physical item imaged by one or more imaging sensors. As yet another example, a virtual object can adopt a shadow that conforms to the positioning of the sun in the physical environment.

[0063] In an augmented reality, mixed reality, or virtual reality environment, a view of a three-dimensional environment is visible to a user. The view of the three-dimensional environment is typically visible to the user through a virtual viewport via one or more display generation components (e.g., a display or a pair of display modules that provide stereoscopic content to different eyes of the same user), the virtual viewport having a viewport boundary that defines the extent of the three-dimensional environment visible to the user via the one or more display generation components. In some embodiments, the region defined by the viewport boundary is smaller than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size, optical properties, or other physical characteristics of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user's eyes). In some embodiments, the region defined by the viewport boundary is larger than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size, optical properties, or other physical characteristics of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user's eyes). The viewport and the viewport boundary typically move as the one or more display generation components move (e.g., for a head-mounted device as the user's head moves, or for a handheld device such as a tablet or smartphone as the user's hand moves). The user's viewpoint determines the content visible in the viewport, the viewpoint typically specifying a position and orientation relative to the three-dimensional environment, and as the viewpoint moves, the view of the three-dimensional environment will also move in the viewport. For a head-mounted device, the viewpoint is typically based on the position and orientation of the user's head, face, and / or eyes to provide a perceptually accurate view of the three-dimensional environment and an immersive experience when the user is using the head-mounted device. For a handheld or stationary device, the viewpoint moves as the handheld or stationary device moves and / or as the user's positioning relative to the handheld or stationary device changes (e.g., the user moves toward, away from, up, down, right, and / or left). For a device that includes a display generation component with virtual passthrough, the portions of the physical environment visible (e.g., displayed and / or projected) via the one or more display generation components are based on the field of view of one or more cameras that communicate with the display generation component, the one or more cameras typically moving as the display generation component moves (e.g., for a head-mounted device as the user's head moves, or for a handheld device such as a tablet or smartphone as the user's hand moves), because the user's viewpoint moves as the field of view of the one or more cameras moves (and the appearance of one or more virtual objects displayed via the one or more display generation components is updated based on the user's viewpoint (e.g., the display position and pose of the virtual object are updated based on the movement of the user's viewpoint)).For a display generation component with optical passthrough, portions of the physical environment that are visible via one or more display generation components (e.g., optically visible through one or more partial or fully transparent portions of the display generation component) are based on the user's field of view through the partial or fully transparent portion of the display generation component (e.g., for a head-mounted device, moving as the user's head moves, or for a handheld device such as a tablet or smartphone, moving as the user's hand moves), because the user's viewpoint moves as the user's field of view through the partial or fully transparent portion of the display generation component moves (and the appearance of one or more virtual objects is updated based on the user's viewpoint).

[0064] In some embodiments, the representation of the physical environment (e.g., via virtual passthrough or optical passthrough display) may be partially or fully occluded by the virtual environment. In some embodiments, the amount of the virtual environment displayed (e.g., the amount of the physical environment not displayed) is based on the immersion level of the virtual environment (e.g., relative to the representation of the physical environment). For example, increasing the immersion level optionally causes more of the virtual environment to be displayed, replacing and / or occluding more of the physical environment, and decreasing the immersion level optionally causes less of the virtual environment to be displayed, thereby revealing portions of the physical environment that were previously not displayed and / or occluded. In some embodiments, at a particular immersion level, one or more first background objects (e.g., in the representation of the physical environment) are visually de-emphasized (e.g., dimmed, blurred, displayed with increased transparency) more than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, the immersion level includes the associated degree to which virtual content (e.g., virtual environment and / or virtual content) displayed by a computer system occludes background content (e.g., content other than the virtual environment and / or virtual content) around / behind the virtual environment, optionally including the number of items of the displayed background content and / or the displayed visual characteristics of the background content (e.g., color, contrast, and / or opacity), the angular range of the virtual content displayed by a display generation component (e.g., 60 degrees of content displayed at low immersion, 120 degrees of content displayed at medium immersion, or 180 degrees of content displayed at high immersion), and / or the proportion of the field of view displayed by the display generation component occupied by the virtual content (e.g., 33% of the field of view occupied by the virtual content at low immersion, 66% of the field of view occupied by the virtual content at medium immersion, or 100% of the field of view occupied by the virtual content at high immersion). In some embodiments, the background content is included in the background on which the virtual content is displayed (e.g., the background content in the representation of the physical environment). In some embodiments, the background content includes a user interface (e.g., a user interface generated by a computer system corresponding to an application), virtual objects not associated with or not included in the virtual environment and / or virtual content (e.g., files or representations of other users generated by a computer system, etc.), and / or real objects (e.g., passthrough objects representing real objects in the physical environment around the user, which are visible such that they are displayed by the display generation component and / or visible via a transparent or translucent component of the display generation component because the computer system does not occlude / hinder their visibility through the display generation component). In some embodiments, at a low immersion level (e.g., a first immersion level), the background, virtual, and / or real objects are displayed in an unoccluded manner. For example, a virtual environment with a low immersion level is optionally displayed simultaneously with background content, which is optionally displayed at full brightness, color, and / or semi-transparency.In some embodiments, at a higher immersion level (e.g., a second immersion level higher than the first immersion level), background, virtual, and / or real objects are displayed in an occluded manner (e.g., dimmed, blurred, or removed from the display). For example, a corresponding virtual environment with a high immersion level is displayed without simultaneously displaying background content (e.g., in full screen or fully immersive mode). As another example, a virtual environment displayed at a medium immersion level is displayed simultaneously with background content that is dimmed, blurred, or otherwise de-emphasized. In some embodiments, the visual characteristics of background objects vary among the background objects. For example, at a particular immersion level, one or more first background objects are more visually de-emphasized (e.g., dimmed, blurred, and / or displayed with increased transparency) than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, zero immersion or a zero immersion level corresponds to a virtual environment that ceases to be displayed, and instead a representation of the physical environment is displayed (optionally with one or more virtual objects, such as applications, windows, or virtual three-dimensional objects), and the representation of the physical environment is not occluded by the virtual environment. Adjusting the immersion level using physical input elements provides a quick and efficient way to adjust immersion, which enhances the operability of the computer system and makes the user-device interface more efficient.

[0065] Viewpoint-locked virtual objects: When a computer system displays a virtual object at the same position and / or orientation in the user's viewpoint, the virtual object is viewpoint-locked even if the user's viewpoint shifts (e.g., changes). In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the forward direction of the user's head (e.g., when the user looks straight ahead, the user's viewpoint is at least a portion of the user's field of view); thus, without moving the user's head, the user's viewpoint remains fixed even when the user's gaze shifts. In embodiments where the computer system has a display generation component (e.g., a display screen) that is repositionable relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the computer system's display generation component. For example, a viewpoint-locked virtual object that is displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's viewpoint even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the position and / or orientation of the viewpoint-locked virtual object displayed in the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head such that the virtual object is also referred to as a "head-locked virtual object".

[0066] Environment-Locked Visual Objects: When a computer system displays a virtual object at a location and / or orientation within a user's field of view, the virtual object is environment-locked (alternatively, "world-locked"), where the location and / or orientation is based on a location and / or object within a three-dimensional environment (e.g., a physical environment or a virtual environment) (e.g., selected and / or anchored relative to the location and / or object). As the user's field of view moves, the location and / or object within the environment relative to the user's field of view changes, which causes the environment-locked virtual object to be displayed at a different location and / or orientation within the user's field of view. For example, an environment-locked virtual object locked to a tree directly in front of the user is displayed at the center of the user's field of view. When the user's field of view shifts to the right (e.g., the user's head turns to the right) such that the tree is now to the left of center within the user's field of view (e.g., the orientation of the tree within the user's field of view has shifted), the environment-locked virtual object locked to the tree is displayed to the left of center within the user's field of view. In other words, the location and / or orientation at which the environment-locked virtual object is displayed within the user's field of view depends on the location and / or orientation of the location and / or object within the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system anchored to a fixed location and / or object within a physical environment) in order to determine the orientation at which the environment-locked virtual object is displayed within the user's field of view. The environment-locked virtual object may be locked to a stationary portion of the environment (e.g., a floor, wall, table, or other stationary object), or may be locked to a movable portion of the environment (e.g., a vehicle, animal, person, or even a representation of a part of the user's body such as the user's hand, wrist, arm, or foot that moves independently of the user's field of view) such that the virtual object moves as the field of view or that portion of the environment moves to maintain a fixed relationship between the virtual object and that portion of the environment.

[0067] In some embodiments, an environment-locked or view-locked virtual object exhibits lazy follow behavior, which reduces or delays the movement of the environment-locked or view-locked virtual object relative to the movement of a reference point that the virtual object follows. In some embodiments, when exhibiting lazy follow behavior, when a movement of a reference point (e.g., a part of the environment, a view point, or a point fixed relative to the view point, such as a point between 5 cm and 300 cm from the view point) that the virtual object is following is detected, the computer system intentionally delays the movement of the virtual object. For example, when the reference point (e.g., the part of the environment or the view point) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point but moves at a second speed that is slower than the first speed (e.g., until the reference point stops moving or slows down, at which point the virtual object begins to catch up to the reference point). In some embodiments, when the virtual object exhibits lazy follow behavior, the device ignores small movements of the reference point (e.g., ignores movements of the reference point below a threshold movement amount, such as moving 0 degrees to 5 degrees or moving 0 cm to 50 cm). For example, when the reference point (e.g., the part of the environment or the view point to which the virtual object is locked) moves a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed to remain fixed or substantially fixed relative to a different view point or part of the environment than the reference point to which the virtual object is locked), and when the reference point (e.g., the part of the environment or the view point to which the virtual object is locked) moves a second amount that is greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is being displayed to remain fixed or substantially fixed relative to a different view point or part of the environment than the reference point to which the virtual object is locked), and then decreases when the movement amount of the reference point increases above a threshold (e.g., the "lazy follow" threshold), because the virtual object is moved by the computer system to remain fixed or substantially fixed relative to the reference point. In some embodiments, the virtual object remaining substantially fixed in position relative to the reference point includes the virtual object being displayed within a threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 15 cm, 20 cm, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward of the positioning relative to the reference point).

[0068] Hardware: There are many different types of electronic systems that enable a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields integrated with display capabilities, windows integrated with display capabilities, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headsets / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smart phones, tablets, and desktop / laptop computers. A head-mounted system can have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system can be configured to accept an external opaque display (e.g., a smart phone). A head-mounted system can incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system can have a transparent or semi-transparent display instead of an opaque display. The transparent or semi-transparent display can have a medium through which light representing an image is directed to a person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light sources, or any combination of these technologies. The medium can be an optical waveguide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In one embodiment, the transparent or semi-transparent display can be configured to selectively become opaque. A projection-based system can employ retinal projection techniques that project graphical images onto a person's retina. The projection system can also be configured to project virtual objects into the physical environment, such as as a hologram or on a physical surface. In some embodiments, the controller 110 is configured to manage and coordinate a user's XR experience. In some embodiments, the controller 110 includes a suitable combination of software, firmware, and / or hardware. Below with respect to Figure 2Controller 110 is described in more detail. In some embodiments, controller 110 is a computing device that is local or remote relative to scene 105 (e.g., a physical environment). For example, controller 110 is a local server located within scene 105. As another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside of scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., an HMD, a display, a projector, a touch screen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, controller 110 is included within the housing (e.g., a physical enclosure) of one or more of display generation component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors, etc.), one or more input devices of input device 125, one or more output devices of output device 155, one or more sensors of sensor 190, and / or one or more peripheral devices of peripheral device 195, or shares the same physical housing or support structure with one or more of the above devices.

[0069] In some embodiments, display generation component 120 is configured to provide an XR experience (e.g., at least a visual component of the XR experience) to a user. In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. More details regarding Figure 3 display generation component 120 are described below. In some embodiments, the functionality of controller 110 is provided by and / or combined with display generation component 120.

[0070] According to some embodiments, when a user is virtually and / or physically present within scene 105, display generation component 120 provides an XR experience to the user.

[0071] In some embodiments, the display generation component is worn on a part of the user's body (e.g., on his / her head, on his / her hand, etc.). Accordingly, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smart phone or a tablet) configured to present XR content, and the user holds the device having a display facing the user's field of view and a camera facing the scene 105. In some embodiments, the handheld device is optionally placed in a housing worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is an XR chamber, housing, or room configured to present XR content, where the user does not wear or hold the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) can be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing an interaction with XR content triggered based on an interaction occurring in the space in front of a handheld device or a tripod-mounted device can be similarly implemented with an HMD, where the interaction occurs in the space in front of the HMD, and the response to the XR content is displayed via the HMD. Similarly, a user interface showing an interaction with XR content triggered based on the movement of a handheld device or a tripod-mounted device relative to the physical environment (e.g., the scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)) can be similarly implemented with an HMD, where the movement is caused by the movement of the HMD relative to the physical environment (e.g., the scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)).

[0072] Although relevant features of the operating environment 100 are shown in Figure 1A for the sake of brevity and to not obscure more relevant aspects of the example embodiments disclosed herein, various other features are not shown.

[0073] Figures 1A to 1PShows various examples of computer systems for performing methods and providing audio, visual, and / or tactile feedback as part of the user interfaces described herein. In some embodiments, the computer system includes one or more display generation components (e.g., a first display component 1-120a and a second display component 1-120b and / or a first optical module 11.1.1-104a and a second optical module 11.1.1-104b) for displaying to a user of the computer system a representation of virtual elements and / or a physical environment optionally generated based on detected events and / or user input detected by the computer system. The user interface generated by the computer system is optionally corrected by one or more corrective lenses 11.3.2-216, which are optionally removably attached to one or more of the optical modules such that the user interface is more easily viewable by users who would otherwise use glasses or contact lenses to correct their vision. Although many of the user interfaces shown herein show a single view of the user interface, the user interface in the HMD optionally uses two optical modules (e.g., a first display component 1-120a and a second display component 1-120b and / or a first optical module 11.1.1-104a and a second optical module 11.1.1-104b) to display, one optical module for the user's right eye and a different optical module for the user's left eye, and presents slightly different images to the two different eyes to create an illusion of stereoscopic depth. The single view of the user interface is typically a right-eye view or a left-eye view, and the depth effect is explained in the text or using other schematic diagrams or views. In some embodiments, the computer system includes one or more external displays (e.g., a display component 1-108) for displaying to a user of the computer system (when the computer system is not being worn) and / or to others in the vicinity of the computer system the status information of the computer system, which is optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more audio output components (e.g., an electronic component 1-112) for generating audio feedback, which is optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting information about the physical environment of the device (e.g., one or more sensors in the sensor assembly 1-356, and / or Figure 1I )), and this information can be used (optionally in combination with one or more illuminators, such as Figure 1IThe illuminator) generates a digital pass-through image, captures visual media corresponding to the physical environment (e.g., photos and / or videos), or determines the pose (e.g., position and / or orientation) of physical objects and / or surfaces in the physical environment, such that virtual objects can be placed based on the detected pose of the physical objects and / or surfaces. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting hand position and / or movement (e.g., sensor assemblies 1-356 and / or Figure 1I one or more of the sensors therein), which can be used (optionally in combination with one or more illuminators, such as Figure 1I the illuminator 6-124 described therein) to determine when one or more air gestures have been performed. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting eye movement (e.g., Figure 1I the eye tracking and gaze tracking sensors therein), which can be used (optionally in combination with one or more lights, such as Figure 1OThe lights in (11.3.2-110) determine the attention or fixation position and / or fixation movement, which can optionally be used to detect only fixation input based on fixation movement and / or dwell. Combinations of the various sensors described above can be used to determine the user's facial expression and / or hand movement for generating an avatar or representation of the user, such as an anthropomorphic avatar or representation for a real-time communication session, where the avatar has facial expressions, hand movements, and / or body movements based on or similar to the detected facial expressions, hand movements, and / or body movements of the user of the device. The fixation and / or attention information is optionally combined with hand tracking information to determine the interaction between the user and one or more user interfaces based on direct and / or indirect input, such as an air gesture or input using one or more hardware input devices, such as one or more buttons (e.g., the first button 1-128, button 11.1.1-114, the second button 1-132, and / or the dial or button 1-328), a knob (e.g., the first button 1-128, button 11.1.1-114, and / or the dial or button 1-328), a digital crown (e.g., the first button 1-128, button 11.1.1-114, and / or the dial or button 1-328 that can be pressed and twisted or rotated), a touchpad, a touch screen, a keyboard, a mouse, and / or other input devices. One or more buttons (e.g., the first button 1-128, button 11.1.1-114, the second button 1-132, and / or the dial or button 1-328) are optionally used to perform system operations, such as re-centering the content in the three-dimensional environment visible to the user of the device, displaying the main user interface for launching an application, starting a real-time communication session, or initiating the display of a virtual three-dimensional background. The knob or digital crown (e.g., the first button 1-128, button 11.1.1-114, and / or the dial or button 1-328 that can be pressed and twisted or rotated) is optionally rotatable to adjust parameters of the visual content, such as the immersion level of the virtual three-dimensional environment (e.g., the extent to which the virtual content occupies the user's viewport in the three-dimensional environment) or other parameters associated with the three-dimensional environment and the virtual content displayed via the optical module (e.g., the first display component 1-120a and the second display component 1-120b and / or the first optical module 11.1.1-104a and the second optical module 11.1.1-104b).

[0074] Figure 1BShows a front view, a top view, and a perspective view of an example of a head-mounted display (HMD) device 1-100 configured to be worn by a user and provide virtual and augmented reality (VR / AR) experiences. The HMD 1-100 may include a display unit 1-102 or component, an electronic strip assembly 1-104 connected to and extending from the display unit 1-102, and a strap assembly 1-106 fixed to the electronic strip assembly 1-104 at either end. The electronic strip assembly 1-104 and the strap 1-106 may be part of a retention assembly configured to wrap around the user's head to hold the display unit 1-102 against the user's face.

[0075] In at least one example, the strap assembly 1-106 may include a first strap 1-116 configured to wrap around the back of the user's head and a second strap 1-117 configured to extend over the top of the user's head. As shown, the second strap may extend between a first electronic strip 1-105a and a second electronic strip 1-105b of the electronic strip assembly 1-104. The strip assembly 1-104 and the strap assembly 1-106 may be part of a fixation mechanism that extends rearward from the display unit 1-102 and is configured to hold the display unit 1-102 against the user's face.

[0076] In at least one example, the fixation mechanism includes a first electronic strip 1-105a that includes a first proximal end 1-134 coupled to the display unit 1-102 (e.g., the housing 1-150 of the display unit 1-102) and a first distal end 1-136 opposite the first proximal end 1-134. The fixation mechanism may also include a second electronic strip 1-105b that includes a second proximal end 1-138 coupled to the housing 1-150 of the display unit 1-102 and a second distal end 1-140 opposite the second proximal end 1-138. The fixation mechanism may also include a first strap 1-116 and a second strap 1-117. The first strap includes a first end 1-142 coupled to the first distal end 1-136 and a second end 1-144 coupled to the second distal end 1-140, and the second strap extends between the first electronic strip 1-105a and the second electronic strip 1-105b. The strips 1-105a-b and the strap 1-116 may be coupled via a connection mechanism or component 1-114. In at least one example, the second strap 1-117 includes a first end 1-146 coupled to the first electronic strip 1-105a between the first proximal end 1-134 and the first distal end 1-136 and a second end 1-148 coupled to the second electronic strip 1-105b between the second proximal end 1-138 and the second distal end 1-140.

[0077] In at least one example, the first and second electronic strips 1-105a-b comprise plastic, metal, or other structural materials forming the shape of the substantially rigid strips 1-105a-b. In at least one example, the first strip 1-116 and the second strip 1-117 are formed of an elastic flexible material including woven textiles, rubber, and the like. The first strip 1-116 and the second strip 1-117 can be flexible to conform to the shape of the user's head when wearing the HMD 1-100.

[0078] In at least one example, one or more of the first and second electronic strips 1-105a-b can define an internal strip volume and include one or more electronic components disposed within the internal strip volume. In one example, as Figure 1B shown, the first electronic strip 1-105a can include the electronic component 1-112. In one example, the electronic component 1-112 can include a speaker. In one example, the electronic component 1-112 can include a computing component, such as a processor.

[0079] In at least one example, the housing 1-150 defines a first front opening 1-152. The front opening is Figure 1B marked as 1-152 in dashed lines therein because the display component 1-108 is arranged to occlude the first opening 1-152 from the field of view when assembling the HMD 1-100. The housing 1-150 can also define a rear second opening 1-154. The housing 1-150 also defines an internal volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display component 1-108, which can include a front cover and a display screen (shown in other figures) disposed in or across the front opening 1-152 to occlude the front opening 1-152. In at least one example, the display screen of the display component 1-108 and generally the display component 1-108 have a curvature configured to follow the curvature of the user's face. The display screen of the display component 1-108 can be curved as shown to complement the user's facial features and the overall curvature from one side of the face to the other, e.g., from left to right and / or from top to bottom, where the display unit 1-102 is pressed.

[0080] In at least one example, the housing 1-150 may define a first aperture 1-126 between a first opening 1-152 and a second opening 1-154, and a second aperture 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 may further include a first button 1-126 disposed in the first aperture 1-128, and a second button 1-132 disposed in the second aperture 1-130. The first button 1-128 and the second button 1-132 can be pressed through the respective apertures 1-126, 1-130. In at least one example, the first button 1-126 and / or the second button 1-132 can be a twist dial as well as a push button. In at least one example, the first button 1-128 is a pushable and twistable dial button, and the second button 1-132 is a push button.

[0081] Figure 1C A rear perspective view of the HMD 1-100 is shown. The HMD 1-100 may include a light seal 1-110 extending rearwardly around a perimeter of the housing 1-150 from the housing 1-150 of the display assembly 1-108, as shown. The light seal 1-110 may be configured to extend from the housing 1-150 to a user's face, around the user's eyes, to block external light from being visible. In one example, the HMD 1-100 may include a first display assembly 1-120a and a second display assembly 1-120b, which are disposed at or within a second opening 1-154 defined by the housing 1-150 and facing rearward and / or disposed within an internal volume of the housing 1-150 and configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a-b may include a respective display screen 1-122a, 1-122b, which are configured to project light through the second opening 1-154 in a rearward direction toward the user's eyes.

[0082] In at least one example, referring Figure 1B and Figure 1C both, the display assembly 1-108 can be a front forward display assembly including a display screen configured to project light in a first forward direction, and the rear display screens 1-122a-b can be configured to project light in a second rearward direction opposite the first direction. As described above, the light seal 1-110 can be configured to block light external to the HMD 1-100 from reaching the user's eyes, including light projected by the forward display screen of the display assembly 1-108 shown in the Figure 1B front perspective view. In at least one example, the HMD 1-100 may further include a curtain 1-124 that obscures the second opening 1-154 between the housing 1-150 and the rear display assemblies 1-120a-b. In at least one example, the curtain 1-124 can be elastic or at least partially elastic.

[0083] Figure 1B and Figure 1C Any one of the feature portions, components, and / or parts shown (including their arrangements and configurations) may be included, either individually or in any combination, in Figures 1D to 1F any other examples of the devices, feature portions, components, and parts shown and described herein. Similarly, with reference to Figures 1D to 1F any one of the feature portions, components, and / or parts shown or described (including their arrangements and configurations) may be included, either individually or in any combination, in Figure 1B and Figure 1C the examples of the devices, feature portions, components, and parts shown.

[0084] Figure 1D A decomposition view showing an example of the HMD 1-200 is presented, which includes individual parts or components separated according to the modular and selective coupling of these parts. For example, the HMD 1-200 may include a band 1-216, which can be selectively coupled to a first electronic strip 1-205a and a second electronic strip 1-205b. The first fixed strip 1-205a may include a first electronic component 1-212a, and the second fixed strip 1-205b may include a second electronic component 1-212b. In at least one example, the first and second strips 1-205a-b can be removably coupled to the display unit 1-202.

[0085] In addition, the HMD 1-200 may include a light seal 1-210 configured to be removably coupled to the display unit 1-202. The HMD 1-200 may also include a lens 1-218, which can be removably coupled to the display unit 1-202, for example, on a first component including a display screen and a second display component. The lens 1-218 may include a custom prescription lens configured for vision correction. As noted, each part shown in the Figure 1D decomposition view and described above can be removably joined, attached, reattached, and replaced to update the parts or swap out parts for different users. For example, bands such as band 1-216, light seals such as light seal 1-210, lenses such as lens 1-218, and electronic strips such as electronic strips 1-205a-b can be swapped out according to the user, such that these parts are customized to fit and correspond to a single user of the HMD 1-200.

[0086] Figure 1D Any one of the feature portions, components, and / or parts shown (including their arrangements and configurations) may be included, either individually or in any combination, in Figure 1B , Figure 1C and Figures 1E to 1Fin any other examples of the devices, features, components, and parts shown and described herein. Similarly, reference Figure 1B , Figure 1C and Figures 1E to 1F any one of the features, components, and / or parts shown or described (including their arrangements and configurations) may be included individually or in any combination in Figure 1D the examples of the devices, features, components, and parts shown.

[0087] Figure 1E A exploded view showing an example of the display unit 1-306 of the HMD is shown. The display unit 1-306 may include a front display assembly 1-308, a frame / housing assembly 1-350, and a curtain assembly 1-324. The display unit 1-306 may further include a sensor assembly 1-356, a logic board assembly 1-358, and a cooling assembly 1-360 disposed between the frame assembly 1-350 and the front display assembly 1-308. In at least one example, the display unit 1-306 may further include a rear display assembly 1-320, which includes a first rear display screen 1-322a and a second rear display screen 1-322b disposed between the frame 1-350 and the curtain assembly 1-324.

[0088] In at least one example, the display unit 1-306 may further include a motor assembly 1-362, which is configured as an adjustment mechanism for adjusting the position of the display screens 1-322a-b of the display assembly 1-320 relative to the frame 1-350. In at least one example, the display assembly 1-320 is mechanically coupled to the motor assembly 1-362, and each display screen 1-322a-b has at least one motor, such that the motor can translate the display screens 1-322a-b to match the pupil distance of the user's eyes.

[0089] In at least one example, the display unit 1-306 may include a dial or button 1-328, which can be pressed relative to the frame 1-350 and can be accessed by a user outside the frame 1-350. The button 1-328 may be electrically connected to the motor assembly 1-362 via a controller, such that the button 1-328 can be manipulated by the user to cause the motors of the motor assembly 1-362 to adjust the position of the display screens 1-322a-b.

[0090] Figure 1E any one of the features, components, and / or parts shown (including their arrangements and configurations) may be included individually or in any combination in Figures 1B to 1D and Figure 1F any other examples of the devices, features, components, and parts shown and described herein. Similarly, reference Figures 1B to 1D and Figure 1FAny of the features, components, and / or parts shown and described (including their arrangements and configurations) may be included, either individually or in any combination, in Figure 1E the examples of the devices, features, components, and parts shown.

[0091] Figure 1F A exploded view of another example of a display unit 1-406 of an HMD device similar to other HMD devices described herein is shown. The display unit 1-406 may include a front display assembly 1-402, a sensor assembly 1-456, a logic board assembly 1-458, a cooling assembly 1-460, a frame assembly 1-450, a rear display assembly 1-421, and a curtain assembly 1-424. The display unit 1-406 may also include a motor assembly 1-462 for adjusting the positions of a first display sub-assembly 1-420a and a second display sub-assembly 1-420b of the rear display assembly 1-421, including a first corresponding display screen and a second corresponding display screen for inter-pupillary adjustment, as described above.

[0092] Figure 1F The various parts, systems, and components shown in the exploded view are described in more detail herein with reference to Figures 1B to 1E and the subsequent figures referred to in this disclosure. Figure 1F The display unit 1-406 shown may be assembled and integrated with Figures 1B to 1E a fixing mechanism shown, which includes an electronic strip, a belt, and other components including a light seal, a connection assembly, etc.

[0093] Figure 1F Any of the features, components, and / or parts shown and described (including their arrangements and configurations) may be included, either individually or in any combination, in Figures 1B to 1E any other examples of the devices, features, components, and parts shown and described herein. Similarly, with reference to Figures 1B to 1E Any of the features, components, and / or parts shown and described (including their arrangements and configurations) may be included, either individually or in any combination, in Figure 1F the examples of the devices, features, components, and parts shown.

[0094] Figure 1G A perspective exploded view of a front cover assembly 3-100 of an HMD device described herein is shown, such as Figure 1G the front cover assembly 3-100 of the HMD 3-100 shown or any other front cover assembly 3-1 of an HMD device shown and described herein. Figure 1GThe front cover assembly 3-100 shown may include a transparent or translucent cover 3-102, a shield 3-104 (or "awning"), an adhesive layer 3-106, a display assembly 3-108 including a bi-convex lens panel or array 3-110, and a structural trim 3-112. The adhesive layer 3-106 may secure the shield 3-104 and / or the transparent cover 3-102 to the display assembly 3-108 and / or the trim 3-112. The trim 3-112 may secure the various components of the front cover assembly 3-100 to the frame or base of the HMD device.

[0095] In at least one example, as Figure 1G shown, the transparent cover 3-102, the shield 3-104, and the display assembly 3-108 including the bi-convex lens array 3-110 may be curved to conform to the curvature of the user's face. The transparent cover 3-102 and the shield 3-104 may be curved in two or three dimensions, e.g., vertically along the Z-direction inside and outside the Z-X plane, and horizontally along the X-direction inside and outside the Z-X plane. In at least one example, the display assembly 3-108 may include a bi-convex lens array 3-110 and a display panel having pixels configured to project light through the shield 3-104 and the transparent cover 3-102. The display assembly 3-108 may be curved in at least one direction (e.g., the horizontal direction) to conform to the curvature of the user's face from one side (e.g., the left side) to the other side (e.g., the right side) of the face. In at least one example, each layer or component of the display assembly 3-108 (which will be shown and described in more detail in subsequent figures, but which may include the bi-convex lens array 3-110 and the display layer) may be similarly or concentrically curved in the horizontal direction to conform to the curvature of the user's face.

[0096] In at least one example, the shield 3-104 may include a transparent or translucent material through which the display assembly 3-108 projects light. In one example, the shield 3-104 may include one or more opaque portions, such as an opaque ink printed portion or other opaque film portion on the back surface of the shield 3-104. When the HMD device is worn, the back surface may be the surface of the shield 3-104 facing the user's eyes. In at least one example, the opaque portion may be on the front surface of the shield 3-104 opposite the back surface. In at least one example, one or more opaque portions of the shield 3-104 may include a peripheral portion that visually hides any components around the outer perimeter of the display screen of the display assembly 3-108. In this way, the opaque portions of the shield hide any other components of the HMD device that would otherwise be visible through the transparent or translucent cover 3-102 and / or the shield 3-104, including electronic components, structural components, etc.

[0097] In at least one example, the shield 3-104 may define one or more apertured transparent portions 3-120 through which the sensor may send and receive signals. In one example, portion 3-120 is an aperture through which the sensor may extend or through which the sensor may send and receive signals. In one example, portion 3-120 is a transparent portion, or a portion that is more transparent than the surrounding translucent or opaque portion of the shield, through which the sensor may send and receive signals through the shield and through the transparent cover 3-102. In one example, the sensor may include a camera, an IR sensor, a LUX sensor, or any other visual or non-visual environmental sensor of the HMD device.

[0098] Figure 1G Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included, either alone or in any combination, in any other example of the devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included, either alone or in any combination, in Figure 1G the examples of the devices, features, components, and parts shown.

[0099] Figure 1H An exploded view of an example of the HMD device 6-100 is shown. The HMD device 6-100 may include a sensor array or system 6-102 that includes one or more sensors, cameras, projectors, etc. mounted to one or more components of the HMD 6-100. In at least one example, the sensor system 6-102 may include a bracket 1-338 to which one or more sensors of the sensor system 6-102 may be fixed / fastened.

[0100] Figure 1I A portion of the HMD device 6-100 is shown, including the front transparent cover 6-104 and the sensor system 6-102. The sensor system 6-102 may include a plurality of different sensors, transmitters, receivers, including cameras, IR sensors, projectors, etc. The transparent cover 6-104 is shown in front of the sensor system 6-102 to illustrate the relative positions of the various sensors and transmitters and the orientation of each sensor / transmitter of the system 6-102. As used herein, the terms "lateral", "side", "transverse", "horizontal", and other similar terms refer to the orientation or direction indicated by the X-axis as Figure 1J shown. Terms such as "vertical", "upward", "downward", and similar terms refer to the orientation or direction indicated by the Z-axis as Figure 1J shown. Terms such as "forward", "backward", "frontward", "backward", and similar terms refer to the orientation or direction indicated by the Y-axis as Figure 1J shown.

[0101] In at least one example, the transparent cover 6-104 may define the front outer surface of the HMD device 6-100, and a sensor system 6-102 including various sensors and their components may be disposed behind the cover 6-104 in the Y-axis / direction. The cover 6-104 may be transparent or translucent to allow light to pass through the cover 6-104, including both light detected by the sensor system 6-102 and light emitted therefrom.

[0102] As described elsewhere herein, the HMD device 6-100 may include one or more controllers that include a processor for electrically coupling the various sensors and transmitters of the sensor system 6-102 to one or more motherboards, processing units, and other electronic devices such as a display screen. Additionally, as will be shown in more detail with reference to other figures below, the various sensors, transmitters, and other components of the sensor system 6-102 may be coupled to Figure 1I various structural frame members, brackets, etc. of the HMD device 6-100 not shown. For clarity, Figure 1I the components of the sensor system 6-102 are shown unattached and unelectrically coupled to other components.

[0103] In at least one example, the device may include one or more controllers having a processor configured to execute instructions stored on a memory component electrically coupled to the processor. The instructions may include or cause the processor to execute one or more algorithms for self-correcting the angles and positions of the various cameras described herein as the initial position, angle, or orientation of the camera changes over time due to an accidental drop event or other event that causes collision or deformation.

[0104] In at least one example, the sensor system 6-102 may include one or more scene cameras 6-106. The system 6-102 may include two scene cameras 6-102 disposed on either side of the bridge or arch structure of the HMD device 6-100 such that each of the two cameras 6-106 generally corresponds to the position of the user's left and right eyes behind the cover 6-103. In at least one example, the scene cameras 6-106 are generally oriented forward in the Y-direction to capture images in front of the user during use of the HMD 6-100. In at least one example, the scene cameras are color cameras and provide images and content for MR video passthrough to a display screen facing the user's eyes when using the HMD device 6-100. The scene cameras 6-106 may also be used for environmental and object reconstruction.

[0105] In at least one example, the sensor system 6-102 may include a first depth sensor 6-108 that generally points forward in the Y direction. In at least one example, the first depth sensor 6-108 may be used for environment and object reconstruction and for hand and body tracking of the user. In at least one example, the sensor system 6-102 may include a second depth sensor 6-110 centered along the width of the HMD device 6-100 (e.g., along the X-axis). For example, the second depth sensor 6-110 may be disposed above the central bridge of the nose or on an adapter structure above the nose when the user wears the HMD 6-100. In at least one example, the second depth sensor 6-110 may be used for environment and object reconstruction and for hand and body tracking. In at least one example, the second depth sensor may include a LIDAR sensor.

[0106] In at least one example, the sensor system 6-102 may include a depth projector 6-112 that generally faces forward to project electromagnetic waves (e.g., in the form of a pre-determined pattern of light points) into the field of view or within the field of view of the user and / or the scene camera 6-106, or into a field of view that includes and extends beyond the field of view of the user and / or the scene camera 6-106. In at least one example, the depth projector is capable of projecting electromagnetic waves of light in the form of a pattern of light points that are reflected from an object and back into the aforementioned depth sensors, including depth sensors 6-108, 6-110. In at least one example, the depth projector 6-112 may be used for environment and object reconstruction and for hand and body tracking.

[0107] In at least one example, the sensor system 6-102 may include a downward-facing camera 6-114 whose field of view generally points downward relative to the HDM device 6-100 along the Z-axis. In at least one example, the downward camera 6-114 may be disposed on the left and right sides of the HMD device 6-100 as shown and is used for hand and body tracking, head-mounted headset tracking, and facial avatar detection and creation for displaying a user avatar on the forward display screen of the HMD device 6-100 described elsewhere herein. For example, the downward camera 6-114 may be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the cheeks, mouth, and chin.

[0108] In at least one example, the sensor system 6-102 can include a jaw camera 6-116. In at least one example, the jaw camera 6-116 can be disposed on the left and right sides of the HMD device 6-100 as shown and is used for hand and body tracking, headset tracking, and face avatar detection and creation for displaying a user avatar on the forward display screen of the HMD device 6-100 described elsewhere herein. For example, the jaw camera 6-116 can be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the user's jaw, cheeks, mouth, and chin. For hand and body tracking, headset tracking, and face avatar

[0109] In at least one example, the sensor system 6-102 can include a side camera 6-118. The side camera 6-118 can be oriented to capture left and right views in the X-axis or in a direction relative to the HMD device 6-100. In at least one example, the side camera 6-118 can be used for hand and body tracking, headset tracking, and face avatar detection and recreation.

[0110] In at least one example, the sensor system 6-102 can include a plurality of eye tracking and gaze tracking sensors for determining the identity, status, and gaze direction of the user's eyes during and / or before use. In at least one example, the eye / gaze tracking sensors can include a nose-eye camera 6-120 that is disposed on either side of the user's nose and is adjacent to the user's nose when the HMD device 6-100 is worn. The eye / gaze sensors can also include a bottom eye camera 6-122 disposed below the respective user's eye for capturing an image of the eye for face avatar detection and creation, gaze tracking, and iris identification functions.

[0111] In at least one example, the sensor system 6-102 can include an infrared illuminator 6-124 that points outward from the HMD device 6-100 to illuminate the external environment and any objects therein with IR light for IR detection using one or more IR sensors of the sensor system 6-102. In at least one example, the sensor system 6-102 can include a flicker sensor 6-126 and an ambient light sensor 6-128. In at least one example, the flicker sensor 6-126 can detect the top light refresh rate to avoid display flicker. In one example, the infrared illuminator 6-124 can include a light-emitting diode and can be particularly used in low-light environments to illuminate the user's hand and other objects in low light for detection by the infrared sensors of the sensor system 6-102.

[0112] In at least one example, multiple sensors (including scene cameras 6-106, downward cameras 6-114, jaw cameras 6-116, side cameras 6-118, depth projectors 6-112, and depth sensors 6-108, 6-110) can be used in combination with an electrically coupled controller to combine depth data with camera data for hand tracking and for sizing, in order to better perform hand tracking and object recognition and tracking functions of the HMD device 6-100. In at least one example, the downward camera 6-114, jaw camera 6-116, and side camera 6-118 described above and shown in Figure 1I can be wide-angle cameras capable of operating in the visible and infrared spectra. In at least one example, these cameras 6-114, 6-116, 6-118 can operate only in black-and-white light detection to simplify image processing and obtain sensitivity.

[0113] Figure 1I Any of the features, components, and / or parts shown (including their arrangement and configuration) can be included, either individually or in any combination, in Figures 1J to 1L any other example of the devices, features, components, and parts shown and described herein. Similarly, any of the features, components, and / or parts shown and described with reference to Figure 1J to Figure 1L can be included, either individually or in any combination, in Fig. 1I the examples of the devices, features, components, and parts shown.

[0114] Figure 1J A lower perspective view of an example of an HMD 6-200 including a cover or shroud 6-204 fixed to a frame 6-230 is shown. In at least one example, the sensors 6-203 of the sensor system 6-202 can be disposed around the perimeter of the HDM 6-200 such that the sensors 6-203 are disposed outwardly around the perimeter of the display area or region 6-232 so as not to obstruct the viewing of the displayed light. In at least one example, the sensors can be disposed behind the shroud 6-204 and aligned with the transparent portion of the shroud, thereby allowing the sensors and projectors to allow light to pass back and forth through the shroud 6-204. In at least one example, an opaque ink or other opaque material or film / layer can be disposed on the shroud 6-204 around the display area 6-232 to hide the components of the HMD 6-200 outside the display area 6-232 rather than the transparent portion defined by the opaque portion, through which the sensors and projectors transmit and receive light and electromagnetic signals during operation. In at least one example, the shroud 6-204 allows light to pass through from the display (e.g., within the display area 6-232), but does not allow light to pass radially outward from the display area around the perimeter of the display and the shroud 6-204.

[0115] In some examples, the shield 6-204 includes a transparent portion 6-205 and an opaque portion 6-207, as described above and elsewhere herein. In at least one example, the opaque portion 6-207 of the shield 6-204 may define one or more transparent regions 6-209 through which the sensors 6-203 of the sensor system 6-202 may send and receive signals. In the example shown, the sensors 6-203 of the sensor system 6-202 send and receive signals through the shield 6-204, or more specifically through the (or defined) transparent regions 6-209 of the opaque portion 6-207 of the shield 6-204. The sensors may include the same or similar sensors as those shown in the examples of Fig. 1I , such as depth sensors 6-108 and 6-110, depth projectors 6-112, first and second scene cameras 6-106, first and second downward cameras 6-114, first and second side cameras 6-118, and first and second infrared illuminators 6-124. These sensors are also shown in the examples of Figure 1K and Figure 1L . Other sensors, sensor types, sensor quantities, and their relative positions may be included in one or more other examples of the HMD.

[0116] Figure 1J Any of the features, components, and / or parts shown (including their arrangements and configurations) may be included, either alone or in any combination, in the Fig. 1I and Figure 1K to Figure 1L shown and any other examples of the devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts shown or described with reference to Fig. 1I and Figure 1K to Figure 1L may be included, either alone or in any combination, in the examples of the devices, features, components, and parts shown in Figure 1J .

[0117] Figure 1K A front view of a portion of an example of the HMD device 6-300 is shown, including a display 6-334, brackets 6-336, 6-338, and a frame or housing 6-330. Figure 1K The example shown does not include a front cover or shield to show the brackets 6-336, 6-338. For example, Figure 1J the shield 6-204 shown includes an opaque portion 6-207 that will visually cover / block the viewing of anything external (e.g., radially / peripherally external) to the display / display area 6-334, including the sensor 6-303 and the bracket 6-338.

[0118] In at least one example, the various sensors of the sensor system 6-302 are coupled to brackets 6-336, 6-338. In at least one example, the scene cameras 6-306 include tight tolerances on the angles relative to each other. For example, the tolerance on the mounting angle between two scene cameras 6-306 can be 0.5 degrees or less, such as 0.3 degrees or less. To achieve and maintain such tight tolerances, in one example, the scene cameras 6-306 can be mounted to bracket 6-338 rather than the shroud. The bracket can include a cantilever on which the scene cameras 6-306 and other sensors of the sensor system 6-302 can be mounted to maintain position and orientation unchanged in the event of a drop event that causes any deformation of the other brackets 6-226, the housing 6-330, and / or the shroud by the user.

[0119] Figure 1K Any of the features, components, and / or parts shown (including their arrangement and configuration) can be included, either individually or in any combination, in Figures 1I to 1J and Figure 1L any other example of the devices, features, components, and parts shown and described herein. Similarly, reference Figures 1I to 1J and Figure 1L to any of the features, components, and / or parts shown or described (including their arrangement and configuration) can be included, either individually or in any combination, in Figure 1K the examples of the devices, features, components, and parts shown.

[0120] Figure 1L A bottom view of an example of the HMD 6-400 is shown, including the front display / cover assembly 6-404 and the sensor system 6-402. The sensor system 6-402 can be similar to other sensor systems described above and elsewhere herein, including reference Figures 1I to 1K described. In at least one example, the chin camera 6-416 can face downward to capture images of the lower facial features of the user. In one example, the chin camera 6-416 can be directly coupled to the frame or housing 6-430 or one or more internal brackets that are directly coupled to the shown frame or housing 6-430. The frame or housing 6-430 can include one or more holes / openings 6-415 through which the chin camera 6-416 can send and receive signals.

[0121] Figure 1L Any of the features, components, and / or parts shown (including their arrangement and configuration) can be included, either individually or in any combination, in Figures 1I to 1K any other example of the devices, features, components, and parts shown and described herein. Similarly, reference Figures 1I to 1KAny of the features, components, and / or parts shown and described (including their arrangements and configurations) may be included, either individually or in any combination, in Figure 1L the examples of the devices, features, components, and parts shown.

[0122] Figure 1M A rear perspective view of an interpupillary distance (IPD) adjustment system 11.1.1 - 102 is shown, the IPD adjustment system including first and second optical modules 11.1.1 - 104a - b that are slidably engaged / coupled to respective guide rods 11.1.1 - 108a - b and motors 11.1.1 - 110a - b of left and right adjustment subsystems 11.1.1 - 106a - b. The IPD adjustment system 11.1.1 - 102 may be coupled to a bracket 11.1.1 - 112 and includes buttons 11.1.1 - 114 that are in electrical communication with motors 11.1.1 - 110a - b. In at least one example, buttons 11.1.1 - 114 may be in electrical communication with first and second motors 11.1.1 - 110a - b via a processor or other circuit components such that the first and second motors 11.1.1 - 110a - b are activated and cause the first and second optical modules 11.1.1 - 104a - b to change positions relative to each other.

[0123] In at least one example, the first and second optical modules 11.1.1 - 104a - b may include respective display screens that are configured to project light toward a user's eyes when wearing the HMD 11.1.1 - 100. In at least one example, a user may manipulate (e.g., press and / or rotate) buttons 11.1.1 - 114 to activate position adjustment of the optical modules 11.1.1 - 104a - b to match the user's interpupillary distance. The optical modules 11.1.1 - 104a - b may also include one or more cameras or other sensor / sensor systems for imaging and measuring the user's IPD such that the optical modules 11.1.1 - 104a - b may be adjusted to match the IPD.

[0124] In one example, the user can manipulate button 11.1.1-114 to cause an automatic position adjustment of the first and second optical modules 11.1.1-104a-b. In one example, the user can manipulate button 11.1.1-114 to cause a manual adjustment such that the optical modules 11.1.1-104a-b move further or closer (e.g., when the user rotates button 11.1.1-114 in one way or another) until the user visually matches her / his own IPD. In one example, the manual adjustment is communicated electronically via one or more circuits, and the power for moving the optical modules 11.1.1-104a-b via motors 11.1.1-110a-b is provided by a power source. In one example, the adjustment and movement of the optical modules 11.1.1-104a-b via manipulation of button 11.1.1-114 is actuated mechanically via movement of button 11.1.1-114.

[0125] Figure 1M Any of the features, components, and / or parts shown (including their arrangement and configuration) can be included, either alone or in any combination, in any other example of the devices, features, components, and parts shown in any other figures and described herein. Similarly, any of the features, components, and / or parts shown or described with reference to any other figure (including their arrangement and configuration) can be included, either alone or in any combination, in Figure 1M the examples of the devices, features, components, and parts shown.

[0126] Figure 1N A front perspective view of a portion of the HMD 11.1.2-100 is shown, including an external structural frame 11.1.2-102 and an internal or intermediate structural frame 11.1.2-104 that define a first aperture 11.1.2-106a and a second aperture 11.1.2-106b. The apertures 11.1.2-106a-b are shown Figure 1N in dashed lines because the view of the apertures 11.1.2-106a-b may be blocked by one or more other components of the HMD 11.1.2-100 that are coupled to the internal frame 11.1.2-104 and / or the external frame 11.1.2-102, as shown. In at least one example, the HMD 11.1.2-100 can include a first mounting bracket 11.1.2-108 that is coupled to the internal frame 11.1.2-104. In at least one example, the mounting bracket 11.1.2-108 is coupled to the internal frame 11.1.2-104 between the first and second apertures 11.1.2-106a-b.

[0127] The mounting bracket 11.1.2-108 may include an intermediate or central portion 11.1.2-109 coupled to the internal frame 11.1.2-104. In some examples, the intermediate or central portion 11.1.2-109 may not be the geometric middle or center of the bracket 11.1.2-108. Instead, the intermediate / central portion 11.1.2-109 may be disposed between a first cantilevered extension arm and a second cantilevered extension arm that extend away from the intermediate portion 11.1.2-109. In at least one example, the mounting bracket 108 includes a first cantilever 11.1.2-112 and a second cantilever 11.1.2-114 that extend away from the intermediate portion 11.1.2-109 of the mounting bracket 11.1.2-108 coupled to the internal frame 11.1.2-104.

[0128] As Figure 1N shown, the outer frame 11.1.2-102 may define a curved geometry on its lower side to accommodate a user's nose when the user wears the HMD 11.1.2-100. The curved geometry may be referred to as a nose bridge 11.1.2-111 and is shown centered on the lower side of the HMD 11.1.2-100 as illustrated. In at least one example, the mounting bracket 11.1.2-108 may be connected to the internal frame 11.1.2-104 between holes 11.1.2-106a-b such that the cantilevers 11.1.2-112, 11.1.2-114 extend downward and laterally outward away from the intermediate portion 11.1.2-109 to be geometrically complementary to the nose bridge 11.1.2-111 geometry of the outer frame 11.1.2-102. In this way, the mounting bracket 11.1.2-108 is configured to accommodate a user's nose, as described above. The geometry of the nose bridge 11.1.2-111 accommodates the nose because the nose bridge 11.1.2-111 provides a curvature that conforms to the shape of the user's nose, providing a comfortable fit from above, over, and around.

[0129] The first cantilever 11.1.2-112 can extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a first direction, and the second cantilever 11.1.2-114 can extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a second direction opposite to the first direction. The first cantilever 11.1.2-112 and the second cantilever 11.1.2-114 are referred to as "cantilevered" or "cantilever" arms because each arm 11.1.2-112, 11.1.2-114 respectively includes free distal ends 11.1.2-116, 11.1.2-118 that are not attached to the inner frame 11.1.2-102 and the outer frame 11.1.2-104. In this way, the arms 11.1.2-112, 11.1.2-114 overhang from the middle portion 11.1.2-109, which can be connected to the inner frame 11.1.2-104, while the distal ends 11.1.2-102, 11.1.2-104 are not attached.

[0130] In at least one example, the HMD 11.1.2-100 can include one or more components coupled to the mounting bracket 11.1.2-108. In one example, the components include a plurality of sensors 11.1.2-110a-f. Each of the plurality of sensors 11.1.2-110a-f can include various types of sensors, including cameras, IR sensors, etc. In some examples, one or more of the sensors 11.1.2-110a-f can be used for object recognition in three-dimensional space, such that it is important to maintain the precise relative positions of two or more of the plurality of sensors 11.1.2-110a-f. The cantilevered nature of the mounting bracket 11.1.2-108 can protect the sensors 11.1.2-110a-f from damage and position change in the event of an accidental drop by the user. Since the sensors 11.1.2-110a-f are cantilevered on the arms 11.1.2-112, 11.1.2-114 of the mounting bracket 11.1.2-108, the stress and deformation of the inner frame and / or the outer frame 11.1.2-104, 11.1.2-102 are not transmitted to the cantilevers 11.1.2-112, 11.1.2-114, and thus do not affect the relative positions of the sensors 11.1.2-110a-f coupled / installed to the mounting bracket 11.1.2-108.

[0131] Figure 1NAny one of the illustrated features, components, and / or parts (including their arrangements and configurations) may be included, either individually or in any combination, in any other example of the devices, features, components described herein. Similarly, any one of the features, components, and / or parts (including their arrangements and configurations) shown and described herein may be included, either individually or in any combination, in Figure 1N the examples of the devices, features, components, and parts shown.

[0132] Fig.1O An example of an optical module 11.3.2-100 for use in an electronic device (such as an HMD, including the HDM devices described herein) is shown. As shown in one or more other examples described herein, the optical module 11.3.2-100 may be one of two optical modules within an HMD, where each optical module is aligned to project light towards the user's eyes. In this manner, a first optical module may project light towards the user's first eye via a display screen, and a second optical module of the same device may project light towards the user's second eye via another display screen.

[0133] In at least one example, the optical module 11.3.2-100 may include an optical frame or housing 11.3.2-102, which may also be referred to as a barrel or an optical module barrel. The optical module 11.3.2-100 may also include a display 11.3.2-104 coupled to the housing 11.3.2-102, the display including one or more display screens. The display 11.3.2-104 may be coupled to the housing 11.3.2-102 such that the display 11.3.2-104 is configured to project light towards the user's eyes when wearing the HMD to which the display module 11.3.2-100 belongs during use. In at least one example, the housing 11.3.2-102 may surround the display 11.3.2-104 and provide connection features for coupling other components of the optical module described herein.

[0134] In one example, the optical module 11.3.2-100 may include one or more cameras 11.3.2-106 coupled to a housing 11.3.2-102. The cameras 11.3.2-106 may be positioned relative to a display 11.3.2-104 and the housing 11.3.2-102 such that the cameras 11.3.2-106 are configured to capture one or more images of a user's eyes during use. In at least one example, the optical module 11.3.2-100 may further include a light strip 11.3.2-108 surrounding the display 11.3.2-104. In one example, the light strip 11.3.2-108 is disposed between the display 11.3.2-104 and the cameras 11.3.2-106. The light strip 11.3.2-108 may include a plurality of lights 11.3.2-110. The plurality of lights may include one or more light emitting diodes (LEDs) or other lights configured to project light towards a user's eyes when wearing the HMD. Each light 11.3.2-110 in the light strip 11.3.2-108 may be spaced apart around the light strip 11.3.2-108 and thus may be spaced apart evenly or unevenly around the display 11.3.2-104 at various positions on the light strip 11.3.2-108 and around the display 11.3.2-104.

[0135] In at least one example, the housing 11.3.2-102 defines a viewing opening 11.3.2-101 through which a user may view the display 11.3.2-104 when wearing the HMD device. In at least one example, the LEDs are configured and arranged to emit light through the viewing opening 11.3.2-101 onto a user's eyes. In one example, the cameras 11.3.2-106 are configured to capture one or more images of a user's eyes through the viewing opening 11.3.2-101.

[0136] As described above, Fig.1O Each of the components and features of the illustrated optical module 11.3.2-100 may be replicated in another (e.g., second) optical module provided with the HMD to interact with the user's other eye (e.g., project light and capture images).

[0137] Fig.1O Any one of the illustrated features, components, and / or parts (including their arrangement and configuration) may be included, either alone or in any combination, in Figure 1P any other example of the devices, features, components, and parts shown or otherwise described herein. Similarly, reference Figure 1P to or any one of the features, components, and / or parts shown or otherwise described herein (including their arrangement and configuration) may be included, either alone or in any combination, in Fig.1OIn the examples of the devices, features, components, and parts shown.

[0138] Figure 1P A cross-sectional view showing an example of an optical module 11.3.2-200 is presented, including a housing 11.3.2-202, a display assembly 11.3.2-204 coupled to the housing 11.3.2-202, and a lens 11.3.2-216 coupled to the housing 11.3.2-202. In at least one example, the housing 11.3.2-202 defines a first hole or passage 11.3.2-212 and a second hole or passage 11.3.2-214. The passages 11.3.2-212, 11.3.2-214 can be configured to slidably engage corresponding tracks or guide rods of an HMD device to allow the optical module 11.3.2-200 to adjust its position relative to the user's eyes to match the user's interpupillary distance (IPD). The housing 11.3.2-202 is capable of slidably engaging the guide rods to fix the optical module 11.3.2-200 in place within the HMD.

[0139] In at least one example, the optical module 11.3.2-200 may further include a lens 11.3.2-216 coupled to the housing 11.3.2-202 and disposed between the display assembly 11.3.2-204 and the user's eyes when the HMD is worn. The lens 11.3.2-216 can be configured to direct light from the display assembly 11.3.2-204 to the user's eyes. In at least one example, the lens 11.3.2-216 can be part of a lens assembly that includes a corrective lens removably attached to the optical module 11.3.2-200. In at least one example, the lens 11.3.2-216 is disposed above a light strip 11.3.2-208 and one or more eye tracking cameras 11.3.2-206 such that the cameras 11.3.2-206 are configured to capture images of the user's eyes through the lens 11.3.2-216, and the light strip 11.3.2-208 includes lights configured to project light through the lens 11.3.2-216 onto the user's eyes during use.

[0140] Figure 1P Any of the features, components, and / or parts shown (including their arrangement and configuration) can be included, either individually or in any combination, in any other examples of the devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) can be included, either individually or in any combination, in Figure 1P the examples of the devices, features, components, and parts shown.

[0141] Figure 2FIG. 0 is a block diagram of an example of controller 110 in accordance with some embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that, for the sake of brevity and to not obscure more relevant aspects of the embodiments disclosed herein, various other features are not shown. To that end, by way of non-limiting example, in some embodiments, controller 110 includes one or more processing units 202 (e.g., microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., universal serial bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, global system for mobile communications (GSM), code division multiple access (CDMA), time division multiple access (TDMA), global positioning system (GPS), infrared (IR), Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these components and various other components.

[0142] In some embodiments, one or more communication buses 204 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0143] Memory 220 includes high-speed random access memory such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes non-transitory computer-readable storage medium. In some embodiments, memory 220 or the non-transitory computer-readable storage medium of memory 220 stores the following programs, modules, and data structures, or subsets thereof, including optionally operating system 230 and XR experience module 240.

[0144] The operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate single or multiple XR experiences of one or more users (e.g., a single XR experience of one or more users, or multiple XR experiences of corresponding groups of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.

[0145] In some embodiments, the data acquisition unit 241 is configured to obtain data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 of Figure 1A and optionally from one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for instructions and heuristics and metadata for heuristics.

[0146] In some embodiments, the tracking unit 242 is configured to map the scene 105 and track the positioning / position of at least the display generation component 120 relative to Figure 1A the scene 105, and optionally track the position of one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for instructions and heuristics and metadata for heuristics. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the positioning / position of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to Figure 1A the scene 105, relative to the display generation component 120, and / or relative to a coordinate system (which is defined relative to the user's hand). The hand tracking unit 244 is described in more detail below with respect to Figure 4 . In some embodiments, the eye tracking unit 243 is configured to track the positioning or movement of the user's gaze (or more generally, the user's eyes, face, or head) relative to the scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hand)) or relative to the XR content displayed via the display generation component 120. The eye tracking unit 243 is described in more detail below with respect to Figure 5 .

[0147] In some embodiments, the coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by the display generation component 120, and optionally by one or more of the output device 155 and / or the peripheral device 195. For this purpose, in various embodiments, the coordination unit 246 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.

[0148] In some embodiments, the data sending unit 248 is configured to send data (e.g., presentation data, location data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For this purpose, in various embodiments, the data sending unit 248 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.

[0149] Although the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 are shown as residing on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 may be located in separate computing devices.

[0150] In addition, Figure 2 Rather, it is more of a functional description of the various features that may be present in a particular implementation, as opposed to the structural schematic of the embodiments described herein. As will be recognized by those of ordinary skill in the art, items shown separately may be combined, and some items may be separated. For example, Figure 2 Some of the functional modules shown separately in may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the specific division of functions, as well as how the features are allocated therein, will vary depending on the implementation, and in some embodiments, will depend in part on the specific combination of hardware, software, and / or firmware selected for the particular implementation.

[0151] Figure 3FIG. 0 is a block diagram of an example of a display generation component 120 according to some embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that, for the sake of brevity and to not obscure more relevant aspects of the embodiments disclosed herein, various other features are not shown. For that purpose, by way of non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., a microprocessor, ASIC, FPGA, GPU, CPU, processing core, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional internal and / or external image sensors 314, a memory 320, and one or more communication buses 304 for interconnecting these components and various other components.

[0152] In some embodiments, one or more communication buses 304 include circuitry for interconnecting and controlling communication between the various system components. In some embodiments, one or more I / O devices and sensors 306 include an inertial measurement unit (IMU), accelerometers, gyroscopes, thermometers, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, and / or one or more depth sensors (e.g., structured light, time-of-flight, etc.).

[0153] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more XR displays 312 correspond to diffractive, reflective, polarization, holographic, and other waveguide displays. For example, the display generation component 120 (e.g., an HMD) includes a single XR display. In another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, one or more XR displays 312 are capable of presenting MR and VR content. In some embodiments, one or more XR displays 312 are capable of presenting MR or VR content.

[0154] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand and optionally the user's arm (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to the scene that the user would see in the absence of the display generation component 120 (e.g., an HMD) (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.

[0155] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes non-transitory computer-readable storage medium. In some embodiments, memory 320 or the non-transitory computer-readable storage medium of memory 320 stores the following programs, modules, and data structures, or subsets thereof, including optionally operating system 330 and XR rendering module 340.

[0156] Operating system 330 includes procedures for handling various basic system services and for performing hardware-related tasks. In some embodiments, XR rendering module 340 is configured to present XR content to a user via one or more XR displays 312. To this end, in various embodiments, XR rendering module 340 includes data acquisition unit 342, XR rendering unit 344, XR mapping generation unit 346, and data transmission unit 348.

[0157] In some embodiments, data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) at least from Figure 1A controller 110 of. For this purpose, in various embodiments, data acquisition unit 342 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0158] In some embodiments, XR rendering unit 344 is configured to present XR content via one or more XR displays 312. For this purpose, in various embodiments, XR rendering unit 344 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0159] In some embodiments, XR mapping generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate extended reality) based on media content data. For this purpose, in various embodiments, XR mapping generation unit 346 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0160] In some embodiments, the data sending unit 348 is configured to send data (e.g., presentation data, location data, etc.) to at least the controller 110, and optionally to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For this purpose, in various embodiments, the data sending unit 348 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.

[0161] Although the data acquisition unit 342, the XR presentation unit 344, the XR mapping generation unit 346, and the data transmission unit 348 are shown as residing on a single device (e.g., Figure 1A the display generation component 120 of), it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR presentation unit 344, the XR mapping generation unit 346, and the data transmission unit 348 may be located in separate computing devices.

[0162] In addition, Figure 3 Rather more serves as a functional description of the various features that may be present in a particular embodiment, as opposed to a schematic diagram of the structure of the embodiments described herein. As will be recognized by those of ordinary skill in the art, the items shown separately may be combined, and some items may be separated. For example, Figure 3 some of the functional modules shown separately in may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the specific division of functions and how the features are allocated therein will vary according to the specific implementation, and in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for the particular implementation.

[0163] Figure 4 is a schematic illustration of an example embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 ( Figure 1A ) is controlled by the hand tracking unit 244 ( Figure 2 ) to track the positioning / position of one or more parts of the user's hand, and / or the movement of one or more parts of the user's hand relative to the scene 105 of FIG. 1 (e.g., relative to a part of the physical environment around the user, relative to the display generation component 120, or relative to a part of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system that is defined relative to the user's hand). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0164] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures hand images at a sufficient resolution to enable the fingers and their corresponding positions to be distinguished. The image sensor 404 typically captures images of other parts of the user's body and may also or possibly capture images of all parts of the body, and may have a zoom capability or a dedicated sensor with increased magnification to capture an image of the hand at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in combination with other image sensors to capture the physical environment of the scene 105, or serves as the image sensor for capturing the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment in such a way that the field of view of the image sensor 404 or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are considered inputs to the controller 110.

[0165] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and in addition, possibly color image data) to the controller 110, which extracts high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API), and the application drives the display generation component 120 accordingly. For example, a user can interact with software running on the controller 110 by moving his hand 406 and changing his hand pose.

[0166] In some embodiments, the image sensor 404 projects a speckle pattern onto the scene containing the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral offset of the speckles in the pattern. This method is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. This method gives the depth coordinates of points in the scene relative to a pre-determined reference plane at a specific distance from the image sensor 404. In the present disclosure, it is assumed that the image sensor 404 defines an orthogonal set of x-axis, y-axis, and z-axis, such that the depth coordinates of points in the scene correspond to the z-component measured by the image sensor. Alternatively, the image sensor 404 (e.g., the hand tracking device) may use other 3D mapping methods, such as stereoscopy or time-of-flight measurement, based on a single or multiple cameras or other types of sensors.

[0167] In some embodiments, the hand tracking device 140 captures and processes a time series of depth maps that include the user's hand as the user moves their hand (e.g., the entire hand or one or more fingers). Software running on a processor in the image sensor 404 and / or the controller 110 processes the 3D map data to extract patch descriptors of the hand in these depth maps. The software can match these descriptors to patch descriptors stored in the database 408 based on a previous learning process to estimate the pose of the hand in each frame. The pose generally includes the 3D positions of the user's hand joints and finger tips.

[0168] The software can also analyze the trajectories of the hand and / or fingers over multiple frames in the sequence to identify gestures. The pose estimation functionality described herein can alternate with the motion tracking functionality such that the patch-based pose estimation is only performed every two (or more) frames, and tracking is used to find changes in the pose that occur on the remaining frames. Pose, motion, and gesture information are provided to applications running on the controller 110 via the aforementioned API. The program can, for example, move and modify the images presented on the display generation component 120 in response to the pose and / or gesture information, or perform other functions.

[0169] In some embodiments, gestures include air gestures. An air gesture is detected when the user does not touch an input element (or independent of an input element that is part of a device (e.g., the computer system 101, one or more input devices 125, and / or the hand tracking device 140)) and is based on the detected movement of a part of the user's body (e.g., the head, one or more arms, one or more hands, one or more fingers, and / or one or more legs) through the air (including the movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), the movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one of the user's hands relative to the other user's hand, and / or the movement of the user's finger relative to another finger or part of the hand of the user), and / or the absolute movement of a part of the user's body (e.g., a tap gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shake gesture that includes a predetermined speed or amount of rotation of a part of the user's body)).

[0170] In some embodiments, according to some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures performed by the movement of a user's finger relative to other fingers or parts of the user's hand for interacting with an XR environment (e.g., a virtual or mixed reality environment). In some embodiments, the air gesture is detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and is based on the detected movement of a part of the user's body through the air, including the movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), the movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one hand of the user relative to the other hand of the user, and / or the movement of the user's finger relative to another finger or part of the user's hand), and / or the absolute movement of a part of the user's body (e.g., a tap gesture including the hand moving a predetermined amount and / or speed in a predetermined pose, or a shake gesture including a predetermined speed or rotation amount of a part of the user's body).

[0171] In some embodiments where the input gesture is an air gesture (e.g., in the absence of physical contact with an input device that provides information to the computer system about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touch screen or contact with a mouse or touchpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Thus, in embodiments involving air gestures, for example, the input gesture is detected in combination with (e.g., simultaneously) the movement of the user's finger and / or hand towards the user interface element, along with the attention (e.g., gaze) towards the user interface element to perform a pinch and / or tap input, as described below.

[0172] In some embodiments, an input gesture directed to a user interface object is performed with direct or indirect reference to the user interface object. For example, user input is performed directly on the user interface object by performing the input at a location corresponding to the positioning of the user's hand relative to the positioning of the user interface object in a three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, when attention (e.g., a gaze) of the user to the user interface object is detected, the input gesture is performed indirectly on the user interface object based on the positioning of the user's hand not being at the location corresponding to the positioning of the user interface object in the three-dimensional environment while the user performs the input gesture. For example, for a direct input gesture, the user can direct the user's input to the user interface object by initiating a gesture at or near a location corresponding to the display positioning of the user interface object (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or between 0 and 5 cm measured from the outer edge of the option or the central portion of the option). For an indirect input gesture, the user can direct the user's input to the user interface object by focusing on the user interface object (e.g., by gazing at the user interface object), and while focusing on the option, the user initiates an input gesture (e.g., at any location detectable by the computer system) (e.g., at a location not corresponding to the display positioning of the user interface object).

[0173] In some embodiments, according to some embodiments, input gestures (e.g., air gestures) used in the various examples and embodiments described herein include pinch inputs and tap inputs for interacting with a virtual or mixed reality environment. For example, the pinch inputs and tap inputs described below are performed as air gestures.

[0174] In some embodiments, the pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch and drag gesture, or a double pinch gesture. For example, a pinch gesture as an air gesture includes the movement of two or more fingers of a hand to contact each other, i.e., optionally followed by an immediate (e.g., within 0 to 1 second) interruption of contact with each other. A long pinch gesture as an air gesture includes the movement of two or more fingers of a hand contacting each other for at least a threshold amount of time (e.g., at least 1 second) before detecting an interruption of contact with each other. For example, a long pinch gesture includes a user maintaining a pinch gesture (e.g., where two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between the two or more fingers is detected. In some embodiments, a double pinch gesture as an air gesture includes two (e.g., or more) pinch inputs (e.g., performed by the same hand) detected consecutively and immediately (e.g., within a predefined time period) with each other. For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., interrupts contact between two or more fingers), and performs a second pinch input within a predefined time period (e.g., within 1 second or within 2 seconds) after releasing the first pinch input.

[0175] In some embodiments, a pinch-and-drag gesture, as an air gesture (e.g., an air drag gesture or an air swipe gesture), includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in combination with (e.g., following) a drag input that changes the position of the user's hand from a first position (e.g., the start position of the drag) to a second position (e.g., the end position of the drag). In some embodiments, the user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., opens their two or more fingers) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., the user pinches two or more fingers together and moves the same hand into a second position in the air using a drag gesture). In some embodiments, the pinch input is performed by the user's first hand, and the drag input is performed by the user's second hand (e.g., while the user continues the pinch input with the user's first hand, the user's second hand moves in the air from a first position to a second position). In some embodiments, an input gesture as an air gesture includes an input performed using both of the user's hands (e.g., a pinch and / or tap input). For example, the input gesture includes two (e.g., or more) pinch inputs performed in combination with each other (e.g., concurrently or within a predefined time period). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch-and-drag input) is performed using the user's first hand, and in combination with performing the pinch input using the first hand, a second pinch input is performed using the other hand (e.g., the second hand of the user's two hands).

[0176] In some embodiments, a tap input performed as an air gesture (e.g., pointing to a user interface element) includes the movement of the user's finger towards the user interface element, the movement of the user's hand towards the user interface element (optionally, the user's finger extends towards the user interface element), the downward movement of the user's finger (e.g., mimicking a mouse click movement or a tap on a touch screen), or other predefined movement of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on the movement characteristics of the finger or hand performing the tap gesture movement, which is the movement of the finger or hand away from the user's viewpoint and / or towards an object that is the target of the tap input, followed by the end of the movement. In some embodiments, the end of the movement is detected based on a change in the movement characteristics of the finger or hand performing the tap gesture (e.g., the end of the movement away from the user's viewpoint and / or towards an object that is the target of the tap input, the reversal of the movement direction of the finger or hand, and / or the reversal of the acceleration direction of the movement of the finger or hand).

[0177] In some embodiments, the user's attention is determined to be directed to a portion of the three-dimensional environment based on the detection of a gaze directed to the portion of the three-dimensional environment (optionally, without requiring other conditions). In some embodiments, the user's attention is determined to be directed to the portion of the three-dimensional environment based on the detection of a gaze directed to the portion of the three-dimensional environment using one or more additional conditions, such as requiring the gaze to be directed to the portion of the three-dimensional environment for at least a threshold duration (e.g., dwell duration) and / or requiring the gaze to be directed to the portion of the three-dimensional environment when the user's viewpoint is within a distance threshold from the portion of the three-dimensional environment, so that the device determines that the user's attention is directed to the portion of the three-dimensional environment, wherein if one of these additional conditions is not met, the device determines that the attention is not directed to the portion of the three-dimensional environment to which the gaze is directed (e.g., until the one or more additional conditions are met).

[0178] In some embodiments, the detection of the ready state configuration of the user or a part of the user is detected by the computer system. The detection of the ready state configuration of the hand is used by the computer system as an indication that the user may be about to use one or more air gesture inputs (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein) performed by the hand to interact with the computer system. For example, based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape where the thumb and one or more fingers are extended and spaced apart to prepare for a pinch or grab gesture, or a pre-tap where one or more fingers are extended and the palm is facing away from the user), based on whether the hand is in a predetermined orientation relative to the user's viewpoint (e.g., below the user's head and above the user's waist and extending at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or based on whether the hand has moved in a particular manner (e.g., moving towards an area in front of the user that is above the user's waist and below the user's head or moving away from the user's body or legs). In some embodiments, the ready state is used to determine whether interactive elements of the user interface respond to attention (e.g., gaze) inputs.

[0179] In scenarios where input is described with reference to an air gesture, it should be understood that a hardware input device attached to one or more of the user's hands or held by one or more of the user's hands can be used to detect a similar gesture, where optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units can be used to track the positioning of the hardware input device in space, and the positioning and / or movement of the hardware input device is used in place of the positioning and / or movement of one or more hands in the corresponding air gesture. In scenarios where input is described with reference to an air pose, it should be understood that a hardware input device attached to one or more of the user's hands or held by one or more of the user's hands can be used to detect a similar pose. User input can be detected using controls contained in the hardware input device, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, one or more hand or finger overlays that can detect the positioning or change in positioning of parts of the hand and / or fingers relative to each other, relative to the user's body, and / or relative to the user's physical environment, and / or other hardware input device controls, where user input using the controls contained in the hardware input device is used in place of hand and / or finger gestures such as an air tap or an air pinch in the corresponding air gesture. For example, a selection input described as being performed using an air tap or an air pinch input can alternatively be detected using a button press, a tap on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input. As another example, a movement input described as being performed using an air pinch and drag (e.g., an air drag gesture or an air swipe gesture) can alternatively be detected based on an interaction with a hardware input control (such as a button press and hold, a touch on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input after the movement of the hardware input device (e.g., along with the hand associated with the hardware input device) through space). Similarly, a two-handed input that includes movement of the hands relative to each other can be performed using an air gesture and a hardware input device in a hand that is not performing an air gesture, two hardware input devices held in different hands, or two air gestures performed using various combinations of an air gesture and / or input detected by one or more of the aforementioned hardware input devices.

[0180] In some embodiments, the software can be downloaded electronically, for example, over a network, to the controller 110, or alternatively can be provided on a tangible non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in a memory associated with the controller 110. Alternatively or in addition, some or all of the described functions of the computer can be implemented in dedicated hardware (such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP)). Although in Figure 4The controller 110 is shown, but by way of example, some or all of the processing functions of the controller, as a unit separate from the image sensor 404, may be performed by a suitable microprocessor and software or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand tracking device) or by other devices associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television receiver, a handheld device, or a head-mounted device) or integrated with any other suitable computerized device (such as a game console or a media player). The sensing function of the image sensor 404 may likewise be integrated into a computer or other computerized device to be controlled by the sensor output.

[0181] Figure 4 Also shown is a schematic illustration of a depth map 410 captured by the image sensor 404 according to some embodiments. As described above, the depth map includes a matrix of pixels having corresponding depth values. The pixels 412 corresponding to the hand 406 have been segmented from the background and the wrist in the figure. The brightness of each pixel within the depth map 410 is inversely proportional to its depth value (i.e., the measured z-distance from the image sensor 404), where the gray shading becomes darker as the depth increases. The controller 110 processes these depth values to identify and segment the components of the image having human hand characteristics (i.e., a group of adjacent pixels). These characteristics may include, for example, the overall size, shape, and movement from frame to frame in a sequence of depth maps.

[0182] Figure 4 Also schematically illustrated is a hand skeleton 414 ultimately extracted by the controller 110 from the depth map 410 of the hand 406. In Figure 4 this figure, the hand skeleton 414 is superimposed on the hand background 416 that has been segmented from the original depth map. In some embodiments, key feature points on the hand and optionally on the wrist or arm connected to the hand (e.g., points corresponding to knuckles, finger tips, palm center, the end of the hand connected to the wrist, etc.) are identified and located on the hand skeleton 414. In some embodiments, the controller 110 uses the positions and movements of these key feature points across multiple image frames to determine, according to some embodiments, the gesture being performed by the hand or the current state of the hand.

[0183] Figure 5 An example embodiment of an eye tracking device 130 ( Figure 1A ) is shown. In some embodiments, the eye tracking device 130 is composed of an eye tracking unit 243 ( Figure 2)Control is used to track the positioning and movement of the user's gaze relative to the scene 105 or relative to the XR content displayed via the display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as, a head-mounted headset, helmet, goggles, or glasses) or a handheld device placed in a wearable frame, the head-mounted device includes both components for generating XR content for the user to view and components for tracking the user's gaze relative to the XR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a handheld device or an XR room, the eye tracking device 130 is optionally a device separate from the handheld device or the XR room. In some embodiments, the eye tracking device 130 is a head-mounted device or a part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used in combination with a display generation component that is also head-mounted or a display generation component that is not head-mounted. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally a part of a non-head-mounted display generation component.

[0184] In some embodiments, the display generation component 120 uses a display mechanism (e.g., a left near-eye display panel and a right near-eye display panel) to display a frame including a left image and a right image in front of the user's eyes, thereby providing the user with a 3D virtual view. For example, a head-mounted display generation component may include a left optical lens and a right optical lens (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display, and virtual objects are displayed on the transparent or translucent display, through which the user can directly view the physical environment. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects may be projected, for example, onto a physical surface or projected as a hologram such that an individual uses the system to observe the virtual objects superimposed over the physical environment. In this case, separate display panels and image frames for the left and right eyes may not be required.

[0185] As Figure 5As shown, in some embodiments, the eye tracking device 130 (e.g., a gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) or near-infrared (NIR) camera), and an illumination source (e.g., an IR or NIR light source, such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera can be directed at the user's eyes to receive the IR or NIR light reflected directly from the eyes by the light source, or alternatively can be directed at a "hot" mirror located between the user's eyes and the display panel, which reflects the IR or NIR light from the eyes to the eye tracking camera while allowing visible light to pass through. The eye tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60 frames - 120 frames per second (fps)), analyzes the images to generate gaze tracking information, and transmits the gaze tracking information to the controller 110. In some embodiments, the two eyes of the user are tracked separately by corresponding eye tracking cameras and illumination sources. In some embodiments, only one eye of the user is tracked by corresponding eye tracking cameras and illumination sources.

[0186] In some embodiments, a device-specific calibration process is used to calibrate the eye tracking device 130 to determine the parameters of the eye tracking device for a particular operating environment 100, such as the 3D geometric relationships and parameters of the LEDs, cameras, hot mirrors (if any), eye lenses, and display screens. The device-specific calibration process can be performed at the factory or another facility before the AR / VR equipment is delivered to the end user. The device-specific calibration process can be an automatic calibration process or a manual calibration process. According to some embodiments, the user-specific calibration process can include an estimation of the eye parameters of a particular user, such as pupil position, fovea position, optical axis, visual axis, interpupillary distance, etc. According to some embodiments, once the device-specific parameters and user-specific parameters are determined for the eye tracking device 130, a flash-assisted method can be used to process the images captured by the eye tracking camera to determine the current visual axis and the user's fixation point relative to the display.

[0187] As Figure 5As shown, the eye tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system that includes at least one eye tracking camera 540 (e.g., an infrared (IR) or near-infrared (NIR) camera) positioned on the side of the user's face where eye tracking is to be performed, and an illumination source 530 (e.g., an IR or NIR light source, such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye 592. The eye tracking camera 540 can be pointed at a mirror 550 (which reflects IR or NIR light from the eye 592 while allowing visible light to pass through) (e.g., as shown in the top portion of Figure 5 ), or alternatively can be pointed at the user's eye 592 to receive reflected IR or NIR light from the eye 592 (e.g., as shown in the bottom portion of Figure 5 ).

[0188] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames for a left display panel and a right display panel) and provides the frames 562 to the display 510. The controller 110 uses the gaze tracking input 542 from the eye tracking camera 540 for various purposes, such as for processing the frames 562 for display. The controller 110 optionally estimates the user's gaze point on the display 510 based on the gaze tracking input 542 obtained from the eye tracking camera 540 using a flash assist method or other suitable method. The gaze point estimated based on the gaze tracking input 542 is optionally used to determine the direction in which the user is currently looking.

[0189] The following describes several possible use cases of the current gaze direction of a user and is not intended to be limiting. As an example use case, the controller 110 may render virtual content differently based on the determined direction of the user's gaze. For example, the controller 110 may generate virtual content at a higher resolution in the foveal region determined according to the current gaze direction of the user than in the peripheral region. As another example, the controller may position or move virtual content in the view at least in part based on the current gaze direction of the user. As another example, the controller may display specific virtual content in the view at least in part based on the current gaze direction of the user. As another example use case in an AR application, the controller 110 may direct an external camera for capturing the physical environment of the XR experience to focus in the determined direction. Then, the autofocus mechanism of the external camera may focus on an object or surface in the environment that the user is currently looking at on the display 510. As another example use case, the eye lens 520 may be a focusable lens, and the controller uses the gaze tracking information to adjust the focus of the eye lens 520 such that the virtual object that the user is currently looking at has an appropriate vergence to match the convergence of the user's eyes 592. The controller 110 may utilize the gaze tracking information to direct the eye lens 520 to adjust the focus such that a nearby object that the user is looking at appears at the correct distance.

[0190] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510) mounted in a wearable housing, two eye lenses (e.g., eye lens 520), an eye tracking camera (e.g., eye tracking camera 540), and a light source (e.g., illumination source 530 (e.g., IR or NIR LED)). The light source emits light (e.g., IR or NIR light) towards the user's eyes 592. In some embodiments, the light source may be arranged in a ring or circle around each of the lenses, as Figure 5 shown. In some embodiments, for example, eight illumination sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer illumination sources 530 may be used, and other arrangements and positions of the illumination sources 530 may be used.

[0191] In some embodiments, the display 510 emits light in the visible light range and does not emit light in the IR or NIR range, and thus does not introduce noise in the gaze tracking system. Note that the position and angle of the eye tracking camera 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, cameras 540 with a wider field of view (FOV) and cameras 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, cameras 540 operating at one wavelength (e.g., 850 nm) and cameras 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.

[0192] As Figure 5 The illustrated embodiments of the gaze tracking system may be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide the user with a computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experience.

[0193] Figure 6 Illustrates a flash-assisted gaze tracking pipeline according to some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a flash-assisted gaze tracking system (e.g., the eye tracking device 130 as shown in Figure 1A and Figure 5 ). The flash-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in the tracking state, when analyzing the current frame to track the pupil contour and flash in the current frame, the flash-assisted gaze tracking system uses the previous information from the previous frame. When not in the tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and flash in the current frame, and if successful, initializes the tracking state to "yes" and continues with the next frame in the tracking state.

[0194] As Figure 6 shown, the gaze tracking camera can capture left and right images of the user's left and right eyes. Then the captured images are input into the gaze tracking pipeline for processing starting at 610. As indicated by the arrow returning to element 600, the gaze tracking system can continue to capture images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images can be input into the pipeline for processing. However, in some embodiments or under some conditions, not all of the captured frames are processed by the pipeline.

[0195] At 610, for the currently captured image, if the tracking state is yes, the method proceeds to element 640. At 610, if the tracking state is no, then as indicated at 620, the image is analyzed to detect the user's pupil and flash in the image. At 630, if the pupil and flash are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eye.

[0196] At 640, if proceeding from element 610, the current frame is analyzed to track the pupil and flash based in part on previous information from a previous frame. At 640, if proceeding from element 630, the tracking state is initialized based on the pupil and flash detected in the current frame. The processing result at element 640 is checked to verify that the result of the tracking or detection can be trusted. For example, the result can be checked to determine if the pupil and a sufficient number of flashes for performing gaze estimation are successfully tracked or detected in the current frame. At 650, if the result cannot be trusted, then at element 660, the tracking state is set to no, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is trusted, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and flash information is passed to element 680 to estimate the user's point of gaze.

[0197] Figure 6 It is intended to be used as an example of an eye-tracking technique that can be used for a particular specific implementation. As will be recognized by those of ordinary skill in the art, according to various embodiments, in the computer system 101 for providing an XR experience to a user, other eye-tracking techniques that currently exist or are developed in the future can be used to replace the flash-assisted eye-tracking technique described herein or used in combination with the flash-assisted eye-tracking technique.

[0198] In some embodiments, a captured portion of the real-world environment 602 is used to provide an XR experience to the user, such as a mixed reality environment in which one or more virtual objects are superimposed over a representation of the real-world environment 602.

[0199] Accordingly, the description herein describes some implementations of a three-dimensional environment (e.g., an XR environment) that includes a representation of real-world objects and a representation of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table that exists in a physical environment, which is captured and displayed in the three-dimensional environment (e.g., actively displayed via a camera and a display of a computer system or passively displayed via a transparent or semi-transparent display of the computer system). As previously described, the three-dimensional environment is optionally a mixed reality system, where the three-dimensional environment is based on the physical environment captured by one or more sensors of the computer system and displayed via a display generation component. As a mixed reality system, the computer system is optionally capable of selectively displaying portions and / or objects of the physical environment such that the corresponding portions and / or objects of the physical environment appear as if they exist in the three-dimensional environment displayed by the computer system. Similarly, the computer system is optionally capable of displaying virtual objects in the three-dimensional environment at corresponding locations that have corresponding positions in the real world such that the virtual objects appear as if they exist in the real world (e.g., the physical environment). For example, the computer system optionally displays a vase such that the vase appears as if a real vase is placed on top of a table in the physical environment. In some implementations, the corresponding locations in the three-dimensional environment have corresponding positions in the physical environment. Thus, when the computer system is described as displaying a virtual object at a corresponding location relative to a physical object (e.g., a location at or near the user's hand or a location at or near a physical table), the computer system displays the virtual object at a specific location in the three-dimensional environment such that it appears as if the virtual object is at or near the physical object in the physical environment (e.g., the virtual object is displayed at a location in the three-dimensional environment that corresponds to the location in the physical environment where the virtual object would be displayed if the virtual object were a real object at that specific location).

[0200] In some implementations, real-world objects that exist in the physical environment and are displayed in the three-dimensional environment (e.g., and / or visible via a display generation component) can interact with virtual objects that only exist in the three-dimensional environment. For example, the three-dimensional environment can include a table and a vase placed on top of the table, where the table is a view (or representation) of a physical table in the physical environment and the vase is a virtual object.

[0201] In a three-dimensional environment (e.g., a real environment, a virtual environment, or an environment that includes a mixture of real and virtual objects), an object is sometimes referred to as having depth or simulated depth, or an object is referred to as being visible, displayed, or placed at different depths. In this context, depth refers to a dimension that is different from height or width. In some embodiments, depth is defined relative to a fixed set of coordinates (e.g., where a room or an object has a height, depth, and width defined relative to a fixed set of coordinates). In some embodiments, depth is defined relative to the position or viewpoint of a user, in which case the depth dimension varies based on the position of the user and / or the position and angle of the user's viewpoint. In some embodiments where depth is defined relative to the position of the user relative to a surface of the environment (e.g., the surface of the floor or ground of the environment), an object that is further away from the user along a line extending parallel to the surface is considered to have a greater depth in the environment, and / or the depth of the object is measured along an axis that extends outward from the position of the user and is parallel to the surface of the environment (e.g., depth is defined in a cylindrical or substantially cylindrical coordinate system, where the position of the user is at the center of a cylinder that extends from the user's head towards the user's feet). In some embodiments where depth is defined relative to the user's viewpoint (e.g., the direction relative to a point in space that determines which part of the environment is visible via a head-mounted device or other display), an object that is further away from the user's viewpoint along a line extending parallel to the direction of the user's viewpoint is considered to have a greater depth in the environment, and / or the depth of the object is measured along an axis that extends outward from the user's viewpoint and along a line parallel to the direction of the user's viewpoint (e.g., depth is defined in a spherical or substantially spherical coordinate system, where the origin of the viewpoint is at the center of a sphere that extends outward from the user's head). In some embodiments, depth is defined relative to a user interface container (e.g., a window or an application where an application and / or system content is displayed), where the user interface container has a height and / or a width, and depth is a dimension that is orthogonal to the height and / or width of the user interface container. In some embodiments where depth is defined relative to a user interface container, when the container is placed in a three-dimensional environment or is initially displayed (e.g., such that the depth dimension of the container extends outward away from the user or the user's viewpoint), the height and / or width of the container is generally orthogonal or substantially orthogonal to a straight line that extends from the position of the user (e.g., the user's viewpoint or the position of the user) to the user interface container (e.g., the center of the user interface container or another feature point of the user interface container). In some embodiments where depth is defined relative to a user interface container, the depth of an object relative to the user interface container refers to the position of the object along the depth dimension of the user interface container. In some embodiments, multiple different containers may have different depth dimensions (e.g., different depth dimensions that extend away from the user or the user's viewpoint in different directions and / or from different starting points).In some embodiments, when defining depth relative to a user interface container, the direction of the depth dimension remains constant for the user interface container as the position of the user interface container, the user, and / or the user's viewing point changes (e.g., or when multiple different viewers are viewing the same container in a three-dimensional environment, such as during an in-person collaboration session and / or when multiple participants are in a real-time communication session with shared virtual content that includes the container). In some embodiments, for a curved container (e.g., including a container having a curved surface or a curved content area), the depth dimension optionally extends into the surface of the curved container. In some cases, z-spacing (e.g., the spacing between two objects in the depth dimension), z-height (e.g., the distance of one object from another in the depth dimension), z-position (e.g., the position of one object in the depth dimension), z-depth (e.g., the position of one object in the depth dimension), or an analog z-dimension (e.g., depth used as a dimension of an object, a dimension of an environment, a direction in space, and / or a direction in an analog space) are used to refer to the concept of depth as described above.

[0202] In some embodiments, the user can optionally interact with virtual objects in a three-dimensional environment using one or both hands as if the virtual objects were real objects in a physical environment. For example, as described above, one or more sensors of a computer system optionally capture one or more hands of the user and display a representation of the user's hand(s) in the three-dimensional environment (e.g., in a manner similar to displaying real-world objects in the three-dimensional environment described above), or in some embodiments, the user's hand(s) can be seen via a display generation component, via the ability to see the physical environment through the user interface, due to the transparency / translucency of a portion of the user interface being displayed by the display generation component, or due to the projection of the user interface onto a transparent / translucent surface or onto the user's eyes or into the user's field of view. Thus, in some embodiments, the user's hands are displayed at corresponding positions in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment that can interact with virtual objects in the three-dimensional environment as if those virtual objects were physical objects in a physical environment. In some embodiments, the computer system can update the display of the representation of the user's hand(s) in the three-dimensional environment in conjunction with the movement of the user's hand(s) in the physical environment.

[0203] In some of the embodiments described below, the computer system is optionally able to determine an “effective” distance between a physical object in the physical world and a virtual object in a three-dimensional environment, e.g., for determining whether the physical object is directly interacting with the virtual object (e.g., whether a hand is touching, grasping, holding, etc. the virtual object or is within a threshold distance of the virtual object). For example, a hand directly interacting with a virtual object optionally includes one or more of the following: a finger of the hand pressing a virtual button, a hand of the user grasping a virtual vase, the hands of the user brought together and pinching / holding the user interface of an application, and two fingers performing any other type of interaction described herein. For example, when determining whether a user is interacting with a virtual object and / or how the user is interacting with the virtual object, the computer system optionally determines the distance between the user's hand and the virtual object. In some embodiments, the computer system determines the distance between the user's hand and the virtual object by determining the distance between the position of the hand in the three-dimensional environment and the position of the virtual object of interest in the three-dimensional environment. For example, the one or more hands of the user are located at a particular location in the physical world, and the computer system optionally captures the one or more hands and displays the one or more hands at a particular corresponding location in the three-dimensional environment (e.g., the location where the hand would be displayed in the three-dimensional environment if the hand were a virtual hand rather than a physical hand). Optionally, the location of the hand in the three-dimensional environment is compared with the location of the virtual object of interest in the three-dimensional environment to determine the distance between the one or more hands of the user and the virtual object. In some embodiments, the computer system optionally determines the distance between the physical object and the virtual object by comparing locations in the physical world (e.g., rather than comparing locations in the three-dimensional environment). For example, when determining the distance between one or more hands of the user and a virtual object, the computer system optionally determines the corresponding position of the virtual object in the physical world (e.g., the location where the virtual object would be located in the physical world if the virtual object were a physical object rather than a virtual object), and then determines the distance between the corresponding physical location and the one or more hands of the user. In some embodiments, the same techniques are optionally used to determine the distance between any physical object and any virtual object. Thus, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the computer system optionally performs any of the techniques described above to map the position of the physical object to the three-dimensional environment and / or to map the position of the virtual object to the physical environment.

[0204] In some embodiments, the same or similar techniques are used to determine where and what the user's gaze is directed at, and / or where and what a physical stylus held by the user is directed at. For example, if the user's gaze is directed at a specific location in the physical environment, the computer system optionally determines the corresponding location in the three-dimensional environment (e.g., the virtual location of the gaze), and if a virtual object is located at the corresponding virtual location, the computer system optionally determines that the user's gaze is directed at the virtual object. Similarly, the computer system is optionally able to determine the direction in the physical environment that the stylus is directed based on the orientation of the physical stylus. In some embodiments, based on this determination, the computer system determines the corresponding virtual location in the three-dimensional environment that corresponds to the location in the physical environment where the stylus is directed, and optionally determines that the stylus is directed at the corresponding virtual location in the three-dimensional environment.

[0205] Similarly, the embodiments described herein may refer to the position of a user (e.g., a user of a computer system) in a three-dimensional environment and / or the position of a computer system in a three-dimensional environment. In some embodiments, the user of the computer system is holding, wearing, or otherwise located at or near the computer system. Thus, in some embodiments, the position of the computer system serves as a proxy for the position of the user. In some embodiments, the position of the computer system and / or the user in the physical environment corresponds to the corresponding position in the three-dimensional environment. For example, the position of the computer system will be the position in the physical environment (and its corresponding position in the three-dimensional environment) from which the user would see the objects in the physical environment at the same location, orientation, and / or size (e.g., in an absolute sense and / or relative to each other) as the objects that are displayed or visible in the three-dimensional environment by the display generation component of the computer system, if the user were standing at that position and facing the corresponding portion of the physical environment visible via the display generation component. Similarly, if the virtual objects displayed in the three-dimensional environment are physical objects in the physical environment (e.g., physical objects placed at the same position in the physical environment as the position of these virtual objects in the three-dimensional environment, and having the same size and orientation in the physical environment as they do in the three-dimensional environment), the position of the computer system and / or the user is the location from which the user would see the virtual objects in the physical environment at the same location, orientation, and / or size (e.g., in an absolute sense and / or relative to each other and real-world objects) as the virtual objects displayed in the three-dimensional environment by the display generation component of the computer system.

[0206] In the present disclosure, various input methods are described with respect to interaction with a computer system. When one input device or input method is used to provide an example and another input device or input method is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the input device or input method described with respect to the other example. Similarly, various output methods are described with respect to interaction with a computer system. When one output device or output method is used to provide an example and another output device or output method is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the output device or output method described with respect to the other example. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment via a computer system. When interaction with a virtual environment is used to provide an example and a mixed reality environment is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the methods described with respect to the other example. Accordingly, the present disclosure discloses embodiments that are combinations of features of multiple examples without exhaustively listing all features of the embodiments in the description of each example embodiment.

[0207] User interface and associated processes

[0208] Attention is now turned to embodiments of a user interface ("UI") and associated processes that can be implemented on a computer system (such as a portable multifunctional device or a head-mounted device) having a display generation component, one or more input devices, and optionally one or more cameras.

[0209] 7A to 7O An example of a computer system that displays a gaze virtual object is shown according to some embodiments, where the gaze virtual object is selectable based on the attention directed to it to perform an operation associated with the selectable virtual object.

[0210] Fig. 7A A computer system 101 is shown displaying a three-dimensional environment 702 via a display generation component 120 (e.g., the display generation component 120 of FIG. 1) from the user's viewpoint (optionally facing the back wall of the physical environment in which the computer system 101 is located). As referred to above with reference to FIGS. 1 to Figure 6 As described, the computer system 101 optionally includes a display generation component 120 (e.g., a touch screen) and a plurality of image sensors (e.g., Figure 3of the image sensor 314). The image sensor optionally includes one or more of the following: a visible light camera; an infrared camera; a depth sensor; or any other sensor that can be used by the computer system 101 to capture one or more images of the user or a part of the user (e.g., one or more hands of the user) when the user interacts with the computer system 101. In some embodiments, the user interfaces illustrated and described below may also be implemented on a head-mounted display that includes a display generation component for displaying the user interface or a three-dimensional environment to the user, and sensors for detecting the movement of the physical environment and / or the user's hand (such as a movement that is interpreted by the computer system as a gesture such as an air gesture) (e.g., an external sensor facing away from the user), and / or sensors for detecting the user's gaze (e.g., an internal sensor facing towards the user's face).

[0211] As Fig. 7A shown, the computer system 101 captures one or more images of the physical environment (e.g., the operating environment 100) around the computer system 101, including one or more objects in the physical environment around the computer system 101. In some embodiments, the computer system 101 displays a representation of the physical environment in a three-dimensional environment 702, or a portion of the physical environment is visible via the display generation component 120 of the computer system 101. For example, the three-dimensional environment 702 includes portions of the left and right walls, ceiling, and floor in the user's physical environment. The three-dimensional environment 702 also optionally includes representations of physical objects such as tables and / or chairs in the physical environment.

[0212] In Fig. 7A it, the three-dimensional environment 702 also includes virtual objects, such as virtual object 704, virtual object 706, and virtual object 708. The virtual objects are optionally one or more of a user interface of an application (e.g., a messaging user interface or a content browsing user interface), a three-dimensional object (e.g., a virtual clock, a virtual ball, or a virtual car), or any other element not included in the physical environment of the computer system 101 that is displayed by the computer system 101. For example, as Fig. 7A shown, the virtual object 704 is optionally a user interface of a game application. In some embodiments, as Fig. 7A shown, the virtual object 704 includes a plurality of selectable virtual objects (e.g., affordances, buttons, toggles, icons, or photos). The virtual object 706 is optionally a menu user interface that includes a plurality of selectable virtual objects. The virtual object 708 is optionally a music player user interface that includes one or more selectable virtual objects for initiating playback of a first media item (“Soul Music Mix”).

[0213] In some embodiments, computer system 101 changes the visual appearance of a virtual object in response to detecting that a user's attention is directed to the virtual object. Additionally, in some embodiments, computer system 101 displays an attention indicator in the three-dimensional environment 702 in response to the user's attention being directed to an optional object in the three-dimensional environment 702, as detailed in reference methods 1200 and / or 2000. For example, from FIG. 7A to FIG. 7B , computer system 101 detects that the user's attention (e.g., attention input 728) is directed to virtual object 706b. In response, computer system 101 has displayed the boundary of a container object that encloses and / or is the container for virtual object 706b (e.g., a box). As Fig. 7A shown, prior to attention input 726 being directed to virtual object 706b, virtual object 706b was not displayed within or with the boundary of the container object. In some embodiments, computer system 101 displays a container object of a suitable size based on the size of the virtual object. For example, in Figure 7B , computer system 101 detects that the user's attention (e.g., attention input 722) is directed to virtual object 704g. In response, computer system 101 displays the boundary of a container object of a suitable size to enclose virtual object 704g. In some embodiments, computer system 101 displays a container object of a suitable size based on the visual appearance of a group of virtual objects. For example, in Figure 7B , computer system 101 detects that the user's attention (e.g., attention input 720) is directed to virtual object 704c. In response, computer system 101 displays the boundary of a container object to enclose virtual object 704c, which is similar or identical in size to other virtual objects in its group (e.g., objects 704a and 704b). In some embodiments, computer system 101 displays a container object of a suitable size based on the size of an associated container object. For example, in Figure 7B , computer system 101 detects that the user's attention (e.g., attention input 730) is directed to virtual object 708a. In response, computer system 101 displays the boundary of a container object to enclose virtual object 708a, which is smaller in size than virtual object 708.

[0214] In some embodiments, computer system 101 does not display or displays the boundary of a container object to enclose a virtual object that already includes, is displayed within, or is the container object. For example, from FIG. 7A to FIG. 7B, computer system 101 detects that the user's attention (e.g., attention input 726) is directed to virtual object 706a. In response, computer system 101 does not display a new container object that would enclose virtual object 706a because virtual object 706a was already enclosed by a container object before attention input 726 was directed to virtual object 706a. Similarly and in another example, computer system 101 detects that the user's attention (e.g., attention input 716) is directed to virtual object 704a. In response, computer system 101 does not display a new container object that would enclose virtual object 704a because virtual object 704a was already enclosed by a container object before attention input 716 was directed to virtual object 704a. Reference is made to methods 800, 900, 1000, 1200, and / or 2000 for additional details regarding changing the visual appearance of a virtual object in response to detecting the user's attention directed to the virtual object. Additionally, although Figure 7B (and other figures) illustrate multiple concurrent attention inputs directed to objects in three-dimensional environment 702, it should be understood that such inputs are optionally alternative inputs rather than concurrent inputs. Additionally, in some embodiments, the input to computer system 101 is provided via an air gesture from hand 710 and / or the user's attention (e.g., as detailed with reference to method 800) or via touchpad 746 from hand 710, and the inputs described herein are optionally received via touchpad 746 or via an air gesture / attention.

[0215] In some embodiments, computer system 101 displays a virtual object with a larger expanded size in response to detecting the user's attention directed to the virtual object. For example, from FIG. 7B to FIG. 7C , computer system 101 detects that the user's attention (e.g., attention input 724) is directed to virtual object 704h. For example, virtual object 704h is a messaging communication module that includes text content corresponding to a game update or a text message from another player within a game application or a link to a website. In response, computer system 101 displays virtual object 704h with a larger size, as Figure 7C shown. In Figure 7C , due to its larger expanded size, virtual object 704h includes more text content (e.g., more portions of the text content corresponding to the game update or text message, or more portions of the URL of the link to the website) as Figure 7B shown when the user's attention is diverted from virtual object 704h. In some embodiments, the user's attention being directed to a virtual object is determined based on the detection of a gaze directed to the virtual object under one or more conditions described in reference methods 800, 900, 1000, 1200, and / or 2000.

[0216] In some embodiments, computer system 101 provides audio feedback in response to detecting that the user's attention is directed to a virtual object. For example, timer 712 corresponds to virtual object 704c and is used to indicate the amount of time that the user's attention is directed to virtual object 704c. For example, from FIG. 7B to FIG. 7C , computer system 101 detects that the user's attention (e.g., attention input 720) is directed to virtual object 704c for a period of time greater than time threshold 712b, as detailed in reference methods 800, 900, 1000, 1200, and / or 2000. In response, computer system outputs audio feedback to indicate that the user's attention is directed to virtual object 704c. Reference methods 800, 900, 1000, 1200, and / or 2000 provide additional details regarding providing audio feedback in response to detecting the user's attention directed to a virtual object.

[0217] In some embodiments, computer system 101 displays a gaze virtual object that can be selected based on the attention directed to a virtual object to perform an operation associated with the selectable virtual object. For example, in Figure 7C , computer system 101 detects that the user's attention (e.g., attention input 716) is directed to virtual object 704a for a period of time greater than a time threshold (e.g., time threshold 712b), as detailed in reference methods 800, 900, 1000, 1200, and / or 2000. In response, computer system 101 displays a gaze virtual object 704a' associated with virtual object 704a within the container object of virtual object 704a, as Figure 7C shown. In response to detecting that the user's attention directed to other selectable objects lasts longer than the time threshold, computer system 101 additionally displays the gaze virtual objects of these selectable objects in the three-dimensional environment 702. In some embodiments, computer system 101 displays the gaze virtual object at a location that is different from but co-located with its associated virtual object. For example, in Figure 7C , computer system 101 displays a gaze virtual object 706b' associated with virtual object 706b below virtual object 706b and within its own container. In some embodiments, as Figure 7C shown, gaze virtual object 706a' is displayed within the container object of the content (e.g., tooltip) associated with the virtual object (e.g., virtual object 706a). As Figure 7CAs shown, the gaze virtual object 706a' is within the tooltip for the virtual object 706a (e.g., the description of the virtual object 706a and / or the function performed by selecting the object 706a) but in the right region, and the virtual object 706a is also displayed in response to the attention being directed to the object for longer than a time threshold. For another example, the computer system 101 places the gaze virtual object associated with the virtual object 704g in the lower right region of the container of the virtual object 704g that is displayed in response to the attention being directed to the virtual object 704g. For yet another example, the computer system 101 displays the gaze virtual object 704g' associated with the virtual object 708a next to the virtual object 708a and within its own container. Different from the gaze virtual object 706b' associated with the virtual object 706b being presented as covering the boundary of the virtual object 706, the gaze virtual object 708a' associated with the virtual object 708a is contained within the virtual object 708. In Figure 7C , the virtual object 704 includes a toggle virtual object 704b that will be described below.

[0218] In some embodiments, the computer system 101 displays the gaze virtual object as including a visual indication of the progress towards the user's attention meeting one or more criteria (described in reference methods 800, 900, 1000, and 1200), the one or more criteria being for using attention to select the gaze virtual object and thus for performing an operation associated with the virtual object corresponding to the gaze virtual object. For example, the computer system 101 optionally indicates the visual indication of the progress by filling the unfilled portion of the gaze virtual object. As Figure 7C shown, the gaze virtual object 704a' associated with the virtual object 704a is unfilled, indicating that the user's attention (e.g., the attention input 716) is away from the gaze virtual object 704a'. From FIG. 7C to FIG. 7D , the user's attention (the attention input 716) changes from being directed to the virtual object 704a to being directed to the gaze virtual object 704a' associated with the virtual object 704a. In response, as Fig.7D shown, the computer system 101 displays that the gaze virtual object 704a' has been filled corresponding to the amount of time indicated by the timer 734, with the user's attention directed to the gaze virtual object 704a' associated with the virtual object 704a. In some embodiments, the computer system 101 optionally similarly displays Figure 7C such a visual indication of the progress in other gaze virtual objects (e.g., the gaze virtual objects 706a', 706b', 706a', 708a', and / or 704g').

[0219] In some embodiments, computer system 101 provides audio feedback indicating the progress of directing the user's attention (pointing the gaze object) towards meeting one or more criteria (as described in reference methods 800, 900, 1000, and 1200) for performing an operation associated with a corresponding virtual object. For example, in Fig.7D , computer system 101 detects that the user's attention (e.g., attention input 716) has been directed towards the gaze virtual object 704a' associated with the virtual object 704a for a period greater than a first threshold 734b and a second threshold 734b. In response, computer system 101 outputs audio feedback having sound characteristics that change (e.g., volume and / or pitch) with the duration of the attention input 716 directed towards the gaze virtual object 704a' associated with the virtual object 704a, as Fig.7D shown. Thus, in some embodiments, computer system 101 continuously outputs audio feedback whose characteristics (e.g., pitch, volume, and / or tone, as detailed in reference methods 800, 900, and / or 1000) change as the duration of the attention 716 directed towards the gaze virtual object 704a changes (e.g., as indicated by timer 734).

[0220] In some embodiments, in response to detecting that the user's attention has become away from the virtual object 704g, as FIG. 7C to FIG. 7D shown, computer system 101 stops displaying the container object that houses the virtual object 704g. If computer system 101 detects that the user's attention (e.g., attention input 716) changes back to being directed towards the virtual object 704g, then computer system 101 optionally redisplay the container object to house the virtual object 704g, as Fig. 7E shown. Fig. 7E A timer 736 below the timer threshold 736b is also shown, which indicates that the period of the user's attention (e.g., attention input 716) directed towards the virtual object 704g is less than the amount of time required to display the gaze virtual object 704g' for the object 704g. Fig. 7E A visual indication of the progress of the gaze virtual object 704a' towards changing (e.g., decreasing in fill amount) as the user's attention towards meeting one or more criteria (as described in reference methods 800, 900, 1000, and 1200) for selecting the gaze virtual object 704a' is also shown. When the period of the user's attention (e.g., attention input 716) is less than the timer threshold 736b, computer system 101 optionally does not display the gaze virtual object 704g' associated with the virtual object 704g.

[0221] In some embodiments, if the attention input indicates that the user's attention has become diverted from the fixation virtual object (or corresponding virtual object) and then moved back to the fixation virtual object (or corresponding virtual object), the computer system 101 updates the visual indication of the progress of the user's attention towards meeting one or more criteria for performing an operation associated with the corresponding virtual object based on whether the attention input moves back to the object within a threshold time (e.g., 0.03 seconds, 0.05 seconds, 0.07 seconds, 0.09 seconds, 0.1 seconds, 0.15 seconds, 0.2 seconds, 0.25 seconds, 0.3 seconds, 0.5 seconds, 1 second, 3 seconds, 5 seconds, 10 seconds, 15 seconds, 20 seconds, or 30 seconds). For example, in Figure 7F because the user's attention (e.g., attention input 716) changes from the fixation virtual object 704a' that is diverted from the virtual object 704a to the fixation virtual object 704a' that points to the virtual object 704a within the time threshold, the visual indication of the progress in the fixation target for the object 704a is maintained (e.g., continues from the filled amount before the change in the user's attention). In some embodiments, the computer system 101 updates the visual indication of the progress, which is different from maintaining the indication of the progress as will be described in the following figures.

[0222] In some embodiments, if the attention input includes an activation input (e.g., as detailed with reference to methods 800, 900, and 1000), the computer system 101 performs an operation associated with the virtual object. For example, in Figure 7F the computer system 101 detects an activation input (e.g., a finger of the hand 710 touches the touchpad 746 and / or an air pinching gesture from the hand 710) before the attention 716 meets one or more criteria for the fixation virtual object 704a' for performing an operation associated with the virtual object as described in reference to methods 800, 900, and 1000 for the object 704a (e.g., before the duration of the attention input 716 that points to the fixation virtual object 704a' associated with the virtual object 704a reaches the threshold 734d as shown by the timer 734 in Figure 7F . In response to detecting the activation input before meeting one or more criteria (e.g., reaching the threshold 734d), the computer system 101 optionally performs an operation associated with the virtual object 704a, such as updating the right side of the virtual object 704 to include content associated with the virtual object 704a, as shown in Figure 7G .

[0223] Figure 7G1 shows concepts similar and / or identical to the concepts shown in Figure 7G (with many of the same reference numerals). It should be understood that unless otherwise indicated below, Figure 7G1 shown has the same as 7A to 7OElements shown with the same reference numeral have one or more or all of the same characteristics. Figure 7G1 The computer system 101 includes a display generation component 120 (or the same as the display generation component 120). In some embodiments, the computer system 101 and the display generation component 120 each have 7A to 7O The computer system 101 shown in FIG. Figure 3 One or more of the characteristics of the display generation component 120 shown, and in some embodiments, 7A to 7O The computer system 101 and display generation component 120 shown have Figure 7G1 One or more of the characteristics of computer system 101 and display generation component 120 are shown.

[0224] exist Figure 7G1 , the display generation component 120 includes one or more internal image sensors 314a oriented toward the user's face (e.g., reference Figure 5 The eye tracking camera 540 is described above. In some embodiments, the internal image sensor 314a is used for eye tracking (e.g., detecting the user's gaze). The internal image sensor 314a is optionally arranged on the left and right portions of the display generation component 120 to enable eye tracking of the user's left and right eyes. The display generation component 120 also includes external image sensors 314b and 314c facing outward from the user to detect and / or capture the physical environment and / or the movement of the user's hand. In some embodiments, the image sensors 314a, 314b, and 314c have reference 7A to 7O One or more of the characteristics of the image sensor 314 described.

[0225] exist Figure 7G1 , the display generation component 120 is shown as displaying optionally corresponding to the reference 7A to 7O Content is described as content displayed and / or visible via display generation component 120. In some embodiments, the content is displayed by a single display (e.g., Figure 5 In some embodiments, display generation component 120 includes a display that is combined (e.g., by the user's brain) to create Figure 7G1 Two or more displays (e.g., left and right display panels for the user's left and right eyes, respectively, as shown in FIG. 1 ) that display outputs of the views of the content shown. Figure 5 described above).

[0226] The display generation component 120 has a corresponding Figure 7G1The field of view of the content shown (e.g., the field of view captured by external image sensors 314b and 314c and / or visible to the user via the display generation component 120, indicated by the dashed line in the top view). Since the display generation component 120 is optionally part of a head-mounted device, the field of view of the display generation component 120 is optionally the same as or similar to the user's field of view.

[0227] In Figure 7G1 , the user is depicted as performing (e.g., with hand 710) an air pinch gesture to provide input to the computer system 101, thereby providing user input that points to the content displayed by the computer system 101. Such descriptions are intended to be exemplary and not restrictive; the user optionally uses different air gestures and / or uses other forms of input as referenced in 7A to 7O to provide user input.

[0228] In some embodiments, the computer system 101 responds to user input as referenced in 7A to 7O .

[0229] In Figure 7G1 's example, since the user's hand is within the field of view of the display generation component 120, it is visible within the three-dimensional environment. That is, the user can optionally see any part of their own body within the field of view of the display generation component 120 in the three-dimensional environment. It should be understood that one or more or all aspects of the present disclosure as shown in 7A to 7O or described and / or referenced in connection with the corresponding methods are optionally implemented on the computer system 101 and the display generation unit 120 in a manner similar or analogous to that shown in Figure 7G1 .

[0230] In some embodiments, in response to detecting that the most recent interaction with the computer system includes only attention input or non-attention input as detailed in reference method 1000, the computer system 101 optionally changes (e.g., shortens or lengthens) the threshold requirement for displaying a virtual object in a gaze. For example, after detecting that the attention input includes only attention input (e.g., does not include non-attention input, such as input from hand 710), the computer system 101 shortens the time threshold for displaying the virtual object in a gaze to 732b' (less than Fig. 7E 's threshold 736b), as shown in Figure 7F . As another example, after detecting that the most recent interaction with the computer system 101 includes non-attention input, the computer system 101 optionally lengthens the time threshold for displaying the virtual object in a gaze to 738b (greater than Fig. 7E 's threshold 736b), as shown in Figure 7G .

[0231] In some embodiments, if an attention input (e.g., without a separate activation input) meets one or more criteria for performing an operation associated with a virtual object, computer system 101 performs the operation associated with the virtual object. For example, in Figure 7H , computer system detects an attention input (e.g., attention input 716 without a separate activation input) that is directed at a gaze virtual object 704b associated with a virtual object 704b. As Figure 7H shown, virtual object 704b is a toggle virtual object that toggles between an invisible mode and a non-invisible mode in a game application user interface 704 presented in a three-dimensional environment 702. In some embodiments, computer system 101 displays the gaze virtual object 704b associated with the toggle virtual object 704b as overlaid on the toggle virtual object 704b, as Figure 7H shown. In some embodiments, computer system 101 displays the gaze virtual object 704b to the right of the toggle virtual object 704b, as Figure 7H shown (e.g., because the toggle object 704b will move to the right when toggled). If computer system 101 detects that the user's attention (e.g., attention input 716 directed at the gaze virtual object 704b) has met the criteria for performing an operation associated with the toggle virtual object 704b, computer system 101 displays the toggle button of the toggle virtual object 704b and / or moves it to the right. In some embodiments, computer system 101 then displays the gaze virtual object 704b to the left of the toggle virtual object 704b (e.g., because the toggle object 704b will move to the left when toggled).

[0232] In Figure 7H , computer system 101 displays that the gaze virtual object 704b associated with the toggle virtual object 704b is filled approximately 75% corresponding to the duration indicated by a timer 740. As Fig.7I shown, the user's attention (e.g., attention input 716) has reached a threshold 734d corresponding to meeting one or more criteria for performing an operation associated with the toggle virtual object 704b as indicated by the timer 740. Fig.7I Also shown is that the gaze virtual object 704b associated with the toggle virtual object 704b is filled 100% to visually represent meeting one or more criteria for performing an operation associated with the toggle virtual object 704b. In response to meeting one or more criteria for performing an operation associated with the toggle virtual object 704b, computer system optionally performs the operation associated with the toggle virtual object 704b, as indicated by a changed state of the toggle virtual object 704b as compared to Fig.7I compared to Figure 7J shown.

[0233] In some embodiments, in response to detecting that the most recent interaction with the computer system 101 includes only an attention input or a non-attention input as detailed in reference method 800, the computer system optionally changes (e.g., shortens or lengthens) the threshold requirements for performing operations associated with a virtual object. For example, instead of waiting for the user's attention to meet the threshold 734d indicated by the timer 740 in Fig.7I to cause the selection of the fixated virtual object 704b', the computer system 101 optionally shortens the threshold for selection to threshold 734c, which is less than threshold 734d, based on determining that the most recent interaction with the computer system 101 includes only an attention input.

[0234] In some embodiments, the virtual object includes scrollable content (e.g., continuous content that cannot all be displayed simultaneously). For example, in Figure 7K , the computer system 101 detects an input including the user's attention 716 directed at the fixated virtual object 704e associated with the virtual object 704e' and an activation input (e.g., a finger of the hand 710 touching the touchpad 746 or an air pinch gesture from the hand 710). From Figures 7K to 7L , the input from the hand 710 corresponds to an input for scrolling the content of the user interface (e.g., a movement of the hand 710 in an upward direction). In response to the input from the hand 710, the computer system 101 scrolls the content of the virtual object 704 to display more content, as shown in Figure 7L . As shown in Figure 7L , in response to scrolling and / or during scrolling, the fixated virtual object 704e' associated with the virtual object 704e is no longer displayed.

[0235] In some embodiments, when the user's attention becomes away from the fixated virtual object and / or the virtual object, the computer system 101 stops displaying the fixated virtual object. For example, in Figure 7M , when the computer system 101 displays the fixated virtual object 704d' of the virtual object 704d in response to the user's attention (e.g., attention input 716) directed at the virtual object 704d and the user's attention meeting one or more criteria (as further indicated by meeting the threshold 748b in the timer 748 shown in Figure 7M ), the computer system 101 detects that the attention input 716 becomes away from the virtual object 704d and directed at the virtual object 708a in Figure 7N . In response, the computer system 101 displays the container housing the virtual object 708a and simultaneously stops displaying the fixated virtual object 704d' of the virtual object 704d, as shown in Fig.7O . From FIG. 7N to FIG. 7O, the timer 752 corresponding to the duration of the attention input 716 pointing to the virtual object 708a has increased but does not meet the threshold 752b, so the computer system 101 does not display the gaze virtual object 708a' of the virtual object 708a.

[0236] FIG. 8A to FIG. 8I is a flowchart showing an exemplary method 800 according to some embodiments, which method displays a gaze virtual object that can be selected based on the attention pointing to the gaze virtual object to perform an operation associated with an optional virtual object. In some embodiments, the method 800 is executed at a computer system (e.g., the computer system 101 in FIG. 1, such as a tablet computer, a smart phone, a wearable computer, or a head-mounted device), the computer system including a display generation component (e.g., the display generation component 120 in FIGS. 1, Figure 3 and Figure 4 such as a head-up display, a display, a touch screen, a projector, etc.) and one or more cameras (e.g., a camera pointing downward at the user's hand (e.g., a color sensor, an infrared sensor, and other depth-sensing cameras) or a camera pointing forward from the user's head). In some embodiments, the method 800 is managed by instructions stored in a non-transitory computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of the computer system 101 (e.g., Figure 1A the control unit 110 in FIG. 1). Some operations in the method 800 are optionally combined, and / or the order of some operations is optionally changed.

[0237] In some embodiments, the method 800 is executed at a computer system (e.g., 101) that communicates with a display generation component (e.g., 120) and one or more input devices (e.g., a gaze tracking device, a hand tracking device, a remote control, one or more touch-sensitive surfaces, one or more buttons, a dial, and / or a knob), the computer system being, for example, a mobile device (e.g., a tablet computer, a smart phone, a media player, or a wearable device) or a computer or other electronic device. In some embodiments, the display generation component is a display integrated with the electronic device (optionally a touch screen display), an external display such as a monitor, a projector, a television, or a hardware component (optionally integrated or external) for projecting a user interface or making the user interface visible to one or more users. In some embodiments, the computer system communicates with a gaze tracking device (e.g., Figure 5communicate with the eye tracking device 130) in. In some embodiments, one or more input devices include those capable of receiving user input (e.g., capturing user input or detecting user input) and sending information associated with the user input to the electronic device. Examples of input devices include touchscreens, mice (e.g., external), touchpads (optionally integrated or external), trackpads (optionally integrated or external), remote control devices (e.g., external), another mobile device (e.g., separate from the electronic device), handheld devices (e.g., external), controllers (e.g., external), cameras, depth sensors, eye tracking or gaze tracking devices, and / or motion sensors (e.g., hand tracking devices or hand motion sensors). In some embodiments, the computer system communicates with a gaze tracking device. In some embodiments, the gaze tracking device is a wearable device, such as a head-mounted device described in more detail herein. In some embodiments, the gaze tracking device does not need to be implemented in a head-mounted or near-eye manner as described herein.

[0238] In some embodiments, the computer system displays (802a) via a display generation component a user interface including a first selectable user interface object for performing a first operation, such as Fig. 7AThe three-dimensional environment 702 therein. In some embodiments, the user interface is a three-dimensional environment (e.g., the three-dimensional environment is an extended reality (XR) environment, such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment). In some embodiments, the user interface is the user interface of an application accessible by a computer system, such as a word processing application with multiple texts, an application launch user interface with multiple application icons, a photo management application with multiple photo representations, a spreadsheet application with multiple data units, a presentation application with multiple slides or other graphical user interface objects, a messaging application with multiple messages, a web browsing application with multiple links, and / or an email application with multiple emails. In some embodiments, the user interface includes multiple user interface objects, including affordance representations, buttons, icons, bubbles, trays, or other containers for text (e.g., hyperlinks and / or graphics), messages (e.g., text and / or graphics), images, or multimedia, and these user interface objects are selectable to display the corresponding user interface (e.g., page) and / or perform an operation associated with selecting the affordance representation (e.g., playing a video or launching an application). In some embodiments, the user interface is a bookmark (e.g., favorites and / or Internet shortcuts) management user interface with multiple web page representations saved by a user of the computer system. In some embodiments, the first selectable user interface object is a representation of a web page and includes the corresponding URL text displayed within the representation of the web page in the user interface. In some embodiments, the user interface is an operation menu with multiple setting representations, and these setting representations define how the computer system operates, what to display, and / or how to display user interface objects. In some embodiments, the first operation associated with the first selectable user interface object includes displaying content, displaying a web page, displaying another user interface, playing multimedia, launching an application, providing a menu, installing a program, or downloading content. In some embodiments, as described in step 812 and method 1000, the first selectable user interface object is selected in response to a combination of user attention and receiving an activation input (e.g., a user input confirming the intention to perform the first operation). In some embodiments, user attention corresponds to user gaze, as referenced Figure 6As detailed. In some embodiments, receiving an activation input includes detecting a portion of a user (e.g., a hand, an arm, and / or a finger) performing an air pinching gesture (e.g., two or more fingers of the user's hand, such as the thumb and index finger, moving together and touching each other) to form a pinched hand shape while the user's attention is directed to the user interface and / or the first optional user interface object, and then moving the hand in the pinched hand shape upward or downward. In some embodiments, the activation input corresponds to a gesture other than the air pinching gesture, such as a forward pointing gesture (e.g., a forward movement of the user's hand when one or more fingers of the user's hand extend toward the first optional user interface object) or a tapping gesture using the fingers of the user's hand (e.g., a forward movement made by the fingers of the user's hand such that the fingers touch the first optional user interface object or the user interface or are close within a threshold distance of the first optional user interface object or the user interface area). In some embodiments, the pinch-and-drag gesture as an air gesture includes a pinch gesture performed in combination with (e.g., following) a drag input that changes the position of the user's hand from a first position (e.g., the start position of the drag) to a second position (e.g., the end position of the drag). In some embodiments, the user maintains the pinched hand shape while performing the drag input and releases the pinch gesture (e.g., spreads two or more of their fingers) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., the user pinches two or more fingers together and moves the same hand to a second position in the air using the drag gesture). In some embodiments, the pinch input is performed by the user's first hand, and the drag input is performed by the user's second hand (e.g., while the user continues the pinch input with the user's first hand, the user's second hand moves from a first position to a second position in the air). In some embodiments, the activation input corresponds to an only-attention input for scrolling through the user interface as described in step 834, such as directing attention to the bottom portion or the top portion of the user interface causing the user interface to scroll its content downward or upward, respectively. In some embodiments, the activation input includes a touchpad input (e.g., a finger touching the touchpad) or an input device input (e.g., a selection via a handheld input device such as a stylus or a remote control). In some embodiments, the activation input is an only-attention and / or only-gaze input (e.g., does not include an input from one or more portions of the user other than those providing the attention input).

[0239] In some embodiments, when displaying the user interface, the computer system detects (802b) the attention of the user of the computer system being directed to the first optional user interface object via one or more input devices, such as Figure 7G and Figure 7G1Attention 716 directed to object 704b. In some embodiments, when the user's attention corresponds to a gaze, the gaze tracking device optionally captures one or more images of the user's eyes and detects pupils and glints in the one or more captured images to track the user's gaze, as detailed in reference Figure 6 described. In some embodiments, the computer system detects the user's gaze directed to a location (or region) of a user interface including a first optional user interface object during a first time period greater than a first time threshold (e.g., 0.02 seconds, 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.25 seconds, 0.3 seconds, 0.5 seconds, 1 second, 2 seconds, 3 seconds, or 5 seconds). In some embodiments, the first optional user interface object is initially displayed with a first visual appearance having a first shape, a first position, a first color, and / or a first effect. In some embodiments, when the computer system detects that the user's attention directed to the location of the user interface including the first optional user interface object persists for greater than the first time threshold, the computer displays the first optional user interface object with a second visual appearance different from the first visual appearance. The second visual appearance optionally makes the first optional user interface object more prominent. For example, the second visual appearance of the first optional user interface object optionally includes a second shape greater than the first shape, a second position closer to the user's (or the computer system's) viewpoint than the first position, a second color brighter than the first color, and / or a second effect in which the first optional user interface object appears to lift more from the backing of the user interface than the first effect (e.g., the second effect visually indicates the first optional user interface object emphasized by a depth effect). In some embodiments, the second visual appearance includes presenting additional information (e.g., a tooltip, an information tip, or an indication) related to the first optional user interface object not shown with the first visual appearance.

[0240] In some embodiments, in response to detecting the user's attention directed to the first optional user interface object, the computer system displays (802c) a first gaze target associated with the first optional user interface object in the user interface, such as Figure 7Hthe fixation target 704b' therein. In some embodiments, in response to detecting that the user's fixation is directed to the first selectable user interface object and the user's fixation is not directed to a location (or region) of the user interface that does not include the first selectable user interface object, the computer system displays a first fixation target associated with the first selectable user interface object. In some embodiments, if the user's fixation is directed to a location (or region) of the user interface that does not include the first selectable user interface object, the computer system does not display (or stops displaying) the first fixation target. In some embodiments, the first fixation target is an entity (element / user interface object) that is different and / or separate from the first selectable user interface object. In some embodiments, the first fixation target is presented at a location different from the location of the first selectable user interface object. In some embodiments, the first fixation target is displayed at a location having a certain spatial relationship (e.g., distance and / or orientation) with the first selectable user interface object, such as 0.1 cm, 0.3 cm, 0.5 cm, 1 cm, 3 cm, 5 cm, 10 cm, 30 cm, or 50 cm above, below, to the left, or to the right of the first selectable user interface object. The location of the first fixation target will be described in more detail later with reference to steps 810, 840, and 842. In some embodiments, the first fixation target is initially displayed with a first visual appearance and / or a first sound effect, which will be described in more detail with reference to method 800. In some embodiments, the first fixation target indicates the fixation duration directed to the first fixation target, and once the fixation duration is greater than a second time threshold (e.g., 0.1 second, 0.5 second, 1 second, 2 seconds, 3 seconds, 5 seconds, 7 seconds, 10 seconds, 20 seconds, 30 seconds, or 60 seconds), the computer system initiates a first operation associated with the first selectable user interface object, as will be described later with reference to method 1000.

[0241] In some embodiments, when the first fixation target is displayed, the computer system detects (802d) that the user's attention (e.g., based on fixation) is directed to the first fixation target, such as Figure 7HAttention 716 therein. In some embodiments, the computer system detects a user's gaze directed to a location (or region) of the user interface that includes a first fixation target during a first time period greater than a first time threshold. In some embodiments, the user's gaze changes from a location of the user interface that includes a first optional user interface object to a location of the user interface that includes the first fixation target. The respective locations of the first optional user interface object and the first fixation target will be described in more detail with reference to steps 810, 840, and 842. In some embodiments, the computer system updates one or more visual characteristics (e.g., size, fill, color, or opacity) of the first fixation target in response to detecting the user's gaze directed to the first fixation target. For example, when the computer system detects the user's gaze directed to a location (or region) of the user interface that includes the fixation target during a second time period (e.g., 0.2 seconds, 0.25 seconds, 0.3 seconds, 0.5 seconds, 1 second, 2 seconds, 3 seconds, 5 seconds, 7 seconds, or 10 seconds) greater than the first time period (e.g., 0.02 seconds, 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.25 seconds, 0.3 seconds, 0.5 seconds, 1 second, 2 seconds, 3 seconds, or 5 seconds), the computer system optionally displays the first fixation target in a second visual appearance different from the first visual appearance (e.g., the second visual appearance includes a larger filled portion than the first visual appearance). For example, the fixation target is optionally a circle having a filled portion and an unfilled portion. The filled portion optionally expands (grows) outward (toward the outer edge of the circle) in response to a sustained gaze toward the fixation target. In some embodiments, the fixation target is similar to a pie chart, where the filled portion includes one or more wedges, and the angle of the wedges grows as the gaze continues to be directed to the fixation target. For example, the wedges optionally expand clockwise (or counterclockwise). In some embodiments, the fixation target is a progress bar having a rectangular shape that includes a filled portion and an unfilled portion, where the filled portion represents the duration of the gaze directed to the fixation target. Additional visual characteristics and sound effects associated with the fixation target will be described in more detail with reference to steps 818, 844 - 848, and method 900.

[0242] In some embodiments, in response to detecting that the user's attention is directed to the first fixation target (802e), and based on determining that the user's attention (directed to the first fixation target) meets one or more criteria, the computer system initiates a first operation (802f) associated with the first optional user interface object, such as in response to attention 716 being directed to Fig.7I and Figure 7JIn some embodiments, the one or more criteria include a criterion that is satisfied when the duration of the user's gaze directed to the first gaze target is greater than a second time threshold (e.g., 0.1 seconds, 0.5 seconds, 1 second, 2 seconds, 3 seconds, 5 seconds, 7 seconds, 10 seconds, 20 seconds, 30 seconds, or 60 seconds). In some embodiments, the first operation is confirmed to be initiated at the first gaze target based on the user's gaze directed to the first gaze target continuing for a time period greater than the second time threshold (e.g., the duration of the gaze in the direction of the first gaze target exceeds the second time threshold).

[0243] In some embodiments, based on determining that the user's attention does not meet one or more criteria, the computer system abandons (802g) initiating a first operation associated with the first selectable user interface object, such as with respect to Figure 7C As shown, the attention 716 moves away from the gaze target 704a'. For example, if the computer system determines that the gaze duration is less than the second time threshold due at least in part to the gaze moving to a position (or area) of the user interface that does not include the first gaze target, the computer system optionally abandons initiating the first operation associated with the first optional user interface object. In some embodiments, the computer system changes the visual appearance of the first gaze target in response to determining that the gaze is no longer directed at the first gaze target, as will be described in detail later with reference to steps 850, 852 and method 900. Displaying the gaze target in response to determining that the user's gaze is directed to the optional user interface object provides confirmation that the user intends to interact with the optional user interface object, thereby reducing errors in the interaction between the user and the computer system (e.g., avoiding accidental activation or deactivation of the optional user interface object due to unintentional gaze) and reducing the input required to correct such errors.

[0244] In some embodiments, determining that the user's attention is directed to the first selectable user interface object includes determining that the user's gaze has been directed to the first selectable user interface object for more than a first threshold time period (804), such as Figure 7CThe threshold 712b therein (e.g., 0.02 seconds, 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.25 seconds, 0.3 seconds, 0.5 seconds, 0.6 seconds, 0.7 seconds, 0.8 seconds, 0.9 seconds, 1 second, 1.2 seconds, 1.5 seconds, 2 seconds, 2.5 seconds, 3 seconds, 5 seconds, 10 seconds, or 30 seconds). In some embodiments, based on the detection of a gaze directed at a first optional user interface object without requiring any conditions, it is determined that the user's attention is directed at the first optional user interface object. In some embodiments, based on the detection of a gaze directed at a first optional user interface object under one or more conditions, such as detecting that the user's gaze is directed at a location (or region) of the first optional user interface object for a time period exceeding the first threshold period described herein for the computer system to display a first gaze target, it is determined that the user's attention is directed at the first optional user interface object. For example, the computer system determines that the user's gaze has been maintained in the region of the first optional user interface object for at least a time amount greater than the first threshold period. In some embodiments, the user's gaze (or the exact position of the user's gaze) directed at the first optional user interface object changes but is determined by the computer system to be continuously within the region of the first optional user interface object during at least that time amount. In some embodiments, the computer system does not display the first gaze target based on determining that the user's attention is not directed at the first optional user interface object (e.g., based on determining that the user's gaze has not been maintained in the region of the first optional user interface object for at least the first threshold period). In some embodiments, one or more additional conditions, such as requiring a gaze to be directed at the first optional user interface object or the first gaze target to perform a corresponding operation, are discussed in more detail with reference to step 802. Requiring the user's attention to be directed at the first optional user interface object for a period exceeding the first threshold period before displaying the gaze target provides additional confirmation of the user's intention to interact with the optional user interface object without cluttering the user interface with additional user interface objects (e.g., gaze targets), thereby enhancing the operability of the computer system and reducing the power consumption of the computer system.

[0245] In some embodiments, displaying the first optional user interface object includes displaying the first optional user interface object (806a) with a first characteristic having a first value, such as Figure 7C the object 704a therein. In some embodiments, displaying the first gaze target includes displaying the first gaze target (806b) with a first characteristic having a second value different from the first value, such as Figure 7CA fixation target 704a' that is smaller than the object 704a. In some embodiments, the first characteristic includes size and / or any visual characteristic having a first corresponding value, such as color, saturation, and / or brightness. In some embodiments, the computer system displays the first fixation target with a first characteristic (e.g., size) having a second value different from the first value (e.g., a size smaller than that of the first alternative user interface object). More details regarding the visual characteristics of the fixation target are described in steps 810, 818, 840, 842 - 852, and method 900. Displaying the fixation target smaller than the alternative user interface object minimizes distractions in the user interface, thereby reducing errors in the interaction between the user and the computer system (e.g., avoiding accidentally activating or deactivating the alternative user interface object due to unintentional fixation) and reducing the input required to correct such errors.

[0246] In some embodiments, in response to detecting the user's attention directed to the first alternative user interface object, such as Figure 7B the attention 724 on the object 704h in, the computer system increases (808) the size of the first alternative user interface object and displays information associated with the first alternative user interface object within the first alternative user interface object, where the information is not displayed before the user's attention is directed to the first alternative user interface object, such as Figure 7Cas shown by object 704h in. For example, the computer system displays the first selectable user interface object with a larger size based on determining that the user's attention is directed to the first selectable user interface object for a first time period greater than a first time threshold, as described in reference step 802. In some embodiments, the information displayed includes supplementary content and / or provides one or more functions not displayed before the user's attention is directed to the first selectable user interface object (e.g., a link to a web page, a dictionary definition, a user interface object, a widget, or a thumbnail preview). In some embodiments, a larger size (or expanded) version of the first selectable user interface object is displayed at the same location of the first selectable user interface object before the user's attention is directed to the first selectable user interface object. In some embodiments, the information associated with the first selectable user interface object is displayed as covering or overlapping the first selectable user interface object. In some embodiments, when the user's attention is not directed to the first selectable user interface object or is not in a region of the first selectable user interface object, the computer system does not increase the size of the first selectable user interface object and displays the information associated with the first selectable user interface object. In some embodiments, before the user's attention is directed to the first selectable user interface object, the first selectable user interface object includes a first portion of information. In some embodiments, the increased size of the first selectable user interface object accommodates a second portion of information that is larger than the first portion of information. Increasing the size of the selectable user interface object to display the information associated with the selectable user interface object provides improved feedback to the user without cluttering the user interface (e.g., by not always displaying the supplementary information associated with the selectable user interface object), thereby enhancing the operability of the computer system and reducing the power consumption of the computer system.

[0247] In some embodiments, before displaying the first fixation target, the first selectable user interface object is included in a first region of the user interface rather than a second region of the user interface (810a), such as Figure 7B the location of object 706b in. In some embodiments, displaying the first fixation target includes displaying a first fixation target associated with the first selectable user interface object (810b) in a second region of the user interface different from the first region, such as Figure 7CThe position of the fixation target 706b' in []. In some embodiments, even if the first fixation target is associated with the first optional user interface object, the first fixation target and its associated first optional user interface object are visually or spatially separated in the user interface (e.g., displayed in different regions of the user interface). For example, the computer system optionally displays the first optional user interface object in the leftmost middle region of the user interface container (e.g., a box) and displays the first fixation target in the rightmost middle region of the same user interface container. As another example, the computer system optionally displays the first optional user interface object in the central region of the user interface container and displays the first fixation target in the rightmost bottom region of the same user interface container. In some embodiments, the computer system requires the user's attention to be directed to the first fixation target or to persist in the second region of the first fixation target for a period of time greater than a second time threshold in order to initiate the first operation associated with the first optional user interface object, as described with reference to step 802. In some embodiments, the computer system displays the first fixation target in another user interface container different from the user interface container associated with the first optional user interface object, as will be described with reference to step 840. In some embodiments, the respective regions in which the computer system displays the first fixation target and the first optional user interface object vary depending on the type and size of the user interface container (e.g., list, table, box, or column) or in the absence of a user interface container associated with the first optional user interface object, as will be detailed with reference to steps 832, 840, and 842. In some embodiments, the computer system displays the first optional user interface object in the first region at the first depth and displays the first fixation target in the second region at a second depth different from the first depth. For example, the second depth is closer to the computer system (e.g., the user's viewing point) than the first depth. In some embodiments, although the computer system displays the first optional user interface object in the first region and displays the first fixation target in a second region different from the first region, the second region selected by the computer system for displaying the first fixation target is within a threshold distance (e.g., 0.05 cm, 0.1 cm, 0.2 cm, 0.3 cm, 0.4 cm, 0.5 cm, 0.8 cm, 1 cm, 1.5 cm, 2 cm, 2.5 cm, 3 cm, 5 cm, 7 cm, or 10 cm) of the first optional user interface object, such that the displayed first fixation target and first optional user interface object appear as a visually distinct group (e.g., the first fixation target is associated with the first optional user interface object and not with other user interface objects of the user interface).Displaying a fixation target in an area different from the first selectable user interface object requires the user's attention to be directed to the fixation target or the area of the fixation target in order to perform an operation associated with the first selectable user interface. This operation provides additional confirmation that the user truly intends to interact with the first fixation target and / or the selectable user interface object, thereby reducing errors in the interaction between the user and the computer system (e.g., avoiding accidentally activating or deactivating a selectable user interface object due to unintentional fixation) and reducing the input required to correct such errors.

[0248] In some embodiments, before detecting that the user's attention is directed to the first selectable user interface object, the first selectable user interface object is displayed at a first distance (812a) from the user's line of sight (e.g., a first visual separation from the backplane behind the first selectable object), such as Fig. 7A object 704a in. In some embodiments, when the user's attention is directed to the first selectable user interface object, the computer system displays the first selectable user interface object at a second distance (812b) different from the first distance from the user's line of sight (e.g., a second, greater visual separation from the backplane behind the first selectable object), such as if object 704 is Figure 7BThe distance of the user's viewing point in [object] changes. In some embodiments, when the user's attention is not directed to or within the area of the first selectable user interface object, the computer system displays the first selectable user interface object at a first distance from the user's viewing point (e.g., the first selectable user interface object abuts the backplane of the user interface, or the first selectable user interface object is presented at a first depth and the user interface is presented at the same first depth). In some embodiments, when the user's attention is directed to or within the area of the first selectable user interface object, the computer system displays the first selectable user interface object at a second distance closer to the user's viewing point than the first distance (e.g., the first selectable user interface object is presented at a second depth closer to the user's viewing point than the first depth of the user interface). In some embodiments, when the user's attention is directed to or within the area of the first selectable user interface object, the computer system displays the first selectable user interface object at the second distance with an analog shadow. In some embodiments, when the user's attention is not directed to or within the area of the first selectable user interface object, the computer system displays the first selectable user interface object at the first distance without an analog shadow. Displaying the first selectable user interface object closer to the user's viewing point when the user's attention is directed to the first selectable user interface object allows the computer system to convey to the user that the user's attention is directed to the first selectable user interface object, thereby reducing errors in the interaction between the user and the computer system (e.g., avoiding accidentally activating or deactivating a selectable user interface object due to unintentional gazing) and reducing the input required to correct such errors.

[0249] In some embodiments, when displaying the first selectable user interface object at the second distance from the user's viewing point, the computer system displays (814) in the user interface in association with the first selectable user interface object a corresponding user interface object including information about the first selectable user interface object, such as Figure 7C"Tooltips" for the object 706a. In some embodiments, the information includes text and / or content associated with the first optional user interface object, such as tooltips, definitions, translations, or other excerpts of information about the first optional user interface object (e.g., information about what actions will occur if the first optional user interface object or its gaze target is selected). In some embodiments, the computer displays within the corresponding user interface object the first gaze target associated with the first optional user interface object as described in step 802. For example, the corresponding user interface object includes information about the first optional user interface object in the left region of the corresponding user interface object and the first gaze target in the right region. In some embodiments, the corresponding user interface object is displayed above, below, to the left, or to the right of the first optional user interface object. In some embodiments, the computer system displays the information before displaying the first gaze target. In some embodiments, the computer system displays the information and the first gaze target simultaneously. In some embodiments, when the user's attention is not directed to the first optional user interface object or is not within the area of the first optional user interface object, the computer system does not display the corresponding user interface object that includes information about the first optional user interface object. Displaying the corresponding user interface object that includes information about the optional user interface object provides improved feedback to the user without cluttering the user interface (e.g., by not always displaying information about the optional user interface object), thereby enhancing the operability of the computer system and reducing the power consumption of the computer system.

[0250] In some embodiments, when the user's attention is directed to the first optional user interface object, the computer system provides (816) a first output indicating that the user's attention is directed to the first optional user interface object, such as a reference Figure 7Bas described for timer 712 in. In some embodiments, the first output includes an audio output, such as one or more tones (e.g., the sounds "ding" or "beep") or chords in a melody. In some embodiments, the tone and / or chord includes one or more acoustic characteristics, such as pitch, volume, timbre, harmony, rhythm, attack, sustain, decay, and / or tempo. In some embodiments, the first output includes a haptic output (e.g., vibration and / or tactility). In some embodiments, the first output includes changing the visual appearance of the first optional user interface object, such as as described in reference step 812. For example, the computer system displays the first optional user interface object with a glowing effect based on the characteristics of one or more simulated light sources located in the three-dimensional environment (e.g., brightness, color, position, size, and / or directionality) and / or based on such characteristics of one or more simulated light sources that are not actually located in the three-dimensional environment, but based on this, the computer system displays the glowing effect as if they were located in the three-dimensional environment. In some embodiments, the first output includes any combination of the audio output, haptic output, and / or changed visual appearance described herein. The audio output, haptic output, and / or visual output are provided based on determining that the user's attention is directed to the first optional user interface object, enhancing the interaction between the user and the computer system by providing the user with improved feedback and reducing the likelihood of errors in the interaction between the user and the computer system.

[0251] In some embodiments, when the user's attention is directed to the first fixation target, the computer system provides a second output (818) indicating that the user's attention is directed to the first fixation target, such as as described in reference Figure 7C for timer 714 in. In some embodiments, the first output includes any combination of the audio output, haptic output, and / or changed visual appearance as described in steps 816 and method 900. In some embodiments, the second output is different from the first output described in step 816 to distinguish the second output associated with the first fixation target from the first output associated with the first optional user interface object. For example, the second output optionally includes a first acoustic characteristic that is higher than the corresponding acoustic characteristic of the first output (e.g., a higher pitch and / or a stronger rhythm). In some embodiments, the computer system provides a third audio output that is similar to or the same as the second output, where the audio output indicates that the user's attention is directed to the first optional user interface object. The audio output, haptic output, and / or visual output are provided based on determining that the user's attention is directed to the first fixation target, enhancing the interaction between the user and the computer system by providing the user with improved feedback and reducing the likelihood of errors in the interaction between the user and the computer system.

[0252] In some embodiments, when the user's attention is directed to the first fixation target, such as Figure 7FThe medium attention 716 is directed to the target 704a', and the computer system detects (820) a first input directed to the first fixation target via one or more input devices (e.g., as described in step 802), where the first input includes an input from a first part of the user's body of the computer system, such as Figure 7F an input from the hand 710. In some embodiments, the first part of the user corresponds to the user's first hand, arm, palm, and / or one or more fingers or the head of the first hand (e.g., the left hand or the right hand). In some embodiments, the first input has one or more of the characteristics of the input described in reference to step 802.

[0253] In some embodiments, in response to detecting the first input directed to the first fixation target, the computer system initiates (820) a first operation associated with the first optional user interface object (e.g., without waiting for the user's attention (directed to the first fixation target) to meet one or more criteria), such as Figure 7G and Figure 7G1 updating the content on the right side of the object 704. In some embodiments, when the computer system detects the first part of the user (e.g., the gesture / input described in step 802) directed to the first fixation target, the computer system initiates the first operation. In some embodiments, the first input includes the air pinch gesture or the selection input (e.g., tap, touch, or click) described in step 802. The first input from the user optionally includes other types of input, such as the touchpad input (e.g., finger touching the touchpad) or the input device input (e.g., selection via a handheld input device such as a stylus or a remote control) described in step 802. In some embodiments, when the computer system detects the first input while the user's attention is directed to the first optional user interface object and does not meet one or more criteria (e.g., the criteria for the user's attention to be directed to the first fixation target to meet one or more criteria described in step 802), the computer system initiates the first operation. Performing an operation in response to the user input directed to the first fixation target provides an effective interaction with the virtual object.

[0254] In some embodiments, determining whether the user's attention is directed to the first optional user interface object includes determining whether the first part of the user's body is in a first state (822) (e.g., the ready state configuration of the user or the first part of the user as described herein), such as Figure 7B the state of the hand 710, and based on determining that the first part of the user's body is in the first state, the computer system abandons displaying (822) the first fixation target associated with the first optional user interface object in the user interface, such as not displaying Figure 7Cthe fixation target 704a' in []. In some embodiments, if the computer system detects that a part of the user's body is in a second state different from the first state and / or in a posture indicating a state different from the first state, the computer system does not display the first fixation target associated with the first optional user interface object in response to detecting that the user's attention is directed to the first optional user interface object, as described in step 802. Abandoning the display of the first fixation target when the input from the first part of the user's body is more likely than the attention-based input when the first part of the user's body is in a ready state reduces the clutter of the user interface.

[0255] In some embodiments, the user interface includes a second optional user interface object (824a) selectable to perform a second operation, such as object 704c. The second optional user interface object optionally has one or more of the characteristics of the first optional user interface object described in step 802. In some embodiments, the second optional user interface object is the same type of object as the first optional user interface object, but performs a second operation different from the first operation associated with the first optional user interface object when selected.

[0256] In some embodiments, the computer system detects (824b) that the attention of the user of the computer system is directed to the second optional user interface object, such as Figure 7B the attention 720 in []. In some embodiments, in response to detecting that the user's attention is directed to the second optional user interface object, the computer system displays (824c) a second fixation target associated with the second optional user interface object in the user interface, where the first fixation target and the second fixation target have the same visual appearance (e.g., color, size, shape, animation, visual effect, and / or movement), such as Figure 7C the fixation target 704a' in [] and the fixation target of object 704c. In some embodiments, the first fixation target and the second fixation target associated with the respective optional user interface objects are the same, even though the respective optional user interface objects are associated with different applications and perform respective different operations from each other when selected. In some embodiments, the first fixation target and the second fixation target have the same visual appearance, including the same visual characteristics described in steps 806, 810, 818, 840, 842 - 852, and method 900. Displaying the fixation targets with a similar visual appearance enhances the interaction between the user and the computer system by providing improved feedback to the user (e.g., by consistently displaying the fixation targets with the same visual appearance) and reducing the likelihood of errors in the interaction between the user and the computer system.

[0257] In some embodiments, displaying the first optional user interface object when the user's attention is directed to the first optional user interface object includes in a first visual appearance such as Figure 7M The appearance of the object 704d in [[ ]] shows a first selectable user interface object (826a). For example, when the computer system displays the first selectable user interface object in a first visual appearance, the first selectable user interface object is expanded to a larger size (e.g., consistent with the increased size of the first selectable user interface object described in step 808, but not limited thereto). In some embodiments, when the computer system displays the first selectable user interface object in a first visual appearance, the first selectable user interface object is displayed in a first fixation target and a first fixation target in a second region, consistent with the first region and the second region described in step 810, but not limited thereto.

[0258] In some embodiments, the user interface includes a second selectable user interface object (826b) selectable to perform a second operation, such as Figure 7M the object 708a in [[ ]]. In some embodiments, the second selectable user interface object selectable to perform a second operation is consistent with the second selectable user interface object described in step 824, but not limited thereto.

[0259] In some embodiments, when the user's attention is directed to the first selectable user interface object, the computer system detects that the attention of the user of the computer system has changed to be directed to a second selectable user interface object (826c), such as from Figures 7M to 7N the attention 716 in [[ ]]. For example, the user's attention becomes away from the first selectable user interface object for a period of time greater than a first time threshold (e.g., as described in step 802), such that the user's attention is directed to a region of the second selectable user interface (e.g., a region other than the region occupied by the first selectable user interface object) or the user's attention is away from the user interface.

[0260] In some embodiments, in response to detecting that the user's attention is directed to the second selectable user interface object, the computer system displays (826d) the first selectable user interface object in a second visual appearance different from the first visual appearance, such as the object 704d from FIG. 7M to FIG. 7OThe changed visual appearance. For example, the second visual appearance optionally includes a second size that is smaller (more compact) than the first size of the first visual appearance. In some embodiments, displaying the first optional user interface object in the second visual appearance includes collapsing the user interface container in response to detecting that the user's attention is directed to the second optional user interface object (e.g., as described in steps 810, 832, 840, and 842). In some embodiments, displaying the first optional user interface object in the second visual appearance includes restoring the visual appearance of the first optional user interface object to its appearance before the user's attention was directed to the first optional user interface object (e.g., a compact size that does not accommodate information associated with the first optional user interface object and / or the first fixation target, as described with reference to steps 808 and 810). Selectively increasing or decreasing (e.g., expanding or collapsing) the size of the optional user interface object in response to whether the attention is directed to the optional user interface object provides improved feedback to the user without cluttering the user interface (e.g., by not always displaying the expanded user interface object), thereby enhancing the operability of the computer system and reducing the power consumption of the computer system.

[0261] In some embodiments, the user interface includes a second optional user interface object (828a) selectable to perform a second operation, such as Figure 7M object 708a in. In some embodiments, when displaying a first fixation target associated with the first optional user interface object, such as Figure 7M target 704d' in, the computer system detects (828b) that the attention of the user of the computer system is directed to a second optional user interface object, such as Fig.7O in. For example, the user's attention becomes diverted from the first optional user interface object for a period of time greater than a first time threshold (e.g., as described in step 802), such that the user's attention is directed to a different region of the user interface (e.g., a region other than the first region occupied by the first optional user interface object) or the user's attention is diverted from the user interface.

[0262] In some embodiments, in response to detecting that the user's attention is directed to the second optional user interface object, the computer system stops (828c) displaying the first fixation target associated with the first optional user interface object, such as in Fig.7OStop displaying 704d'. For example, the first fixation target is optionally configured by the computer system to be transient, because the computer system displays the first fixation target when the user's attention is directed to the first fixation target and / or the first selectable object, and stops displaying the first fixation target or reduces the visual salience of the first fixation target when the user's attention is directed to a second selectable user interface object (e.g., away from the first selectable user interface object). Details regarding changing the visual salience of the first fixation target are described with reference to steps 850 and 852. Displaying or not displaying the first fixation target associated with the first selectable user interface object in response to whether attention is directed to the first selectable user interface object provides improved feedback to the user without cluttering the user interface (e.g., by not always displaying the first fixation target), thereby enhancing the operability of the computer system and reducing the power consumption of the computer system.

[0263] In some embodiments, the user interface includes a second selectable user interface object (830a) for performing a second operation, such as Fig.7I the role 001 object in. In some embodiments, the computer system detects (830b) that the attention of a user of the computer system is directed to the second selectable user interface object (e.g., as described in step 802), such as Fig.7I between Figure 7J the attention 716 to the role 001 object. In some embodiments, in response to detecting that the user's attention is directed to the second selectable user interface object, the computer system displays (830c) a second fixation target associated with the second selectable user interface object in the user interface, such as Figure 7J the fixation target 704e' in. In some embodiments, the computer system displays the second fixation target associated with the second selectable user interface object based on determining that the user's attention is directed to the second selectable user interface object for a first time period greater than a first time threshold as described in reference step 802. For example, the second fixation target is optionally the same as the first fixation target described in step 802 but is not limited thereto. In some embodiments, as described in reference step 824, the first fixation target and the second fixation target have the same visual appearance. Displaying a fixation target for a second selectable user interface object different from the first selectable user interface object provides confirmation of the user's intention to interact with the second selectable user interface object, thereby reducing errors in the interaction between the user and the computer system (e.g., avoiding accidentally activating or deactivating a selectable user interface object due to unintentional fixation) and reducing the input required to correct such errors.

[0264] In some embodiments, the first selectable user interface object is displayed as being composed of a simulated material having a certain thickness (832) (e.g., non-zero thickness), such as Figure 7JThe object 704e in. In some embodiments, the computer system displays a first optional user interface object having an analog three-dimensional depth or thickness. For example, the first optional user interface object is optionally displayed on, against, and / or in front of the user interface, projected into a forward projection. In some embodiments, the thickness of the analog material optionally corresponds to the forward projection of the first optional user interface object. In some embodiments, the first optional user interface object is displayed as being composed of an analog glass or other material.

[0265] In some embodiments, displaying the first fixation target includes displaying a first fixation target embedded in the surface of an analog material of the first optional user interface object (832), such as in Figure 7J where the target 704e is embedded in the surface of the object 704e. For example, the first fixation target is optionally manifested as being etched into the first optional user interface object (optionally, the front surface) (e.g., into the area of the forward projection of the first optional user interface object). In some embodiments, the analog material is transparent such that the first fixation target is viewed at different viewing angles. Displaying the first optional user interface object as being composed of an analog material having a thickness conveys to the user the relative placement and / or orientation of the first optional user interface object, thereby reducing errors in the interaction between the user and the computer system (e.g., avoiding accidentally activating or deactivating an optional user interface object due to unintentional fixation) and reducing the input required to correct such errors.

[0266] In some embodiments, when displaying a user interface that includes a scrollable region (e.g., a region that can be manipulated / scrolled), where the first optional user interface object is included in the scrollable region (such as Figure 7J the right side region of the object 704 in), the computer system detects (834a) the user's attention directed to the first edge region of the scrollable region via one or more input devices, such as whether the fixation 716 is directed to Figure 7K the top edge of the right side region of the object 704 in. In some embodiments, in response to detecting the user's attention directed to the first edge region, the computer system scrolls (834b) the scrollable region according to the user's attention being directed to the first edge region, including scrolling the first optional user interface object from a first position to a second position in the user interface, such as Figure 7LThe scrolling shown. For example, a computer system detects a user's attention directed to a first edge region (e.g., top, bottom, left, or right) of a scrollable region to scroll a first selectable user interface object within a user interface and scrolls the first selectable user interface object accordingly. For example, if the computer system detects that the user's attention is directed to the top edge of the scrollable region, the first selectable user interface object, which includes any other displayed user interface objects and / or content, moves upward to optionally display from the bottom of the user interface user interface objects and / or content that were not previously displayed before the scroll. As another example, if the computer system detects that the user's attention is directed to the bottom edge of the scrollable region, the first selectable user interface object, which includes any other displayed user interface objects and / or content, moves downward to optionally display from the top of the user interface user interface objects and / or content that were not previously displayed before the scroll. In some embodiments, when the user's attention is not directed to any edge including the first edge of the scrollable region, the computer system does not scroll the first selectable user interface object within the user interface. Scrolling the first selectable user interface object in response to detecting the user's attention directed to the edge region of the scrollable region provides quick access to the user interface object without requiring the user to provide further input to navigate within the user interface, thereby reducing the amount of input and providing a more efficient interaction between the user and the computer system.

[0267] In some embodiments, when it is detected that the user's attention is directed to the first edge region of the scrollable region, a first fixation target (836a) associated with the first selectable user interface object is displayed, such as target 704e' is displayed in Figure 7K In some embodiments, when the scrollable region is scrolled, the computer system stops displaying the first fixation target (836b) associated with the first selectable user interface object, such as target 704e' is stopped being displayed in Figure 7L . Optionally, regardless of whether the first fixation target reaches the boundary of the user interface such that continuing to scroll would cause the first fixation target to scroll out of the user interface, the computer system stops displaying the first fixation target. In some embodiments, when scrolling, the computer system reduces the visual saliency of the first fixation target before stopping displaying it. In some embodiments, if the computer system detects stopping the scrolling of the scrollable region, the computer system displays the first fixation target in a manner consistent with but not limited to the description in steps 850 and 852. Details regarding changing the visual saliency of the first fixation target are described with reference to steps 850 and 852. Stopping the display of the first fixation target associated with the first selectable user interface object in response to scrolling provides improved feedback to the user without cluttering the user interface (e.g., by not always displaying the first fixation target), thereby enhancing the operability of the computer system and reducing the power consumption of the computer system.

[0268] In some embodiments, when displaying a user interface that includes a scrollable region (such as the right region of object 704 in Figure 7K , the computer system detects (838a) a first input directed to the scrollable region via one or more input devices, where the first input includes a corresponding gesture performed by a corresponding part of the user's body of the computer system that corresponds to a request to scroll the scrollable region, such as the input from hand 710 in Figure 7K . In some embodiments, in response to detecting the first input, the computer system scrolls (838b) the scrollable region according to the first input, including scrolling a first selectable user interface object from a first position to a second position different from the first position in the user interface, such as the scrolling shown in Figure 7L . In some embodiments, the first input from the user includes an air pinch gesture performed by the user's hand while the user's attention is directed to a first edge region of the scrollable region, in which the user's index finger and the user's thumb come together and touch, and then the hand in a pinching shape moves in one direction and / or by a certain amount. The computer system optionally scrolls the first selectable user interface object within the user interface by an amount and / or in a direction corresponding to the movement of the user's hand (e.g., scrolling the first selectable user interface object upward if the hand moves upward and scrolling the first selectable user interface object downward if the hand moves downward). The first input from the user optionally includes other types of input, such as touchpad input (e.g., a finger touches the touchpad and moves in one direction and / or by a certain amount) or input device input (e.g., the movement of a handheld input device that detects the direction and / or amount of movement of the input device when held in the user's hand). Scrolling the first selectable user interface object in response to detecting the first input when the user's attention is directed to the edge region of the scrollable region provides quick access to the user interface object without requiring the user to provide further input to navigate within the user interface, thereby reducing the number of inputs and providing a more efficient interaction between the user and the computer system.

[0269] In some embodiments, displaying the first selectable user interface object includes displaying the first selectable user interface object as a first element (840a) in the user interface, such as object 706b in Figure 7C . In some embodiments, displaying the first fixation target includes displaying the first fixation target as a second element (840b) outside the first element in the user interface, such as Figure 7Cas shown by target 706b' in []. For example, the user interface includes a first selectable user interface object and a first fixation target, also referred to as a first element and a second element, respectively, separated by a visible or invisible boundary. In some embodiments, since the first selectable user interface object is a first type of object as described in method 1200 (e.g., a photo or a video), the first selectable user interface object fills the available area up to the boundary. Thus, in some embodiments, the first fixation target is displayed outside the first selectable user interface object in an area adjacent to the first selectable user interface object, consistent with but not limited to the first area and the second area described in step 810. Displaying the fixation target outside the first selectable user interface object provides a more efficient use of the display space, thereby reducing errors in the interaction between the user and the computer system (e.g., avoiding accidentally activating or deactivating a selectable user interface object due to unintentional fixation) and reducing the input required to correct such errors.

[0270] In some embodiments, displaying the first selectable user interface object includes displaying the first selectable user interface object as a first element (842a) in the user interface, such as Figure 7C object 704a in []. In some embodiments, displaying the first fixation target includes displaying the first fixation target within the first element (842b) in the user interface, such as Figure 7C as shown by target 704a' in []. For example, the first fixation target is displayed within the first selectable user interface object (also referred to as the first element). In some embodiments, when the first fixation target is displayed within the first element, the user interface includes a user interface container (e.g., a frame) and is configured to display the first fixation target and the first selectable user interface object within a visible or invisible boundary of the user interface container. In some embodiments, the first fixation target is displayed in a first area of the user interface container, and the first selectable user interface object is displayed in a second area of the user interface container, consistent with but not limited to the first area and the second area described in step 810. In some embodiments, the first fixation target is displayed on, above, and / or covering the surface of the first selectable user interface object and / or the first selectable user interface object. Displaying the fixation target within the first selectable user interface object provides a more efficient use of the display space, thereby reducing errors in the interaction between the user and the computer system (e.g., avoiding accidentally activating or deactivating a selectable user interface object due to unintentional fixation) and reducing the input required to correct such errors.

[0271] In some embodiments, when the computer system detects the user's attention directed to the first fixation target, the computer system outputs a first feedback (844) based on the user's attention directed to the first fixation target, such as with reference to Figure 7Cas described with respect to timer 714 in. For example, the first feedback includes any combination of audio outputs, where the audio output includes audio characteristics as described in steps 816, 818, and method 900. The audio output is provided based on determining that the user's attention is directed to the first fixation target, enhancing the interaction between the user and the computer system by providing improved feedback to the user and reducing the likelihood of errors in the interaction between the user and the computer system.

[0272] In some embodiments, outputting the first feedback includes outputting feedback (846) having corresponding characteristics that change based on the progress of the user's attention towards meeting one or more criteria, such as the audio characteristics changing as Figure 7C timer 714 in elapses. In some embodiments, when outputting the first feedback having corresponding characteristics, the computer system detects the duration during which the user's attention directed to the first fixation target changes from a first attention duration to a second attention duration, and in response to detecting that the user's attention directed to the first fixation target changes from the first attention duration to the second attention duration, the computer system outputs a first feedback indicator having corresponding characteristics that change corresponding to the duration of the user's attention directed to meeting one or more criteria as described in step 802. For example, the audio characteristics optionally correspond to the progress made by the user's attention towards meeting one or more criteria (e.g., the volume increases as the progress increases, the pitch becomes higher as the progress increases, and / or the melody / tone increases as the progress increases, or alternatively, the volume, pitch, and / or melody / tone decrease when the computer system detects that the user's attention moves away from the first fixation target and thus stops the progress of the user's attention towards meeting one or more criteria). Providing a changing audio output based on the progress of the user's attention towards meeting one or more criteria provides improved feedback to the user and allows the computer system to convey to the user their progress towards meeting one or more criteria in order to perform an operation associated with the first optional user interface object.

[0273] In some embodiments, when outputting the first feedback, based on determining that the user's attention directed to the first fixation target meets one or more criteria, the computer system outputs (848) a second feedback indicating that the one or more criteria are met, such as in Fig.7IAudio output when timer 714 reaches threshold 714d or when timer 740 reaches threshold 734d. For example, the second feedback includes an audible output, such as one or more tones (e.g., the sounds "ding" or "beep") or chords in a melody. In some embodiments, the second feedback is consistent with the first audio feedback described in method 900. In some embodiments, the second feedback is different from the first audio feedback described in step 846 such that the second feedback includes one or more different sound characteristics (e.g., different pitch or reverberant sound). The second feedback is provided based on determining that the user's attention directed to the first fixation target meets one or more criteria for initiating a first operation associated with the first optional user interface object, enhancing the user's interaction with the computer system by providing the user with improved feedback and reducing the likelihood of errors in the interaction between the user and the computer system.

[0274] In some embodiments, when the first fixation target is displayed, the computer system detects (850a) that the user's attention is diverted from the first optional user interface object, such as when attention 716 moves away from object 704d, from Figure 7M moving to Fig.7O . In some embodiments, in response to detecting that the user's attention is diverted from the first optional user interface object, and based on determining that the user's attention meets one or more second criteria, the computer system reduces (850b) the visual saliency of the first fixation target relative to the three-dimensional environment, such as fixation target 704d' from FIG. 7M to FIG. 7O as shown. In some embodiments, the one or more second criteria include criteria that are met when the user's attention is not directed to the first fixation target. In some embodiments, reducing the visual saliency of the first fixation target relative to the three-dimensional environment includes gradually changing the degree of visibility (e.g., decreasing opacity, increasing transparency, decreasing brightness, decreasing color saturation, and / or increasing blur) until the appearance of the first fixation target ceases to be displayed in the user interface.

[0275] In some embodiments, based on determining that the user's attention does not meet one or more second criteria, the computer system abandons (850b) reducing the visual salience of the first fixation target relative to the three-dimensional environment. In some embodiments, the foregoing reducing the visual salience of the first fixation target relative to the three-dimensional environment includes maintaining the visual appearance of the first fixation target at the appearance of the first fixation target when the user's attention is diverted from the first alternative user interface object. Reducing the visual salience of the first fixation target associated with the first alternative user interface object in response to detecting that the attention is diverted from the first alternative user interface object provides confirmation that the user no longer intends to interact with the first alternative user interface object, thereby reducing errors in the interaction between the user and the computer system (e.g., avoiding accidentally activating or deactivating an alternative user interface object due to unintentional fixation) and reducing the input required to correct such errors.

[0276] In some embodiments, when the first fixation target is displayed with reduced visual salience relative to the three-dimensional environment, the computer system detects (852a) the user's attention directed to the first alternative user interface object, such as detecting that the attention 716 moves back to the target 704a', from Fig. 7E moving to Figure 7F . In some embodiments, in response to detecting the user's attention directed to the first alternative user interface object, the computer system increases (852b) the visual salience of the first fixation target relative to the three-dimensional environment, such as increasing the salience of the target 704a' in Figure 7F . In some embodiments, when the computer system detects that the user's attention has become diverted from the first alternative user interface and has returned to the first alternative user interface object for less than a first threshold time period, the computer stops reducing the visual salience of the first fixation target relative to the three-dimensional environment, as described in step 852, and increases the visual salience of the first fixation target relative to the three-dimensional environment by reversing the degree of changed visibility (e.g., increased opacity, decreased transparency, increased brightness, increased color saturation, and / or decreased blur) until the appearance of the first fixation before the user's attention was diverted from the first alternative user interface object is the same as the appearance of the first fixation target. In such a case, the computer system optionally displays the first fixation target according to some embodiments described with reference to steps 808, 816, and / or method 900. Stopping the transition of the visual appearance of the user interface object in response to detecting that the attention has returned to the first alternative user interface object after being diverted from the first alternative user interface object provides confirmation that the user does indeed intend to interact with the first alternative user interface object, thereby reducing errors in the interaction between the user and the computer system (e.g., avoiding accidentally activating or deactivating an alternative user interface object due to unintentional fixation) and reducing the input required to correct such errors.

[0277] In some embodiments, one or more criteria include criteria that are satisfied when the user's attention is directed to a first fixation target for a duration exceeding a corresponding threshold time period (854a), such as Figure 7C the threshold 712b in

[0278] In some embodiments, the corresponding threshold time period is the second time threshold described in step 802. In some embodiments, the corresponding time threshold varies based on the user's previous interactions with one or more optional user interface objects. Figure 7F In some embodiments, based on determining that one or more previous user interactions with one or more optional user interface objects or one or more fixation targets satisfy one or more second criteria, the corresponding threshold time period is a first threshold time period (854b), such as

[0279] the threshold 732b' in Figure 7G and Figure 7G1 In some embodiments, based on determining that one or more previous user interactions with one or more optional user interface objects or one or more fixation targets do not satisfy one or more second criteria, the corresponding threshold time period is a second threshold time period (854c) different from the first threshold time period, such as

[0280] It should be understood that the specific order in which the operations in method 800 are described is merely exemplary and is not intended to indicate that the described order is the only order in which these operations can be performed. Those of ordinary skill in the art will envision various ways to reorder the operations described herein.

[0281] FIG. 9A to FIG. 9J is a flowchart showing an exemplary method 900 of displaying a gaze virtual object according to some embodiments, where the gaze virtual object is selectable based on the attention directed to it to perform an operation associated with an optional virtual object. In some embodiments, method 900 is executed at a computer system (e.g., computer system 101 in FIG. 1, such as a tablet computer, a smart phone, a wearable computer, or a head-mounted device), which includes a display generation component (e.g., the display generation component 120 in FIGS. 1, Figure 3 and Figure 4 such as a head-up display, a display, a touch screen, a projector, etc.) and one or more cameras (e.g., a camera pointing downward at the user's hand (e.g., a color sensor, an infrared sensor, and other depth-sensing cameras) or a camera pointing forward from the user's head). In some embodiments, method 900 is managed by instructions stored in a non-transitory computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., Figure 1A the control unit 110 in FIG. 1). Some operations in method 900 are optionally combined, and / or the order of some operations is optionally changed.

[0282] In some embodiments, method 900 is executed at a computer system that communicates with a display generation component and one or more input devices. For example, the computer system includes the devices described in reference to method 800. In some embodiments, the display generation component includes the display described in reference to method 800. In some embodiments, one or more input devices have one or more of the characteristics of the one or more input devices described in reference to method 800.

[0283] In some embodiments, the computer system displays (902a) via the display generation component a user interface including a first optional user interface object that is selectable to perform a first operation, such as Figure 7CThe target 704a' in it. In some embodiments, the user interface is displayed in the three-dimensional environment described in reference method 800. In some embodiments, the user interface is the user interface described in reference method 800. In some embodiments, the first optional user interface object corresponds to the first optional user interface object described in method 800. In some embodiments, the first optional user interface object corresponds to the first fixation target or other fixation targets of the parent optional object, such as those described in method 800. In some embodiments, the first operation is associated with the first optional user interface object (e.g., the first optional user interface object is not a fixation target, and if the computer system detects a non-attention-based selection of the first optional user interface object, such as a selection via an air pinching gesture, where the user's attention is directed to the first optional user interface object while the user's hand performs an air pinching gesture that includes bringing the fingertips of the thumb and index finger of the hand together and touching, then the first operation will be executed by the computer system). In some embodiments, the first operation is associated with a user interface object other than the first optional user interface object. For example, if the first optional user interface object corresponds to the fixation target described in method 800, then the first optional user interface object is optionally configured to initiate a process of executing the first operation associated with the parent object of the first optional user interface object (e.g., the parent object of the first optional user interface object is not a fixation target, and if the computer system detects a non-attention-based selection of the parent object, such as the above-mentioned selection via a pinching air gesture, then the first operation will be executed by the computer system).

[0284] In some embodiments, when the user interface is displayed, the computer system detects (902b) via one or more input devices the attention of the user of the computer system being directed to the first optional user interface object, such as Figure 7C the attention 716 directed to the target 704a' in it. In some embodiments, the gaze tracking device optionally captures one or more images of the user's eyes and detects the pupils and glints in the one or more captured images to track the user's gaze, as detailed in reference Figure 6 as described. In some embodiments, the user's attention becomes the location (or part) of the user interface that includes the first optional user interface object. In some embodiments, the user interface includes a visual representation (indication) of the user's gaze as described in reference methods 1200 and 2000. ...

Claims

1. A method, the method comprising: At a computer system in communication with a display generation component and one or more input devices: Display, via the display generation component, a user interface including a first selectable user interface object selectable to perform a first operation; When the user interface is displayed, detect, via the one or more input devices, that the attention of a user of the computer system is directed to the first selectable user interface object; In response to detecting that the user's attention is directed to the first selectable user interface object, display, in the user interface, a first fixation target associated with the first selectable user interface object; When the first fixation target is displayed, detect that the user's attention is directed to the first fixation target; And In response to detecting that the user's attention is directed to the first fixation target: Initiate, based on determining that the user's attention meets one or more criteria, the first operation associated with the first selectable user interface object; And Abandon initiating, based on determining that the user's attention does not meet the one or more criteria, the first operation associated with the first selectable user interface object.

2. The method according to claim 1, wherein determining that the user's attention is directed to the first selectable user interface object includes determining that the user's gaze has been directed to the first selectable user interface object for a period exceeding a first threshold time period.

3. The method according to any one of claims 1 to 2, wherein: Displaying the first selectable user interface object includes displaying the first selectable user interface object with a first characteristic having a first value; and Displaying the first fixation target includes displaying the first fixation target with the first characteristic having a second value different from the first value.

4. The method according to any one of claims 1 to 3, the method further comprising: In response to detecting the user's attention directed to the first selectable user interface object, increase the size of the first selectable user interface object and display information associated with the first selectable user interface object in the first selectable user interface object, wherein the information is not displayed before the user's attention was directed to the first selectable user interface object.

5. The method according to any one of claims 1 to 4, wherein: Before the first fixation target is displayed, the first selectable user interface object is included in a first area of the user interface rather than a second area of the user interface, and Displaying the first fixation target includes displaying the first fixation target associated with the first selectable user interface object in a second area of the user interface different from the first area.

6. The method according to any one of claims 1 to 5, wherein before detecting that the user's attention is directed to the first selectable user interface object, the first selectable user interface object is displayed at a first distance from the user's viewing point, the method further comprising: When the user's attention points to the first optional user interface object, display the first optional user interface object at a second distance from the user's viewpoint that is different from the first distance.

7. The method according to claim 6, the method further comprising: When the first optional user interface object is displayed at the second distance from the user's viewpoint, display, in the user interface, a corresponding user interface object associated with the first optional user interface object and including information about the first optional user interface object.

8. The method according to any one of claims 1 to 7, the method further comprising: When the user's attention points to the first optional user interface object, provide a first output indicating that the user's attention points to the first optional user interface object.

9. The method according to claim 8, the method further comprising: When the user's attention points to the first fixation target, provide a second output indicating that the user's attention points to the first fixation target.

10. The method according to any one of claims 1 to 9, the method further comprising: When the user's attention points to the first fixation target, detect, via the one or more input devices, a first input pointing to the first fixation target, wherein the first input includes an input from a first part of the user's body of the computer system; and in response to detecting the first input pointing to the first fixation target, initiate the first operation associated with the first optional user interface object.

11. The method according to claim 10, wherein determining whether the user's attention points to the first optional user interface object includes determining whether the first part of the user's body is in a first state, and the method includes, based on determining that the first part of the user's body is in the first state, abandoning the display of the first fixation target associated with the first optional user interface object in the user interface.

12. The method according to any one of claims 1 to 11, wherein the user interface includes a second optional user interface object that can be selected to perform a second operation, and the method further comprises: Detecting that the user's attention of the computer system points to the second optional user interface object; and In response to detecting that the user's attention points to the second optional user interface object, display, in the user interface, a second fixation target associated with the second optional user interface object, wherein the first fixation target and the second fixation target have the same visual appearance.

13. The method according to any one of claims 1 to 12, wherein: Displaying the first selectable user interface object when the user's attention is directed to the first selectable user interface object includes displaying the first selectable user interface object in a first visual appearance; and the user interface includes a second selectable user interface object selectable to perform a second operation, and the method further includes: when the user's attention is directed to the first selectable user interface object, detecting that the user's attention of the computer system has changed to be directed to the second selectable user interface object; and In response to detecting that the user's attention is directed to the second selectable user interface object, displaying the first selectable user interface object in a second visual appearance different from the first visual appearance.

14. The method according to any one of claims 1 to 13, wherein the user interface includes a second selectable user interface object selectable to perform a second operation, and the method further includes: When displaying the first fixation target associated with the first selectable user interface object, detecting that the user's attention of the computer system is directed to the second selectable user interface object; And In response to detecting that the user's attention is directed to the second selectable user interface object, stopping displaying the first fixation target associated with the first selectable user interface object.

15. The method according to any one of claims 1 to 14, wherein the user interface includes a second selectable user interface object selectable to perform a second operation, and the method further includes: Detecting that the user's attention of the computer system is directed to the second selectable user interface object; And In response to detecting that the user's attention is directed to the second selectable user interface object, displaying a second fixation target associated with the second selectable user interface object in the user interface.

16. The method according to any one of claims 1 to 15, wherein: The first selectable user interface object is displayed as being composed of an analog material having a thickness; and displaying the first fixation target includes displaying the first fixation target embedded in the surface of the analog material of the first selectable user interface object.

17. The method according to any one of claims 1 to 16, the method further includes: When displaying the user interface including a scrollable area, detecting the user's attention directed to the first edge area of the scrollable area via the one or more input devices, wherein the first selectable user interface object is included in the scrollable area; And In response to detecting the user's attention directed to the first edge area, scrolling the scrollable area according to the user's attention directed to the first edge area, including scrolling the first selectable user interface object from a first position to a second position in the user interface.

18. The method according to claim 17, wherein when it is detected that the user's attention is directed to the first edge area of the scrollable area, the first fixation target associated with the first optional user interface object is displayed, and the method further includes: When scrolling the scrollable area, stop displaying the first fixation target associated with the first optional user interface object.

19. The method according to any one of claims 1 to 18, the method further includes: When displaying the user interface including the scrollable area, detect a first input directed to the scrollable area via the one or more input devices, wherein the first input includes a corresponding gesture performed by a corresponding part of the user's body of the computer system corresponding to a request to scroll the scrollable area; and In response to detecting the first input, scroll the scrollable area according to the first input, including scrolling the first optional user interface object from a first position to a second position different from the first position in the user interface.

20. The method according to any one of claims 1 to 19, wherein: Displaying the first optional user interface object includes displaying the first optional user interface object as a first element in the user interface; and Displaying the first fixation target includes displaying the first fixation target as a second element outside the first element in the user interface.

21. The method according to any one of claims 1 to 19, wherein: Displaying the first optional user interface object includes displaying the first optional user interface object as a first element in the user interface; and Displaying the first fixation target includes displaying the first fixation target within the first element in the user interface.

22. The method according to any one of claims 1 to 21, the method further includes: When detecting the user's attention directed to the first fixation target, output a first feedback based on the user's attention directed to the first fixation target.

23. The method according to claim 22, wherein outputting the first feedback includes outputting a feedback having a corresponding characteristic that changes based on the progress of the user's attention towards meeting the one or more criteria.

24. The method according to any one of claims 22 to 23, the method further includes: When outputting the first feedback, output a second feedback indicating that the one or more criteria are met according to a determination that the user's attention directed to the first fixation target meets the one or more criteria.

25. The method according to any one of claims 1 to 24, the method further includes: When displaying the first fixation target, detect that the user's attention deviates from the first optional user interface object; And In response to detecting that the user's attention deviates from the first optional user interface object: According to a determination that the user's attention meets one or more second criteria, reduce the visual salience of the first fixation target relative to the three-dimensional environment; and in response to determining that the user's attention does not meet the one or more second criteria, forgoing reducing the visual salience of the first fixation target relative to the three-dimensional environment.

26. The method of claim 25, the method further comprising: when displaying the first fixation target with reduced visual salience relative to the three-dimensional environment, detecting the user's attention directed to the first alternative user interface object; and in response to detecting the user's attention directed to the first alternative user interface object, increasing the visual salience of the first fixation target relative to the three-dimensional environment.

27. The method of any one of claims 1-26, wherein the one or more criteria include criteria that are met when the user's attention is directed to the first fixation target for longer than a corresponding threshold time period, wherein: in response to determining that one or more previous user interactions with one or more alternative user interface objects or one or more fixation targets meet one or more second criteria, the corresponding threshold time period is a first threshold time period, and in response to determining that the one or more previous user interactions with the one or more alternative user interface objects or the one or more fixation targets do not meet the one or more second criteria, the corresponding threshold time period is a second threshold time period different from the first threshold time period.

28. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, a user interface including a first alternative user interface object selectable to perform a first operation; when displaying the user interface, detecting, via the one or more input devices, the user's attention of the computer system directed to the first alternative user interface object; in response to detecting the user's attention directed to the first alternative user interface object, displaying, in the user interface, a first fixation target associated with the first alternative user interface object; when displaying the first fixation target, detecting the user's attention directed to the first fixation target; and in response to detecting the user's attention directed to the first fixation target: initiating, in response to determining that the user's attention meets one or more criteria, the first operation associated with the first alternative user interface object; and forgoing initiating, in response to determining that the user's attention does not meet the one or more criteria, the first operation associated with the first alternative user interface object.

29. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method comprising the following operations: Display, via the display generation component, a user interface including a first selectable user interface object for performing a first operation; When the user interface is displayed, detect, via the one or more input devices, that the attention of a user of the computer system is directed to the first selectable user interface object; In response to detecting that the user's attention is directed to the first selectable user interface object, display, in the user interface, a first fixation target associated with the first selectable user interface object; When the first fixation target is displayed, detect that the user's attention is directed to the first fixation target; And In response to detecting that the user's attention is directed to the first fixation target: Initiate, based on determining that the user's attention meets one or more criteria, the first operation associated with the first selectable user interface object; And Abstain from initiating, based on determining that the user's attention does not meet the one or more criteria, the first operation associated with the first selectable user interface object.

30. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; Means for: displaying, via the display generation component, a user interface including a first selectable user interface object for performing a first operation; Means for: when the user interface is displayed, detecting, via the one or more input devices, that the attention of a user of the computer system is directed to the first selectable user interface object; Means for: in response to detecting that the user's attention is directed to the first selectable user interface object, displaying, in the user interface, a first fixation target associated with the first selectable user interface object; Means for: when the first fixation target is displayed, detecting that the user's attention is directed to the first fixation target; And Means for: in response to detecting that the user's attention is directed to the first fixation target: Initiating, based on determining that the user's attention meets one or more criteria, the first operation associated with the first selectable user interface object; And Abstaining from initiating, based on determining that the user's attention does not meet the one or more criteria, the first operation associated with the first selectable user interface object.

31. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any one of the methods according to claims 1 to 27.

32. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 1 to 27.

33. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: One or more processors; Memory; And Means for performing any one of the methods according to claims 1 to 27.

34. A method, the method comprising: At a computer system in communication with a display generation component and one or more input devices: Displaying, via the display generation component, a user interface including a first selectable user interface object selectable to perform a first operation; When the user interface is displayed, detecting, via the one or more input devices, that a user of the computer system directs attention to the first selectable user interface object; And When it is detected that the user's attention is directed to the first selectable user interface object, displaying, in the user interface, a visual indication of the progress towards the user's attention meeting one or more criteria for activating the first selectable user interface object to perform the first operation based on the user's attention.

35. The method according to claim 34, the method further comprising: When it is detected that the user's attention is directed to the first selectable user interface object, displaying a visual indication of the user's attention in a region corresponding to the first selectable user interface object, wherein before the user's attention was directed to the first selectable user interface object, the visual indication of the user's attention was not displayed in the region corresponding to the first selectable user interface object.

36. The method according to claim 35, wherein the visual indication of the user's attention includes a virtual illumination effect applied to at least a portion of the first selectable user interface object.

37. The method according to any one of claims 35 to 36, wherein displaying the visual indication of the user's attention in the area corresponding to the first optional user interface object includes displaying the visual indication of the user's attention at a first position in the visual indication of the user's attention with a first visual intensity, and displaying the visual indication of the user's attention at a second position in the visual indication of the user's attention with a second visual intensity different from the first visual intensity, wherein the first position and the second position are at different distances from the center of the visual indication of the user's attention.

38. The method according to any one of claims 34 to 37, wherein displaying the visual indication of progress includes displaying the visual indication of progress in an area around the virtual object.

39. The method according to any one of claims 34 to 38, wherein displaying the visual indication of progress includes displaying a filling corresponding to the virtual object towards which the user's attention is directed to the progress satisfying the one or more criteria.

40. The method according to any one of claims 34 to 39, the method further comprising: When it is detected that the user's attention is directed to the first optional user interface object, outputting a first audio feedback indicating the progress of the user's attention towards satisfying the one or more criteria for activating the first optional user interface object.

41. The method according to claim 40, wherein outputting the first audio feedback includes changing the corresponding characteristics of the first audio feedback, including the volume changed based on the progress of the user's attention towards satisfying the one or more criteria.

42. The method according to any one of claims 40 to 41, wherein outputting the first audio feedback includes changing the corresponding characteristics of the first audio feedback, including the pitch changed based on the progress of the user's attention towards satisfying the one or more criteria.

43. The method according to any one of claims 40 to 42, the method further comprising: When outputting the first feedback, detecting that the user's attention directed to the first optional user interface object satisfies the one or more criteria; and In response to detecting that the user's attention directed to the first optional user interface object satisfies the one or more criteria, outputting a second audio feedback indicating the activation of the first optional user interface object to perform the first operation.

44. The method according to any one of claims 34 to 43, wherein: The first optional user interface object is a first type of object; and the user interface includes a second optional user interface object selectable to perform a second operation, wherein the second optional user interface object is a second type of object different from the first type of object, and the method further comprises: Detecting that the user's attention is directed to the second optional user interface object; and In response to detecting that the user's attention is directed to the second selectable user interface object, display a third selectable user interface object in the user interface, where the third selectable user interface target is of the first type of object and is selectable to perform the second operation.

45. The method according to any one of claims 34 to 44, the method further comprising: When displaying the indication of the progress of the user's attention meeting the one or more criteria, detect that the user's attention deviates from the first selectable user interface object, where the current state of the indication of the progress corresponds to a first progress amount towards meeting the one or more criteria; In response to detecting that the user's attention deviates from the first selectable user interface object, update the current state of the indication of the progress to correspond to a second progress amount towards meeting the one or more criteria, where the second progress amount is less than the first progress amount; When the current state of the indication of the progress corresponds to the second progress amount, detect that the user's attention is directed to the first selectable user interface object via the one or more input devices; And in response to detecting that the user's attention is directed to the first selectable user interface object, update the current state of the indication of the progress to resume from the second progress amount towards meeting the one or more criteria.

46. The method according to any one of claims 34 to 45, the method further comprising: When displaying the indication of the progress of the user's attention meeting the one or more criteria, detect a change in the position of the user's attention including the user's attention moving away from the first selectable user interface object and then moving back to the first selectable user interface object, where the current state of the indication of the progress corresponds to a first progress amount towards meeting the one or more criteria; And In response to detecting the change in the position of the user's attention: According to determining that the change in the position of the user's attention meets one or more second criteria, update the current state of the indication of the progress to continue from the first progress amount towards meeting the one or more criteria, the one or more second criteria including criteria met when the user's attention changes back to the first selectable object within a threshold time of changing away from the first selectable object; And According to determining that the change in the position of the user's attention does not meet the one or more second criteria, update the current state of the indication of the progress to resume from a second progress amount less than the first progress amount towards meeting the one or more criteria.

47. The method according to any one of claims 34 to 46, the method further comprising: When displaying the user interface including a scrollable area, detect the user's attention directed to the first edge area of the scrollable area via the one or more input devices, where the first selectable user interface object is included in the scrollable area; and in response to detecting the user's attention directed to the first edge region, scrolling the scrollable region according to the user's attention directed to the first edge region, including scrolling the first selectable user interface object from a first position to a second position in the user interface.

48. The method according to any one of claims 34 to 47, the method further comprising: when displaying the user interface including the scrollable region, detecting a first input directed to the scrollable region via the one or more input devices, wherein the first input includes a corresponding gesture performed by a corresponding part of the body of the user of the computer system corresponding to a request to scroll the scrollable region; and in response to detecting the first input, scrolling the scrollable region according to the first input, including scrolling the first selectable user interface object from a first position to a second position different from the first position in the user interface.

49. The method according to any one of claims 34 to 48, wherein the user interface includes a second selectable user interface object for performing the first operation, the method further comprising: when displaying the user interface and before displaying the first selectable user interface object, detecting the user's attention directed to the second selectable user interface object; and in response to detecting the user's attention directed to the second selectable user interface object, displaying the first selectable user interface object as a fixation target of the second selectable user interface object in the user interface.

50. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, a user interface including a first selectable user interface object for performing a first operation; when displaying the user interface, detecting, via the one or more input devices, the attention of the user of the computer system directed to the first selectable user interface object; and when detecting the user's attention directed to the first selectable user interface object, displaying, in the user interface, a visual indication of the progress towards the user's attention meeting one or more criteria for activating the first selectable user interface object to perform the first operation based on the user's attention.

51. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method including: displaying, via the display generation component, a user interface including a first selectable user interface object for performing a first operation; When displaying the user interface, detect, via the one or more input devices, that the attention of a user of the computer system is directed to the first selectable user interface object; And When it is detected that the attention of the user is directed to the first selectable user interface object, display, in the user interface, a visual indication of the progress towards the user's attention meeting one or more criteria for activating the first selectable user interface object to perform the first operation based on the user's attention.

52. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; Means for: displaying, via the display generation component, a user interface including a first selectable user interface object selectable to perform a first operation; Means for: when displaying the user interface, detecting, via the one or more input devices, that the attention of a user of the computer system is directed to the first selectable user interface object; And Means for: when it is detected that the attention of the user is directed to the first selectable user interface object, displaying, in the user interface, a visual indication of the progress towards the user's attention meeting one or more criteria for activating the first selectable user interface object to perform the first operation based on the user's attention.

53. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one of the methods according to claims 34 to 49 and 246 to 252.

54. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 34 to 49 and 246 to 252.

55. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And Means for performing any one of the methods according to claims 34 to 49 and 246 to 252.

56. A method, the method comprising: At a computer system communicating with a display generation component and one or more input devices: Displaying, via the display generation component, a user interface including a first selectable user interface object selectable to perform a first operation; When displaying the user interface, detecting, via the one or more input devices, that the attention of a user of the computer system is directed to the first selectable user interface object; When it is detected that the user's attention is directed to the first optional user interface object: Based on determining that the user's attention meets one or more criteria, display, in the user interface, a second optional user interface object that is selectable based on the user's attention and is used to perform the first operation associated with the first optional user interface object, where the one or more criteria include criteria that are met when the user's attention remains directed to the first optional user interface object for a duration greater than a time threshold; And Based on determining that an activation input that is different from the user's attention and is directed to the first optional user interface object is received before the user's attention meets the one or more criteria, perform the first operation associated with the first optional user interface object.

57. The method according to claim 56, the method further comprising: When the second optional user interface object is displayed, detect that the user's attention is directed to the second optional user interface object; And In response to detecting that the user's attention is directed to the second optional user interface object: Based on determining that the user's attention directed to the second optional user interface object meets one or more second criteria, perform the first operation associated with the first optional user interface object, where the one or more second criteria include criteria that are met when the user's attention remains directed to the second optional user interface object for a duration longer than a second time threshold; And Based on determining that the user's attention directed to the second optional user interface object does not meet the one or more second criteria, abandon performing the first operation associated with the first optional user interface object.

58. The method according to claim 57, the method further comprising: When it is detected that the user's attention is directed to the second optional user interface object: Based on determining that an activation input that is different from the user's attention and is directed to the second optional user interface object is received before the user's attention meets the one or more second criteria, perform the first operation associated with the first optional user interface object.

59. The method according to any one of claims 57 to 58, the method further comprising: When it is detected that the user's attention is directed to the second optional user interface object, display, in the user interface, a visual indication of the progress towards the user's attention directed to the second optional user interface object meeting the one or more second criteria.

60. The method according to any one of claims 56 to 59, where the first optional user interface object is a first type of object, and the user interface includes a third optional user interface object, the method further comprising: When the third optional user interface object is displayed, detect that the user's attention is directed to the third optional user interface object and the user's attention directed to the third optional user interface object meets one or more second criteria; And In response to detecting that the user's attention directed to the third optional user interface object satisfies the one or more second criteria: Based on determining that the third optional user interface object is the first type of object, display in the user interface a fourth optional user interface object that is selectable based on the user's attention for performing a second operation associated with the third optional user interface object; And Based on determining that the third optional user interface object is a second type of object different from the first type of object, perform the second operation associated with the third optional user interface object.

61. The method according to claim 60, the method further comprising: When the user's attention is directed to the third optional user interface object and before the user's attention directed to the third optional user interface object satisfies the one or more second criteria, detect an activation input received that is different from the user's attention and is directed to the third optional user interface object; And In response to detecting the activation input directed to the third user interface object: Based on determining that the third optional user interface object is the second type of object, perform the second operation associated with the third optional user interface object.

62. The method according to any one of claims 60 to 61, the method further comprising: When the user's attention is directed to the third optional user interface object, display in the user interface a visual indication of the progress towards the user's attention directed to the third optional user interface object satisfying the one or more second criteria.

63. The method according to any one of claims 56 to 62, wherein: The user interface includes a third optional user interface object that is selectable for performing a second operation, and the method further comprising: When the user's attention is directed to the first optional user interface object and when the third optional user interface object is displayed with a first visual appearance, detect that the user's attention of the computer system has become directed to the third optional user interface object; and In response to detecting the user's attention directed to the third optional user interface object, display the third optional user interface object with a second visual appearance different from the first visual appearance.

64. The method according to any one of claims 56 to 63, the method further comprising: When detecting that the user's attention is directed to the first optional user interface object, display in the user interface a visual indication of the user's attention directed to the first optional user interface object.

65. The method according to any one of claims 56 to 64, wherein: The first optional user interface object is the first type of object; The second optional user interface object is a second type of object different from the first type of object.

66. The method according to any one of claims 56 to 65, wherein: The user interface includes a first set of user interface objects, the first set of user interface objects including the first optional user interface object and the third optional user interface object, the third optional user interface object being selectable for displaying a fourth optional user interface object based on the user's attention, and the fourth optional user interface object being selectable for performing a second operation associated with the third optional user interface object based on the user's attention.

67. The method according to any one of claims 56 to 66, wherein the user interface includes a first set of user interface objects, the first set of user interface objects including the second optional user interface object and the third optional user interface object, the third optional user interface object being associated with a fourth optional user interface object, and the method further includes: When displaying the first set of user interface objects, detecting the user's attention directed to the first set of user interface objects; And In response to detecting the user's attention directed to the first set of user interface objects: Performing the first operation associated with the first optional user interface object according to determining that the user's attention directed to the second optional user interface object meets one or more second criteria; And Performing the second operation associated with the fourth optional user interface object according to determining that the user's attention directed to the third optional user interface object meets the one or more second criteria, wherein the second operation is different from the first operation.

68. The method according to any one of claims 56 to 67, wherein: The first optional user interface object is a first type of object; The second optional user interface object is a second type of object different from the first type of object; and The user interface includes a first set of user interface objects of the first type and a second set of user interface objects of the second type, wherein the first set of user interface objects includes the first optional user interface object and the third optional user interface object, and the second set of user interface objects includes the second optional user interface object and the fourth optional user interface object.

69. The method according to any one of claims 56 to 68, wherein when detecting the user's attention directed to the first optional user interface object: According to determining that the user mainly uses only attention input to interact with the device, the time threshold is a first time threshold, and According to determining that the user uses input including non-attention activation input to interact with the computer system, the time threshold is a second time threshold greater than the first time threshold.

70. A computer system communicating with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: Display, via the display generation component, a user interface including a first selectable user interface object for performing a first operation; When the user interface is displayed, detect, via the one or more input devices, the attention of a user of the computer system being directed to the first selectable user interface object; When the attention of the user being directed to the first selectable user interface object is detected: Based on determining that the attention of the user meets one or more criteria, display, in the user interface, a second selectable user interface object for performing the first operation associated with the first selectable user interface object based on the attention of the user, the one or more criteria including criteria that are met when the attention of the user remains directed to the first selectable user interface object for a duration greater than a time threshold; And Based on determining that an activation input different from the attention of the user and directed to the first selectable user interface object is received before the attention of the user meets the one or more criteria, perform the first operation associated with the first selectable user interface object.

71. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: Display, via the display generation component, a user interface including a first selectable user interface object for performing a first operation; When the user interface is displayed, detect, via the one or more input devices, the attention of a user of the computer system being directed to the first selectable user interface object; When the attention of the user being directed to the first selectable user interface object is detected: Based on determining that the attention of the user meets one or more criteria, display, in the user interface, a second selectable user interface object for performing the first operation associated with the first selectable user interface object based on the attention of the user, the one or more criteria including criteria that are met when the attention of the user remains directed to the first selectable user interface object for a duration greater than a time threshold; And Based on determining that an activation input different from the attention of the user and directed to the first selectable user interface object is received before the attention of the user meets the one or more criteria, perform the first operation associated with the first selectable user interface object.

72. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; Means for: Displaying, via the display generation component, a user interface including a first selectable user interface object for performing a first operation; Apparatus for: when displaying the user interface, detecting, via the one or more input devices, the attention of a user of the computer system being directed to the first selectable user interface object; Apparatus for: when detecting that the attention of the user is directed to the first selectable user interface object: displaying, in the user interface, a second selectable user interface object that is selectable based on the attention of the user to perform the first operation associated with the first selectable user interface object, according to determining that the attention of the user meets one or more criteria, the one or more criteria including criteria that are met when the attention of the user remains directed to the first selectable user interface object for a duration greater than a time threshold; and performing the first operation associated with the first selectable user interface object according to determining that an activation input that is different from the attention of the user and is directed to the first selectable user interface object is received before the attention of the user meets the one or more criteria.

73. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods according to claims 56 to 69.

74. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform any of the methods according to claims 56 to 69.

75. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and means for performing any of the methods according to claims 56 to 69.

76. A method, the method comprising: at a computer system in communication with a display generation component and one or more input devices: displaying, via the display generation component, a user interface including a selectable object having a region; when displaying the user interface including the selectable object, detecting, via the one or more input devices, the attention of a user of the computer system; in response to detecting the attention of the user, displaying, via the display generation component, a visual indication in the region of the selectable object that emphasizes a first position in the region of the selectable object according to determining that the attention of the user is directed to the first position in the region of the selectable object; when the visual indication is displayed in the region of the selectable object, detecting, via the one or more input devices, a movement of the attention of the user; and In response to detecting the movement of the user's attention, according to determining a second position in the area of the optional object where the user's attention points, change the appearance of the visual indication in the area of the optional object so as to emphasize the second position in the area of the optional object.

77. The method according to claim 76, wherein the visual indication simultaneously emphasizes the first position in the area of the optional object and the second position in the area of the optional object.

78. The method according to claim 76, wherein: Displaying, by the display generation component, the visual indication that emphasizes the first position in the area of the optional object in the area of the optional object includes displaying the visual indication that emphasizes the first position in the area of the optional object without emphasizing the second position in the area of the optional object; and Changing the appearance of the visual indication in the area of the optional object so as to emphasize the second position in the area of the optional object includes displaying the visual indication that emphasizes the second position in the area of the optional object without emphasizing the first position in the area of the optional object.

79. The method according to any one of claims 76 to 78, wherein the user's attention is based on the user's finger.

80. The method according to any one of claims 76 to 79, wherein the user's attention is based on the user's gaze.

81. The method according to any one of claims 76 to 80, wherein the user's attention is detected at least in part via an input from a touch-sensitive surface communicating with the computer system.

82. The method according to any one of claims 76 to 81, wherein the visual indication in the area of the optional object has a shape obtained by masking a first shape corresponding to the user's attention and a second shape of the optional object.

83. The method according to any one of claims 76 to 82, wherein displaying, by the display generation component, the visual indication that emphasizes the first position in the area of the optional object in the area of the optional object includes: According to determining that the size of the optional object is a first size, the visual indication has a second size based on the first size; And According to determining that the size of the optional object is a third size different from the first size, the visual indication has a fourth size based on the third size, wherein the fourth size is different from the second size.

84. The method according to claim 83, wherein displaying, by the display generation component, the visual indication that emphasizes the first position in the area of the optional object in the area of the optional object includes: Based on determining that the size of the optional object is greater than a threshold size, wherein the first size and the third size are less than the threshold size, the visual indication has a maximum size that is not based on the size of the optional object.

85. The method according to claim 83 or 84, wherein displaying the visual indication emphasizing the first position in the region of the optional object via the display generation component includes: Based on determining that the size of the optional object is less than a threshold size, wherein the first size and the third size are greater than the threshold size, the visual indication has a minimum size that is not based on the size of the optional object.

86. The method according to any one of claims 83 to 85, wherein the size of the visual indication is greater than the size of the minimum dimension of the optional object.

87. The method according to any one of claims 76 to 86, wherein the visual indication is partially transparent.

88. The method according to any one of claims 76 to 87, wherein in response to detecting the movement of the user's attention, based on determining that the user's attention points to the second position in the region of the optional object, changing the appearance of the visual indication in the region of the optional object to emphasize the second position in the region of the optional object includes smoothly moving the visual indication from the first position in the region of the optional object to the second position in the region of the optional object, wherein one or more characteristics of the movement of the visual indication from the first position in the region of the optional object to the second position in the region of the optional object are different from one or more characteristics of the movement of the user's attention from the first position to the second position.

89. The method according to any one of claims 76 to 88, wherein: Displaying the visual indication emphasizing the first position in the region of the optional object includes displaying the visual indication, wherein a first part of the visual indication is obscured by the optional object, and changing the appearance of the visual indication in the region of the optional object to emphasize the second position in the region of the optional object includes displaying the visual indication, wherein a second part of the visual indication is obscured by the optional object, and the second part is different from the first part of the visual indication.

90. The method according to any one of claims 76 to 89, the method includes: In response to detecting the movement of the user's attention, and based on determining that the user's attention points to a third position in a second region of a second optional object different from the optional object: Stop displaying the visual indication in the optional object; And The visual indication emphasizing the third position in the second region of the second alternative object is displayed at the third position in the second region of the second alternative object via the display generation component.

91. The method according to claim 90, wherein: The alternative object is a first alternative object, Displaying the visual indication emphasizing the first position in the region of the first alternative object in the region of the first alternative object includes displaying the visual indication, wherein a first part of the visual indication is masked by the first alternative object, and displaying the visual indication emphasizing the third position in the second region of the second alternative object at the third position in the second region of the second alternative object includes displaying the visual indication, wherein a second part of the visual indication different from the first part of the visual indication is obscured by the second alternative object.

92. The method according to any one of claims 76 to 91, wherein the alternative object is a key on a keyboard in the user interface.

93. The method according to any one of claims 76 to 92, wherein the alternative object is a button in the user interface.

94. The method according to any one of claims 76 to 93, wherein the alternative object is an alternative tray in the user interface.

95. The method according to any one of claims 76 to 94, the method comprising: In response to detecting the user's attention and based on determining that the user's attention is directed to the first position in the region of the alternative object, changing the visual appearance of the alternative object outside the region in which the visual indication is displayed for the alternative object.

96. The method according to claim 95, wherein changing the appearance of the alternative object outside the region in which the visual indication is displayed includes displaying a simulated specular highlight around or on at least one edge of the alternative object.

97. The method according to claim 96, wherein displaying the simulated specular highlight around or on at least one edge of the alternative object is performed based on the physical illumination of the physical environment of the user of the computer system.

98. The method according to claim 96 or 97, wherein displaying the simulated specular highlight around at least one edge of the alternative object is performed based on the simulated illumination of a three-dimensional environment displayed by the computer system.

99. The method according to any one of claims 96 to 98, the method comprising: Detecting, via the one or more input devices, the user's attention directed to the corresponding user interface including the alternative object and to a part of the corresponding user interface different from the alternative object; and in response to detecting the user's attention directed to the corresponding user interface, displaying the simulated mirror highlight on the portion of the corresponding user interface according to determining that the user's attention is directed to the portion of the corresponding user interface, wherein one or more of the characteristics of the simulated mirror highlight on the portion of the corresponding user interface are the same as one or more of the characteristics of the simulated mirror highlight on the optional option.

100. The method according to any one of claims 95 to 99, wherein changing the appearance of the optional object outside the region in which the visual indication is displayed includes highlighting the optional object.

101. The method according to any one of claims 95 to 100, the method comprising: in response to detecting the user's attention and according to determining that the user's attention is directed to the first position in the region of the optional object, changing the visual separation between the optional object and the corresponding portion of the user interface.

102. The method according to any one of claims 76 to 101, wherein the visual appearance of the visual indication in the region of the optional object changes based on the distance from the position of the user's attention.

103. The method according to any one of claims 76 to 102, wherein: detecting the user's attention of the computer system via the one or more input devices includes detecting the user's gaze, and the method further comprises: in response to detecting the user's attention and according to determining that one or more criteria are met, changing the visual separation between the optional object and the corresponding portion of the user interface, the one or more criteria including criteria met when the user's attention is directed to the optional object for a predetermined period of time.

104. The method according to claim 103, the method further comprising displaying a fixation target associated with the optional object via the display generating component in response to detecting the user's attention and according to determining that the one or more criteria are met.

105. The method according to any one of claims 76 to 104, wherein the user interface includes at least the optional object and a second optional object, and wherein the method comprises: in response to detecting the user's attention and according to determining that the user's attention is directed to the optional object: displaying the optional object with a first visual characteristic having a first value and displaying the second optional object with the first visual characteristic having the first value according to determining that the interaction with the user interface is according to a first mode; and displaying the optional object with the first visual characteristic having the first value and displaying the second optional object with the first visual characteristic having a second value different from the first value according to determining that the interaction with the user interface is according to a second mode different from the first mode.

106. The method according to claim 105, the method comprising: In response to detecting the user's attention: Based on determining that the user's attention is directed to the selectable object, display the selectable object with the first visual characteristic having the first value and the second visual characteristic having the third value, and display the second selectable object with the second visual characteristic having a fourth value different from the third value; and Based on determining that the user's attention is directed to the second selectable object, display the second selectable object with the first visual characteristic having the first value and the second visual characteristic having the third value, and display the selectable object with the second visual characteristic having the fourth value.

107. The method according to any one of claims 76 to 106, wherein displaying the visual indication in the region of the selectable object includes: When the interaction with the user interface is according to a first mode, wherein detecting the user's attention of the computer system via the one or more input devices includes detecting a spatial interaction between a part of the user and the selectable object: Based on determining that the part of the user is at a first distance from the selectable object, display the visual indication with a first visual appearance; and Based on determining that the part of the user is at a second distance different from the first distance from the selectable object, display the visual indication with a second visual appearance different from the first visual appearance.

108. The method according to claim 107, the method including: When the interaction with the user interface is according to a second mode different from the first mode, wherein detecting the user's attention of the computer system via the one or more input devices includes detecting the user's gaze: Based on determining that the part of the user is at a third distance from the selectable object, display the visual indication with a third visual appearance; and Based on determining that the part of the user is at a fourth distance different from the third distance from the selectable object, display the visual indication with the third visual appearance.

109. The method according to any one of claims 76 to 108, wherein displaying the visual indication in the region of the selectable object includes: Based on determining that the movement of the user's attention is higher than a threshold speed, reducing the visual saliency of the visual indication during the movement of the user's attention.

110. The method according to any one of claims 76 to 109, wherein displaying the visual indication in the region of the selectable object includes: Based on determining that the movement of the user's attention is lower than a threshold speed, increasing the visual saliency of the visual indication during the movement of the user's attention.

111. The method according to any one of claims 76 to 110, the method including: When displaying the visual indication in the region of the selectable object, detecting a selection input corresponding to a request to select the selectable object via the one or more input devices; and In response to detecting the selection input corresponding to the request for selecting the optional object, perform an operation associated with the optional object.

112. The method according to claim 111, wherein: the selection input corresponding to the request for selecting the optional object includes a first part and a second part, and the method includes: in response to detecting the first part of the selection input, increase the visual saliency of the visual indication.

113. The method according to claim 111 or 112, wherein: the selection input corresponding to the request for selecting the optional object includes a first part and a second part, and the method includes: in response to detecting the first part of the selection input, change the visual appearance of the optional object.

114. The method according to any one of claims 76 to 113, the method including: in response to detecting the selection input corresponding to the request for selecting the optional object: depending on determining that the optional object has a size smaller than a threshold size, display, via the display generation component, a visual selection indication corresponding to the selection of the optional object.

115. The method according to claim 114, wherein the visual selection indication corresponding to the selection of the optional object at least partially surrounds the optional object.

116. The method according to claim 115, wherein: the selection input corresponding to the request for selecting the optional object includes a first part and a second part, and the method includes: in response to detecting the first part of the selection input, display, with a first size, the selection indication surrounding the optional object corresponding to the selection of the optional object, and in response to detecting the second part of the selection input, display, with a second size different from the second size, the selection indication surrounding the optional object corresponding to the selection of the optional object.

117. A computer system communicating with a display generation component and one or more input devices, the computer system including: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: display, via the display generation component, a user interface including an optional object having a region; when displaying the user interface including the optional object, detect, via the one or more input devices, the attention of a user of the computer system; in response to detecting the attention of the user, depending on determining that the attention of the user points to a first position in the region of the optional object, display, via the display generation component, a visual indication emphasizing the first position in the region of the optional object in the region of the optional object; when displaying the visual indication in the region of the optional object, detect, via the one or more input devices, the movement of the attention of the user; and In response to detecting the movement of the user's attention, according to determining that the user's attention points to a second position in the region of the selectable object, change the appearance of the visual indication in the region of the selectable object to emphasize the second position in the region of the selectable object.

118. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: Display, via the display generation component, a user interface including a selectable object having a region; When displaying the user interface including the selectable object, detect the attention of a user of the computer system via the one or more input devices; In response to detecting the user's attention, according to determining that the user's attention points to a first position in the region of the selectable object, display, via the display generation component, a visual indication in the region of the selectable object that emphasizes the first position in the region of the selectable object; When the visual indication is displayed in the region of the selectable object, detect the movement of the user's attention via the one or more input devices; And In response to detecting the movement of the user's attention, according to determining that the user's attention points to a second position in the region of the selectable object, change the appearance of the visual indication in the region of the selectable object to emphasize the second position in the region of the selectable object.

119. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; Means for: displaying, via the display generation component, a user interface including a selectable object having a region; Means for: when displaying the user interface including the selectable object, detecting the attention of a user of the computer system via the one or more input devices; Means for: in response to detecting the user's attention, according to determining that the user's attention points to a first position in the region of the selectable object, displaying, via the display generation component, a visual indication in the region of the selectable object that emphasizes the first position in the region of the selectable object; Means for: when the visual indication is displayed in the region of the selectable object, detecting the movement of the user's attention via the one or more input devices; And Means for: in response to detecting the movement of the user's attention, according to determining that the user's attention points to a second position in the region of the selectable object, changing the appearance of the visual indication in the region of the selectable object to emphasize the second position in the region of the selectable object.

120. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods according to claims 76 to 116.

121. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, cause the computer system to perform any of the methods according to claims 76 to 116.

122. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And Means for performing any of the methods according to claims 76 to 116.

123. A method, the method comprising: At a computer system that communicates with a display generation component and one or more input devices: Displaying a user interface via the display generation component; When the user interface is displayed, detecting an input pointing to a first region of the user interface via the one or more input devices; and In response to detecting the input pointing to the first region of the user interface: Displaying a magnified view of the first region of the user interface via the display generation component according to determining that the first region of the user interface includes at least two selectable objects that meet a first criterion; And Abandoning displaying the magnified view of the first region of the user interface via the display generation component according to determining that the first region of the user interface does not include the at least two selectable objects that meet the first criterion.

124. The method according to claim 123, the method comprising: In response to detecting the input pointing to the first region of the user interface: Performing an operation associated with the one selectable object according to determining that the first region of the user interface includes no more than one selectable object that meets the first criterion.

125. The method according to claim 123 or 124, the method comprising: In response to detecting the input pointing to the first region of the user interface: Abandoning performing one or more operations associated with the at least two selectable objects in the first region of the user interface according to determining that the first region of the user interface includes the at least two selectable objects that meet the first criterion.

126. The method according to any one of claims 123 to 125, the method comprising: When the magnified view of the first region of the user interface is displayed, detecting a selection input pointing to the content of the magnified view of the first region of the user interface via the one or more input devices; and In response to detecting the selection input for the content of the magnified view of the first region of the user interface: Perform an operation associated with the first selectable object according to determining that the attention of the user of the computer system is directed to the first selectable object among the at least two selectable objects in the magnified view of the first region; And Perform an operation associated with the second selectable object according to determining that the attention of the user of the computer system is directed to a second selectable object different from the first selectable object among the at least two selectable objects in the magnified view of the first region.

127. The method according to any one of claims 123 to 126, the method comprising: When displaying the magnified view of the first region of the user interface, detect the attention of the user of the computer system via the one or more input devices; And In response to detecting the attention of the user: Stop displaying the magnified view of the first region according to determining that the attention of the user is not directed to the magnified view of the first region.

128. The method according to any one of claims 123 to 127, wherein displaying the magnified view of the first region of the user interface comprises: Simultaneously display: The magnified view of the first region of the user interface; And At least one region of the user interface that was displayed when the input directed to the first region of the user interface was detected.

129. The method according to any one of claims 123 to 128, the method comprising: When displaying the user interface, detect a second input directed to a second region of the user interface via the one or more input devices; And In response to detecting the second input directed to the second region of the user interface: Display a magnified view of the second region of the user interface without displaying the magnified view of the first region of the user interface according to determining that the second region includes a set of two or more selectable objects that meet the first criterion.

130. The method according to any one of claims 123 to 129, the method comprising: When displaying the user interface, detect a second input directed to a second region of the user interface via the one or more input devices; And In response to detecting the second input directed to the second region of the user interface: Abandon displaying a magnified view of the second region of the user interface via the display generation component according to determining that the second region of the user interface does not include a set of two or more selectable objects that meet the first criterion.

131. The method according to any one of claims 123 to 130, wherein the input directed to the first region of the user interface includes the attention of the user of the computer system directed to the first region and the corresponding part of the user performing a corresponding air gesture.

132. The method according to any one of claims 123 to 131, wherein the input pointing to the first region of the user interface comprises the attention of a user of the computer system pointing to the first region of the user interface for at least a threshold time period.

133. The method according to any one of claims 123 to 132, wherein the user interface comprises a website and the at least two selectable objects are links in the website.

134. The method according to any one of claims 123 to 133, wherein the first criterion comprises a criterion that is satisfied when the at least two selectable objects are within a threshold angular distance of each other with respect to the viewpoint of the user of the computer system.

135. The method according to any one of claims 123 to 134, wherein the content of the user interface in the magnified view of the first region is based on the position of the attention of the user of the computer system in the user interface when the input pointing to the first region of the user interface is detected.

136. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for: displaying a user interface via the display generation component; when the user interface is displayed, detecting an input pointing to a first region of the user interface via the one or more input devices; and in response to detecting the input pointing to the first region of the user interface: displaying a magnified view of the first region of the user interface via the display generation component based on determining that the first region of the user interface comprises at least two selectable objects that satisfy a first criterion; and abandoning displaying the magnified view of the first region of the user interface via the display generation component based on determining that the first region of the user interface does not comprise the at least two selectable objects that satisfy the first criterion.

137. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method comprising: displaying a user interface via the display generation component; when the user interface is displayed, detecting an input pointing to a first region of the user interface via the one or more input devices; and in response to detecting the input pointing to the first region of the user interface: displaying a magnified view of the first region of the user interface via the display generation component based on determining that the first region of the user interface comprises at least two selectable objects that satisfy a first criterion; and Discard the enlarged view of the first region of the user interface via the display generating component based on determining that the first region of the user interface does not include the at least two selectable objects that meet the first criterion.

138. A computer system in communication with a display generating component and one or more input devices, the computer system comprising: One or more processors; A memory; Means for: displaying a user interface via the display generating component; Means for: when the user interface is displayed, detecting an input pointing to a first region of the user interface via the one or more input devices; And Means for: in response to detecting the input pointing to the first region of the user interface: Display an enlarged view of the first region of the user interface via the display generating component based on determining that the first region of the user interface includes at least two selectable objects that meet a first criterion; And Discard the enlarged view of the first region of the user interface via the display generating component based on determining that the first region of the user interface does not include the at least two selectable objects that meet the first criterion.

139. A computer system in communication with a display generating component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods according to claims 123 to 135.

140. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generating component and one or more input devices, cause the computer system to perform any of the methods according to claims 123 to 135.

141. A computer system in communication with a display generating component and one or more input devices, the computer system comprising: One or more processors; A memory; And Means for performing any of the methods according to claims 123 to 135.

142. A method, the method comprising: At a computer system in communication with a display generating component and one or more input devices: Display a user interface including a slider element via the display generating component, the slider element including a first visual indication of a current value associated with the slider element; When the user interface including the slider element is displayed, wherein the first visual indication of the current value of the slider element is at a first position in the slider element, detect that the attention of a user of the computer system is directed to a second position in the slider element different from the first position; And In response to detecting the attention of the user being directed to the second position in the slider element: Initiate a process of adjusting the corresponding value of the slider element to a first value based on the second position of the attention in the slider element, according to the determination that before detecting the user's attention directed to the second position, the user's attention had been directed to the first visual indication at the first position, and the user's attention directed to the first visual indication at the first position had met one or more criteria in a first set; and Abandon initiating the process of adjusting the corresponding value of the slider element, according to the determination that before detecting the user's attention directed to the second position, the user's attention did not meet the one or more criteria in the first set regarding the first visual indication, and move the first visual indication in the slider element based on the second position of the attention in the slider element, where the user's attention had met the one or more criteria in the first set.

143. The method according to claim 142, the method further comprising, in response to detecting the user's attention directed to the second position in the slider element, moving the first visual indication in the slider element based on the second position of the attention in the slider element, according to the determination that before detecting the user's attention directed to the second position, the user's attention had been directed to the first visual indication at the first position, and the user's attention directed to the first visual indication at the first position had met the one or more criteria in the first set.

144. The method according to any one of claims 142 to 143, the method further comprising, in response to detecting the user's attention directed to the second position in the slider element, abandoning moving the first visual indication in the slider element based on the second position of the attention in the slider element, according to the determination that before detecting the user's attention directed to the second position, the user's attention did not meet the one or more criteria in the first set regarding the first visual indication.

145. The method according to any one of claims 142 to 144, the method further comprising, in response to detecting the user's attention directed to the second position in the slider element, according to the determination that before detecting the user's attention directed to the second position, the user's attention had been directed to the first visual indication at the first position, and the user's attention directed to the first visual indication at the first position had met the one or more criteria in the first set: Reduce the visual salience of a second visual indication of the current value of the slider element at the first position in the slider element; and move the first visual indication in the slider element based on the second position of the attention in the slider element without adjusting the current value of the slider element.

146. The method according to claim 145, wherein the process of initiating the adjustment of the corresponding value of the slider element to the first value based on the attention at the second position in the slider element includes enhancing the visual saliency of the first value of the slider element at the second position in the slider element based on the current value of the slider value corresponding to the first value in the first visual indication.

147. The method according to any one of claims 142 to 146, the method further comprising: When the process of adjusting the corresponding value of the slider element to the first value is in progress, detecting that the attention of the user is directed to the second position in the slider element for a duration longer than a threshold time amount; And In response to detecting that the attention of the user is directed to the second position in the slider element for a duration longer than the threshold time amount and meeting a second set of one or more criteria, displaying, via the display generation component, a fixation target for updating the current value of the slider element to the first value.

148. The method according to claim 147, the method further comprising: When the fixation target for updating the current value of the slider element to the first value is displayed, detecting that the attention of the user is directed to the fixation target; And In response to detecting that the attention of the user is directed to the fixation target, and based on determining that the attention of the user meets a third set of one or more criteria, adjusting the corresponding value of the slider element to the first value based on the second position of the attention in the slider element.

149. The method according to any one of claims 147 to 148, the method further comprising: When detecting that the attention of the user is directed to the fixation target, displaying, via the display generation component, a third visual indication indicating the progress of the attention of the user towards meeting the third set of one or more criteria.

150. The method according to claim 149, the method further comprising: When detecting that the attention of the user is directed to the fixation target, and when the third visual indication indicating the progress of the attention of the user towards meeting the third set of one or more criteria is displayed via the display generation component and before meeting the third set of one or more criteria, detecting that the attention of the user is no longer directed to the fixation target; And In response to detecting that the attention of the user is no longer directed to the fixation target, updating the third visual indication to indicate that the progress towards meeting the third set of criteria is regressing.

151. The method according to any one of claims 147 to 150, the method further comprising: After detecting that the attention of the user is directed to the second position in the slider element for a duration longer than the threshold time amount, detecting that the attention of the user has moved away from the second position and towards a third position in the slider element different from the second position; And In response to detecting that the user's attention has moved away from the second position and towards the third position, move the first visual indication towards the third position in the slider element that is different from the second position.

152. The method according to claim 151, the method further comprising: In response to detecting that the user's attention has moved away from the second position and towards the third position, stop displaying the fixation target associated with the second position in the slider element.

153. The method according to any one of claims 147 to 152, wherein the second set of one or more criteria includes criteria that are satisfied when the first visual indication is within a threshold distance of the second position in the slider element, the method further comprising: In response to detecting that the user's attention is directed at the second position in the slider element for longer than the threshold amount of time: Based on determining that the second set of one or more criteria is not satisfied because the first visual indication is not within the threshold distance of the second position in the slider element, abandon displaying the fixation target associated with the second position in the slider element.

154. The method according to claim 153, wherein initiating the process of adjusting the corresponding value of the slider element to the first value based on the second position of the attention in the slider element includes moving the first visual indication in the slider element from the first position to the second position, the method further comprising: When the user's attention is directed at the second position in the slider element: During the first part of the movement of the first visual indication in the slider element, abandon displaying the fixation target associated with the second position; And After the first part of the movement of the first visual indication in the slider element, display the fixation target associated with the second position.

155. The method according to any one of claims 142 to 154, when the process of adjusting the corresponding value of the slider element to the first value based on the second position of the attention in the slider element is in progress: Detect that the user's attention has moved away from the slider element; and In response to detecting that the user's attention has moved away from the slider element, stop the process of adjusting the corresponding value of the slider element.

156. The method according to any one of claims 155, wherein stopping the process of adjusting the corresponding value of the slider element includes stopping the display of the first visual indication.

157. The method according to claim 156, wherein stopping the process of adjusting the corresponding value of the slider element includes moving the first visual indication to the first position in the slider element.

158. The method according to any one of claims 155 to 157, wherein stopping the process of adjusting the corresponding value of the slider element includes: Based on determining that the fixation target associated with the second position is being displayed, stop the display of the fixation target.

159. The method according to any one of claims 142 to 158, the method further comprising, in response to detecting that the user's attention is directed to the second position, displaying a fourth visual indication indicating the user's attention at the second position in the slider element.

160. The method according to any one of claims 142 to 159, the method further comprising: When the user's attention is directed to the second position in the slider element: Detect an input from a first part of the user's body via the one or more input devices; And In response to detecting the input from the first part of the user's body, initiate a second process of adjusting the corresponding value of the slider element to the first value based on the input from the first part of the user's body, regardless of whether the user's attention has ever met the first set of one or more criteria.

161. The method according to claim 160, wherein when the process of adjusting the corresponding value of the slider element to the first value based on the second position of the attention in the slider element is in progress and before adjusting the corresponding value of the slider element to the first value, detecting the input from the first part of the user's body, the method further comprising: In response to detecting the input from the first part of the user's body, adjust the corresponding value of the slider element to the first value based on the second position of the attention in the slider element.

162. The method according to any one of claims 159 to 161, wherein the second process of adjusting the corresponding value of the slider element includes adjusting the corresponding value of the slider element based on the movement of the first part of the user's body.

163. The method according to any one of claims 142 to 162, wherein the slider element includes a volume slider, and adjusting the corresponding value of the slider element corresponds to adjusting the volume level of the computer system.

164. The method according to any one of claims 142 to 163, wherein the slider element includes a color slider, and adjusting the corresponding value of the slider element corresponds to adjusting the color settings of the computer system.

165. The method according to any one of claims 142 to 164, wherein the slider element includes a content playback control slider, adjusting the corresponding value of the slider element corresponds to adjusting the current playback position within a content item, and initiating the process of adjusting the corresponding value of the slider element to the first value based on the second position of the attention in the slider element includes displaying, via the display generation component, a preview of the corresponding playback position corresponding to the current position of the first visual indication in the slider element.

166. The method according to any one of claims 142 to 165, the method further comprising: In response to detecting that the user's attention is directed to the second position, displaying a fourth visual indication indicating the user's attention at the second position in the slider element; When displaying the fourth visual indication indicating the user's attention at the second position in the slider element, detecting the movement of the user's attention relative to the slider element; And In response to detecting the movement of the user's attention relative to the slider element, moving the fourth visual indication indicating the user's attention in the slider element based on the movement of the user's attention relative to the slider element.

167. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for: Displaying via the display generation component a user interface including a slider element, the slider element including a first visual indication of a current value associated with the slider element; When displaying the user interface including the slider element, wherein the first visual indication of the current value of the slider element is at a first position in the slider element, detecting that the attention of a user of the computer system is directed to a second position in the slider element different from the first position; And In response to detecting that the user's attention is directed to the second position in the slider element: According to determining that before detecting the user's attention directed to the second position, the user's attention had been directed to the first visual indication at the first position, and the user's attention directed to the first visual indication at the first position had met a first set of one or more criteria, initiating a process of adjusting a corresponding value of the slider element to a first value based on the second position of the attention in the slider element; And According to determining that before detecting the user's attention directed to the second position, the user's attention had not met the first set of one or more criteria regarding the first visual indication, and the user's attention had met the first set of one or more criteria, moving the first visual indication in the slider element based on the second position of the attention in the slider element and abandoning initiating the process of adjusting the corresponding value of the slider element.

168. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: Display, via the display generation component, a user interface including a slider element, the slider element including a first visual indication of a current value associated with the slider element; When displaying the user interface including the slider element, where a first visual indication of the current value of the slider element is at a first position within the slider element, detect that the attention of a user of the computer system is directed to a second position within the slider element that is different from the first position; and in response to detecting that the user's attention is directed to the second position in the slider element: Initiate a process of adjusting the corresponding value of the slider element to a first value based on the second position of the attention in the slider element according to a determination that before detecting the user's attention directed to the second position, the user's attention had been directed to the first visual indication at the first position, and the user's attention directed to the first visual indication at the first position had met a first set of one or more criteria; and According to a determination that before detecting the user's attention directed to the second position, the user's attention did not meet the first set of one or more criteria regarding the first visual indication, abandon initiating the process of adjusting the corresponding value of the slider element, and move the first visual indication in the slider element based on the second position of the attention in the slider element where the user's attention had met the first set of one or more criteria.

169. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; Means for: displaying, via the display generation component, a user interface including a slider element, the slider element including a first visual indication of a current value associated with the slider element; Means for: when displaying the user interface including the slider element, where the first visual indication of the current value of the slider element is at a first position in the slider element, detecting that the attention of a user of the computer system is directed to a second position in the slider element different from the first position; and Means for: in response to detecting that the user's attention is directed to the second position in the slider element: Initiate a process of adjusting the corresponding value of the slider element to a first value based on the second position of the attention in the slider element according to a determination that before detecting the user's attention directed to the second position, the user's attention had been directed to the first visual indication at the first position, and the user's attention directed to the first visual indication at the first position had met a first set of one or more criteria; and According to a determination that before detecting the user's attention directed to the second position, the user's attention did not meet the first set of one or more criteria regarding the first visual indication, abandon initiating the process of adjusting the corresponding value of the slider element, and move the first visual indication in the slider element based on the second position of the attention in the slider element where the user's attention had met the first set of one or more criteria.

170. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods according to claims 142 to 166.

171. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, cause the computer system to perform any of the methods according to claims 142 to 166.

172. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And Means for performing any of the methods according to claims 142 to 166.

173. A method, the method comprising: At a first computer system that communicates with a display generation component and one or more input devices: Displaying, via the display generation component, a user interface including user interface elements; when the user interface elements are displayed, detecting that the attention of a user of the computer system is directed to a first portion of the user interface; And in response to detecting that the attention of the user of the computer system is directed to the first portion of the user interface: Moving the user interface element towards the first portion of the user interface at a first rate based on determining that the user interface element is greater than a first threshold distance from the first portion of the user interface to which the attention of the user of the computer system is directed; And After moving the user interface element towards the first position, stopping the movement of the user interface element relative to the user interface based on determining that the user interface element is less than a second threshold distance from the first portion of the user interface to which the attention of the user of the computer system is directed, wherein the second threshold distance is less than the first threshold distance.

174. The method according to claim 173, wherein the position of the user interface element relative to the user interface corresponds to a corresponding state associated with the user interface.

175. The method according to any one of claims 173 to 174, the method further comprising, in response to detecting that the attention of the user of the computer system is directed to the first portion of the user interface, moving the user interface element towards the first portion of the user interface at a corresponding rate different from the first rate based on determining that the user interface element is less than the first threshold distance and greater than the second threshold distance from the first portion of the user interface to which the attention of the user of the computer system is directed.

176. The method according to claim 175, wherein: Based on determining that the user interface element is at a first distance from the first part of the user interface, the corresponding rate is a second rate, and Based on determining that the user interface element is at a second distance from the first part of the user interface that is different from the first distance, the corresponding rate is a third rate that is different from the second rate.

177. The method according to claim 176, wherein the second distance is less than the first distance, and the third rate is less than the second rate.

178. The method according to any one of claims 175 to 177, wherein the first rate is based on a first amount of change in the distance between the user interface element and the first part of the user interface, and the corresponding rate is based on a second amount of change in the distance between the user interface element and the first part of the user interface, wherein the first amount is less than the second amount.

179. The method according to claim 178, wherein the first rate is a fixed rate.

180. The method according to any one of claims 175 to 177, wherein moving the user interface element towards the first part of the user interface at the first rate includes: Moving the user interface element towards the first part of the user interface at the first rate until the user interface element is at a second threshold distance from the first part of the user interface; And In response to the user interface element reaching the second threshold distance from the first part of the user interface, stopping the movement of the user interface element.

181. The method according to any one of claims 175 to 177, wherein moving the user interface element towards the first part of the user interface at the corresponding rate includes: Moving the user interface element towards the first part of the user interface at the corresponding rate until the user interface element is at a second threshold distance from the first part of the user interface; And In response to the user interface element reaching the second threshold distance from the first part of the user interface, stopping the movement of the user interface element.

182. The method according to any one of claims 175 to 181, the method further comprising: When moving the user interface element, detecting that the attention of the user interface is directed to the user interface element; And In response to detecting that the user's attention is directed to the user interface element, stopping the movement of the user interface element.

183. The method according to any one of claims 173 to 181, wherein the user interface element is a user interface object for adjusting a value of the user interface.

184. The method according to claim 183, the method further comprising: Based on determining that the user interface element is at the first part of the user interface, displaying a fixation target for adjusting the value of the user interface; When displaying the fixation target for adjusting the value of the user interface, detect the user's attention directed at the fixation target via the one or more input devices; And In response to detecting the user's attention directed at the fixation target and based on determining that one or more criteria are met, adjust the value of the user interface based on the current position of the user interface element.

185. The method according to claim 184, wherein the fixation target is displayed within a second threshold distance of the user interface element.

186. The method according to any one of claims 173 to 185, wherein in response to detecting the user's attention of the computer system directed at the first part of the user interface: Based on determining that the user interface element has a first size, the second threshold distance is a first distance, and Based on determining that the user interface element has a second size different from the first size, the second threshold is a second distance different from the first distance.

187. The method according to any one of claims 173 to 186, wherein: Displaying the user interface element includes displaying the content associated with the user interface element at a position corresponding to the user interface element, Based on determining that the content has a first size, the second threshold distance is a first distance, and Based on determining that the content has a second size different from the first size, the second threshold is a second distance different from the first distance.

188. The method according to any one of claims 173 to 187, wherein moving the user interface element at the first rate based on determining that the user interface element is greater than the first threshold distance from the first part of the user interface element and moving the user interface element at the second rate based on determining that the user interface element is less than the second threshold distance from the first part of the user interface element is independent of whether the first part of the user interface has a first spatial relationship with respect to the user interface element or a second spatial relationship different from the first spatial relationship with respect to the user interface element.

189. The method according to any one of claims 173 to 188, wherein displaying the user interface element includes displaying a representation of a portion of a content item at a position different from the user interface element, wherein the representation of the portion of the content item corresponds to the current position of the user interface element in the user interface, and the method further includes: When displaying the user interface element and simultaneously displaying the representation of the portion of the content item at a position different from the user interface element, detect the user's attention directed at a second part of the user interface via the one or more input devices; And In response to detecting the user's attention directed at the second part of the user interface: Based on determining that the second portion of the user interface does not correspond to the representation of the content item, move the user interface element according to the distance between the second portion of the user interface and the user interface element; and Based on determining that the second portion of the user interface corresponds to the representation of the content item, abandon moving the user interface element, regardless of the distance between the second portion of the user interface and the user interface element.

190. The method according to claim 189, wherein the user interface includes a content playback scrubber element, and the representation of the content item corresponds to a preview of the content item at a playback position corresponding to the current position of the user interface element in the content playback scrubber element.

191. The method according to any one of claims 189 to 190, wherein the size of the user interface element is smaller than the size of the representation of the content item.

192. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, a user interface including a user interface element; when the user interface element is displayed, detecting that the attention of a user of the computer system is directed to a first portion of the user interface; and in response to detecting that the attention of the user of the computer system is directed to the first portion of the user interface: based on determining that the user interface element is greater than a first threshold distance from the first portion of the user interface to which the attention of the user of the computer system is directed, move the user interface element towards the first portion of the user interface at a first rate; and after moving the user interface element towards the first position, based on determining that the user interface element is less than a second threshold distance from the first portion of the user interface to which the attention of the user of the computer system is directed, stop the movement of the user interface element relative to the user interface, wherein the second threshold distance is less than the first threshold distance.

193. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, cause the computer system to perform a method including: displaying, via the display generation component, a user interface including a user interface element; When displaying the user interface element, detect that the attention of the user of the computer system is directed to a first portion of the user interface; and in response to detecting that the attention of the user of the computer system is directed to the first portion of the user interface: Determine that the user interface element is greater than a first threshold distance from the first part of the user interface to which the attention of the user of the computer system is directed, and move the user interface element towards the first part of the user interface at a first rate; And After moving the user interface element towards the first position, determine that the user interface element is less than a second threshold distance from the first part of the user interface to which the attention of the user of the computer system is directed, and stop the movement of the user interface element relative to the user interface, where the second threshold distance is less than the first threshold distance.

194. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; Means for: displaying a user interface including a user interface element via the display generation component; Means for: when the user interface element is displayed, detecting that the attention of the user of the computer system is directed to a first part of the user interface; And in response to detecting that the attention of the user of the computer system is directed to the first part of the user interface: Determine that the user interface element is greater than a first threshold distance from the first part of the user interface to which the attention of the user of the computer system is directed, and move the user interface element towards the first part of the user interface at a first rate; And After moving the user interface element towards the first position, determine that the user interface element is less than a second threshold distance from the first part of the user interface to which the attention of the user of the computer system is directed, and stop the movement of the user interface element relative to the user interface, where the second threshold distance is less than the first threshold distance.

195. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include instructions for performing any one of the methods according to claims 173 to 191.

196. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 173 to 191.

197. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And Means for performing any one of the methods according to claims 173 to 191.

198. A method, the method comprising: At a computer system in communication with a display generation component and one or more input devices: Display a user interface via the display generation component, the user interface including an attention indicator at a first location in the user interface, wherein the first location corresponds to the location of the attention of a user of the computer system pointing to the user interface; When displaying the user interface including the attention indicator, detect movement of the attention of the user via the one or more input devices; In response to detecting the movement of the attention of the user, move the attention indicator from the first location to a second location in the user interface that is different from the first location and corresponds to the movement of the attention of the user; And After displaying the attention indicator at the second location in the user interface and when the attention of the user continues to point to the second location, stop displaying the attention indicator in the user interface according to a determination that one or more criteria are met.

199. The method according to claim 198, the method further comprising: After displaying the attention indicator at the second location in the user interface and when the attention of the user continues to point to the second location, maintain the display of the attention indicator at the second location in the user interface according to a determination that the one or more criteria are not met.

200. The method according to any one of claims 198 to 199, wherein: Displaying the attention indicator at the first location in the user interface includes displaying the attention indicator at a location of the content of a first application; Displaying the attention indicator at the second location in the user interface includes displaying the attention indicator at a location of the content of a second application different from the first application.

201. The method according to any one of claims 198 to 200, wherein: The first location of the user interface corresponds to the content of a corresponding application; and No indication that the attention of the user of the computer system points to the corresponding application is provided to the corresponding application.

202. The method according to any one of claims 198 to 201, wherein the one or more criteria include criteria that are met when a content playback state of the content at the second location in the user interface is a first state.

203. The method according to any one of claims 198 to 202, wherein the one or more criteria include criteria that are met when an input for navigating in the content of the user interface is received at the second location.

204. The method according to any one of claims 198 to 203, wherein: The one or more criteria include criteria that are met when the content at the second location in the user interface is of a first type of content and not met when the content at the second location in the user interface is of a second type of content different from the first type of content.

205. The method according to any one of claims 198 to 204, wherein the one or more criteria include a criterion that is satisfied when the movement of the user's attention is less than a threshold movement.

206. The method according to any one of claims 198 to 205, wherein the one or more criteria include a criterion that is satisfied when the user's attention is directed to the second position in the user interface for a duration greater than a threshold duration.

207. The method according to any one of claims 198 to 206, the method further comprising: After stopping displaying the attention indicator in the user interface and when the attention indicator is not displayed in the user interface, detecting, via the one or more input devices, a movement of the user's attention in the user interface from the second position to a third position, wherein the second position corresponds to a first user interface object in the user interface; In response to detecting the movement of the user's attention: According to determining that the third position corresponds to a second user interface object different from the first user interface object in the user interface, displaying the attention indicator at the third position in the user interface via the display generation component; And According to determining that the third position corresponds to the first user interface object, refraining from displaying the attention indicator at the third position in the user interface.

208. The method according to any one of claims 206 to 207, wherein: According to determining that the second position corresponds to a first type of content, the threshold duration is a first duration, and According to determining that the second position corresponds to a second type of content different from the first type of content, the threshold duration is a second duration different from the first duration.

209. The method according to any one of claims 198 to 208, wherein stopping displaying the attention indicator in the user interface includes: Displaying a progressive animation transition between displaying the attention indicator at the second position in the user interface and stopping displaying the attention indicator in the user interface, the method further comprising: When the attention indicator is not displayed in the user interface, detecting that the user's attention is directed to a third position in the user interface; and In response to detecting that the user's attention is directed to the third position in the user interface and according to determining that one or more second criteria are satisfied, displaying a progressive animation transition between not displaying the attention indicator in the user interface and displaying the attention indicator at the third position in the user interface.

210. The method according to any one of claims 198 to 209, the method further comprising: When the attention indicator is displayed at the second position in the user interface, detecting a selection input via the one or more input devices, wherein the second position corresponds to an optional user interface object; And In response to detecting the selection input, initiate an operation associated with the selectable user interface object.

211. The method according to any one of claims 198 to 210, wherein the second position corresponds to a user interface object, and the attention indicator at the second position includes a visual indication that emphasizes the second position in the area of the object and is displayed in the area of the object.

212. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the following operations: Display a user interface via the display generation component, the user interface including an attention indicator at a first position in the user interface, wherein the first position corresponds to the position of the attention of a user of the computer system pointing to the user interface; When displaying the user interface including the attention indicator, detect the movement of the attention of the user via the one or more input devices; In response to detecting the movement of the attention of the user, move the attention indicator from the first position to a second position in the user interface that is different from the first position and corresponds to the movement of the attention of the user; And After displaying the attention indicator at the second position in the user interface and when the attention of the user continues to point to the second position, stop displaying the attention indicator in the user interface according to a determination that one or more criteria are met.

213. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: Display a user interface via the display generation component, the user interface including an attention indicator at a first position in the user interface, wherein the first position corresponds to the position of the attention of a user of the computer system pointing to the user interface; When displaying the user interface including the attention indicator, detect the movement of the attention of the user via the one or more input devices; In response to detecting the movement of the attention of the user, move the attention indicator from the first position to a second position in the user interface that is different from the first position and corresponds to the movement of the attention of the user; And After displaying the attention indicator at the second position in the user interface and when the attention of the user continues to point to the second position, stop displaying the attention indicator in the user interface according to a determination that one or more criteria are met.

214. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; Display a user interface via the display generation component, the user interface including an attention indicator at a first position in the user interface, wherein the first position corresponds to the position of the attention of a user of the computer system pointing to the user interface; When displaying the user interface including the attention indicator, detect the movement of the user's attention via the one or more input devices; In response to detecting the movement of the user's attention, move the attention indicator from the first position to a second position in the user interface that is different from the first position and corresponds to the movement of the user's attention; And After displaying the attention indicator at the second position in the user interface and when the user's attention continues to point to the second position, stop displaying the attention indicator in the user interface according to a determination that one or more criteria are met.

215. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods according to claims 198 to 211.

216. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, cause the computer system to perform any of the methods according to claims 198 to 211.

217. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And Means for performing any of the methods according to claims 198 to 211.

218. A method, the method comprising: At a computer system that communicates with a display generation component and one or more input devices: Display a user interface including a text entry field via the display generation component; when displaying the user interface including the text entry field, detect the attention of a user pointing to the text entry field via the one or more input devices; after detecting the attention of the user pointing to the text entry field, detect voice input; And In response to detecting the voice input: Based on determining that the user's attention directed to the text entry field meets one or more criteria of a first set, text is entered into the text entry field based on the speech input, wherein the one or more criteria of the first set require that the user's attention has been directed to a first portion of the text entry field to meet the one or more criteria of the first set; and Based on determining that the user's attention directed to the text entry field does not meet the one or more criteria of the first set, entering the text into the text entry field based on the speech input is abandoned.

219. The method according to claim 218, the method further comprising: In response to detecting the speech input: Based on determining that the user's attention directed to the text entry field has met the one or more criteria of the first set, providing an output indicating that subsequent speech input will be entered as text into the text entry field; and Based on determining that the user's attention directed to the text entry field does not meet the one or more criteria of the first set, abandoning providing the output indicating that the subsequent speech input will be entered as text into the text entry field.

220. The method according to any one of claims 218 to 219, the method further comprising: When detecting the user's attention directed to the text entry field: Based on determining that the user's attention directed to the text entry field does not meet the one or more criteria of the first set, displaying, via the display generation component, a visual indication of the dictation process of entering text in the text entry field based on subsequent speech input.

221. The method according to claim 220, wherein the visual indication of the dictation process includes an indication of progress towards meeting the one or more criteria of the first set.

222. The method according to any one of claims 220 to 221, the method further comprising: When displaying the visual indication of the dictation process, determining that the user's attention directed to the text entry field meets the one or more criteria of the first set; and Based on determining that the user's attention directed to the text entry field has met the one or more criteria of the first set, stopping the display of the visual indication of the dictation process.

223. The method according to any one of claims 220 to 222, wherein displaying the visual indication of the dictation process includes: Displaying an image corresponding to the dictation process, When displaying the image corresponding to the dictation process, determining that the user's attention is directed to the image corresponding to the dictation process, and Based on determining that the user's attention is directed to the image corresponding to the dictation process, displaying an animation in which the image corresponding to the dictation process changes to a geometric shape different from the image corresponding to the dictation process.

224. The method according to any one of claims 220 to 222, wherein displaying the visual indication of the dictation process includes: Displaying a geometric shape, When displaying the image corresponding to the oral process, determine that the user's attention is directed to the geometric shape, and In response to detecting that the user's attention is directed to the geometric shape, display an animation in which the geometric shape changes to an image different from the geometric shape corresponding to the oral process.

225. The method according to any one of claims 220 to 224, the method further comprising: When detecting the user's attention directed to the text entry field: Based on determining that the user's attention directed to the text entry field does not satisfy one or more of the first set of criteria: Based on determining that the text entry field satisfies a second set of criteria, display the visual indication of the oral process; And Based on determining that the text entry field does not satisfy the second set of criteria, abandon displaying the visual indication of the oral process.

226. The method according to any one of claims 218 to 225, wherein the computer system displays the visual indication of the oral process in the first portion of the text entry field.

227. The method according to any one of claims 218 to 226, the method further comprising: When detecting the user's attention directed to the text entry field: Based on determining that the user's attention directed to the text entry field does not satisfy one or more of the first set of criteria, display the first portion of the text entry field via the display generating component in a first appearance; And based on determining that the user's attention directed to the text entry field has satisfied one or more of the first set of criteria, display the first portion of the text entry field via the display generating component in a second appearance different from the first appearance.

228. The method according to any one of claims 218 to 227, the method further comprising: When detecting the user's attention directed to the text entry field: Based on determining that the user's attention directed to the text entry field does not satisfy one or more of the first set of criteria, display the text entry field via the display generating component in a first appearance; And Based on determining that the user's attention directed to the text entry field has satisfied one or more of the first set of criteria, display the text entry field via the display generating component in a second appearance different from the first appearance.

229. The method according to claim 228, wherein: Displaying the text entry field in the first appearance includes displaying the text entry field with a first background, and Displaying the text entry field in the second appearance includes displaying the text entry field with a second background different from the first background.

230. The method according to claim 229, wherein: Displaying the text entry field with the first background includes displaying the background in a color that does not change according to the audio level of the voice input, and Displaying the text entry field with the second background includes displaying the background in a color that changes according to the audio level of the voice input.

231. The method according to any one of claims 229 to 230, wherein: Displaying the text entry field with the first background includes displaying the background whether or not the voice input is detected but not animating the background, and displaying the text entry field with the second background includes displaying the animated background whether or not the voice input is detected.

232. The method according to any one of claims 228 to 231, wherein: Displaying the text entry field with the first appearance includes displaying an image associated with the text entry field in the text entry field, and Displaying the text entry field with the second appearance includes displaying the text entry field without the image associated with the text entry field.

233. The method according to any one of claims 228 to 232, wherein: Displaying the text entry field with the first appearance includes displaying an image associated with the text entry field at a corresponding position in the text entry field, and not displaying an image associated with the dictation process of entering text in the text entry field based on a subsequent voice input at the corresponding position in the text entry field, and Displaying the text entry field with the second appearance includes displaying the text entry field having the image associated with the dictation process at the corresponding position in the text entry field, and not displaying the image associated with the text entry field at the corresponding position in the text entry field.

234. The method according to any one of claims 228 to 233, wherein: Displaying the text entry field with the first appearance includes displaying a first text associated with the text entry field, and not displaying a second text associated with the dictation process of entering text in the text entry field based on a subsequent voice input in the text entry field, and Displaying the text entry field with the second appearance includes displaying the second text in the text entry field and not displaying the first text in the text entry field.

235. The method according to any one of claims 218 to 234, wherein entering the text into the text entry field based on the voice input includes displaying an animation of a portion of the text that changes over time according to a change in the audio level of the voice input.

236. The method according to any one of claims 218 to 235, the method further comprising: According to determining that the user's attention directed to the text entry field has satisfied one or more of the first set of criteria, displaying a caret in the text entry field via the display generation component, the caret including an animation of the caret that changes over time according to a change in the audio level of the voice input; And Based on determining that the user's attention directed to the text entry field does not meet the one or more criteria of the first set, abandon displaying the insertion marker in the text entry field, the insertion marker including an animation of the insertion marker that changes over time according to the change in the audio level of the voice input.

237. The method according to any one of claims 218 to 236, wherein the one or more criteria of the first set require that the user's attention be directed to the first portion of the text entry field for at least a predefined threshold duration.

238. The method according to any one of claims 218 to 237, the method further comprising: Based on determining that the user's attention directed to the text entry field does not meet the one or more criteria of the first set: Display, via the display generation component, the position in the text entry field to which the user's attention is directed in a first appearance; and display, via the display generation component, the position in the text entry field to which the user's attention is not directed in a second appearance.

239. The method according to any one of claims 218 to 238, the method further comprising: Based on determining that the user's attention directed to the text entry field meets the one or more criteria of the first set: Display, via the display generation component, the position in the text entry field to which the user's attention is directed in a first appearance; and display, via the display generation component, the position in the text entry field to which the user's attention is not directed in a second appearance.

240. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: Display, via the display generation component, a user interface including a text entry field; When displaying the user interface including the text entry field, detect the user's attention directed to the text entry field via the one or more input devices; after detecting the user's attention directed to the text entry field, detect a voice input; And In response to detecting the voice input: Based on determining that the user's attention directed to the text entry field meets the one or more criteria of the first set, enter text into the text entry field based on the voice input, wherein the one or more criteria of the first set require that the user's attention has been directed to the first portion of the text entry field to meet the one or more criteria of the first set; And Based on determining that the user's attention directed to the text entry field does not meet the one or more criteria of the first set, abandon entering the text into the text entry field based on the voice input.

241. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: Display, via the display generation component, a user interface including a text entry field; When displaying the user interface including the text entry field, detect the user's attention directed to the text entry field via the one or more input devices; after detecting the user's attention directed to the text entry field, detect voice input; And In response to detecting the voice input: Based on determining that the attention of the user directed to the text entry field satisfies a first set of one or more criteria, enter text into the text entry field based on the voice input, wherein the first set of one or more criteria requires that the attention of the user has been directed to a first portion of the text entry field to satisfy the first set of one or more criteria; And Based on determining that the attention of the user directed to the text entry field does not satisfy the first set of one or more criteria, refrain from entering the text into the text entry field based on the voice input.

242. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; Means for: displaying, via the display generation component, a user interface including a text entry field; Means for: when the user interface including the text entry field is displayed, detecting the attention of the user directed to the text entry field via the one or more input devices; after detecting the attention of the user directed to the text entry field, detecting a voice input; And Means for: in response to detecting the voice input: Based on determining that the attention of the user directed to the text entry field satisfies a first set of one or more criteria, enter text into the text entry field based on the voice input, wherein the first set of one or more criteria requires that the attention of the user has been directed to a first portion of the text entry field to satisfy the first set of one or more criteria; And Based on determining that the attention of the user directed to the text entry field does not satisfy the first set of one or more criteria, refrain from entering the text into the text entry field based on the voice input.

243. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods according to claims 218 to 239.

244. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform any of the methods according to claims 218 to 239.

245. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and means for performing any of the methods according to claims 218 to 239.

246. The method according to claim 34, wherein: the first optional user interface object is associated with a plurality of values and includes a plurality of positions associated with the plurality of values, and the method further comprises: when it is detected that the user's attention is directed to a corresponding position among the plurality of positions associated with a corresponding value among the plurality of values, displaying a corresponding visual indication of progress towards switching the first optional user interface object to the corresponding value in the user interface.

247. The method according to claim 246, the method further comprising: when the first optional user interface object is displayed: initiating a process of changing the current value of the first optional user interface object to the first value according to determining that the user's attention is directed to a first position corresponding to the first value of the first optional user interface object; and abandoning initiating the process of changing the current value of the first optional user interface object to the first value according to determining that the user's attention is not directed to the first position of the first optional user interface object.

248. The method according to claim 246, the method further comprising: when the first optional user interface object is displayed: initiating a process of changing the current value of the first optional user interface object to the first value according to determining that the user's attention directed to the first position of the first optional user interface object meets one or more second criteria; and abandoning initiating the process of changing the current value of the first optional user interface object to the first value according to determining that the user's attention directed to the first position of the first optional user interface object does not meet the one or more second criteria.

249. The method according to any one of claims 246 to 248, the method further comprising: when the first optional user interface object is displayed: initiating a process of changing the current value of the first optional user interface object to the first value according to determining that the user's attention is directed to a first position corresponding to the first value of the first optional user interface object; and initiating a process of changing the current value of the first optional user interface object to the second value according to determining that the user's attention is directed to a second position corresponding to a second value different from the first value of the first optional user interface object.

250. The method according to any one of claims 246 to 249, wherein the corresponding visual indication of the progress towards switching the first selectable user interface object to the corresponding value displayed in the user interface comprises: displaying the corresponding visual indication at a second position in the first selectable user interface object based on determining that the current value of the first selectable user interface object is a first value, wherein the second position corresponds to a second value different from the first value; and displaying the corresponding visual indication at a first position in the first selectable user interface object based on determining that the current value of the first selectable user interface object is the second value, wherein the first position corresponds to the first value.

251. The method according to any one of claims 246 to 250, wherein the first selectable user interface object comprises a selectable toggle affordance for changing the current value of the first selectable user interface object between a first value and a second value.

252. The method according to claim 251, wherein the corresponding visual indication of the progress towards switching the first selectable user interface object to the corresponding value displayed in the user interface comprises displaying the corresponding visual indication in the area of the toggle affordance.

253. A method, the method comprising: at a computer system in communication with a display generation component and one or more input devices: displaying, via the display generation component, a user interface including a value selection user interface object for selecting a corresponding value having a plurality of components associated with the value selection user interface object, wherein the value selection user interface object comprises: a first constituent element corresponding to a first set of value options for a first component of the corresponding value associated with the value selection user interface object; and a second constituent element corresponding to a second set of value options for a second component of the corresponding value associated with the value selection user interface object, wherein the second constituent element is different from the first constituent element and the second constituent element is displayed simultaneously with the first constituent element; detecting, via the one or more input devices, the attention of a user of the computer system directed to the first constituent element when the value selection user interface object including the first constituent element and the second constituent element is displayed; visually emphasizing the first constituent element relative to the second constituent element in response to detecting the attention of the user directed to the first constituent element when the first constituent element corresponds to a first value of the first constituent element; and when the attention of the user is directed to the first constituent element and when the first constituent element is visually emphasized relative to the second constituent element: navigating from a first value of the first constituent element to a second value of the first constituent element in a first set of value options of the first constituent part for the corresponding value based on the user's attention directed to the first constituent element; and after navigating in the first set of value options of the first constituent part for the corresponding value based on the user's attention directed to the first constituent element, initiating a process to update the corresponding value based on a selection of the second value of the first constituent element.

254. The method according to claim 253, the method further comprising: when navigating in the first set of value options based on the user's attention directed to the first constituent element, detecting via the one or more input devices that when a current navigation position in the first constituent element corresponds to the second value of the first constituent element different from the first value of the first constituent element, the user's attention directed to the first constituent element satisfies one or more first criteria; and wherein initiating the process of updating the first constituent part of the corresponding value to the second value of the first constituent element is in response to detecting that the user's attention directed to the first constituent element satisfies the one or more first criteria when the current navigation position in the first constituent element corresponds to the second value.

255. The method according to any one of claims 253 to 254, wherein the user's attention is based on the user's gaze.

256. The method according to any one of claims 253 to 255, wherein visually emphasizing the first constituent element relative to the second constituent element includes displaying a first user interface object for selecting a value for the first constituent part of the corresponding value, wherein the first user interface object is not displayed before the user's attention has been directed to the first constituent element.

257. The method according to any one of claims 253 to 256, wherein visually emphasizing the first constituent element relative to the second constituent element includes increasing the visual salience of the first constituent element relative to the second constituent element.

258. The method according to any one of claims 253 to 257, wherein the first set of value options of the first constituent part for the corresponding value corresponds to a date.

259. The method according to any one of claims 253 to 258, wherein the first set of value options of the first constituent part for the corresponding value corresponds to a time.

260. The method according to any one of claims 253 to 259, wherein navigating in the first set of value options based on the user's attention directed to the first constituent element includes navigating in the first set of value options based on the user's gaze directed to the first constituent element.

261. The method according to claim 260, wherein navigating in the first set of value options based on the user's gaze directed to the first constituent element includes: Based on determining that the position of the user's gaze in the first set of constituent elements is a first position, navigate in the first set of value options in a first manner; and Based on determining that the position of the user's gaze in the first set of constituent elements is a second position different from the first position, navigate in the first set of value options in a second manner different from the first manner.

262. The method according to any one of claims 260 to 261, wherein navigating in the first set of value options based on the user's gaze directed to the first set of constituent elements includes: Based on determining that the duration of the user's gaze in the first set of constituent elements is a first duration, navigate in the first set of value options in a first manner; and Based on determining that the duration of the user's gaze in the first set of constituent elements is a second duration different from the second duration, navigate in the first set of value options in a second manner different from the first manner.

263. The method according to any one of claims 260 to 261, wherein navigating in the first set of value options for the first component includes scrolling the first set of value options in a selection area of the first constituent element, and selecting a corresponding value in the first set of value options for the first component based on being positioned within the selection area of the first constituent element.

264. The method according to any one of claims 260 to 261, the method further comprising: When navigating in the first set of value options based on the user's gaze directed to the first set of constituent elements and when the first constituent element is navigated to a third value different from the first value of the first set of value options, detecting, via the one or more input devices, that the user's gaze deviates from the first constituent element; and In response to detecting that the user's gaze deviates from the first constituent element, based on determining that the user's gaze directed to the first constituent element did not meet one or more criteria before deviating from the first constituent element: Stop navigating in the first set of value options for the first component for the corresponding value; and Maintain the first value of the first component for the corresponding value.

265. The method according to any one of claims 253 to 264, wherein initiating the process of updating the corresponding value based on the selection of the second value of the first constituent element is performed in response to: Detecting, via the one or more input devices, that the user's attention directed to the first constituent element meets one or more first criteria when the current navigation position in the first constituent element corresponds to the second value of the first constituent element different from the first value of the first constituent element.

266. The method according to claim 265, wherein the process of updating the corresponding value based on the selection of the second value of the first constituent element includes: In response to detecting that the user's attention directed to the first component element meets the one or more first criteria, display a fixation target in the user interface that is selectable using the attention to update the first component of the corresponding value to the second value of the first component element.

267. The method according to claim 266, the method further comprising: When displaying the fixation target that is selectable to update the first component of the corresponding value to the second value of the first component element, detect the user's attention directed to the fixation target; And In response to detecting the user's attention directed to the fixation target, and based on determining that the user's attention directed to the fixation target meets one or more second criteria, update the first component of the corresponding value to the second value of the first component element, wherein the one or more second criteria include criteria that are met when the user's attention is directed to the fixation target for a duration exceeding a threshold time period.

268. The method according to any one of claims 265 to 267, wherein navigating among the first set of value options for the first component of the corresponding value based on the user's attention is performed without displaying the fixation target.

269. The method according to any one of claims 253 to 268, the method further comprising: When the user's attention is directed to the first component element, detect a first input directed to the first component element via the one or more input devices, wherein the first input includes an input from a first part of the user's body of the computer system, the input performing a corresponding air gesture, followed by movement of the first part of the user's body while in a corresponding posture; And In response to detecting the first input directed to the first component element, navigate among the first set of value options for the first component of the corresponding value based on the movement of the first part of the user's body.

270. The method according to claim 269, the method further comprising: When navigating among the first set of value options, detect that the first part of the user's body is no longer in the corresponding posture via the one or more input devices; And in response to detecting that the first part of the user's body is no longer in the corresponding posture: Based on determining that the current navigation position in the first component element corresponds to a third value, update the first component of the corresponding value to the third value of the first component element; And Based on determining that the current navigation position in the first component element corresponds to a fourth value different from the third value, update the first component of the corresponding value to the fourth value of the first component element.

271. The method according to any one of claims 253 to 270, the method further comprising: When displaying the value selection user interface object including the first component element and the second component element, detect the user's attention directed to the second component element via the one or more input devices; And In response to detecting the user's attention directed to the second component element when the second component element corresponds to a first value of the second component element, Visually emphasize the second component element relative to the first component element.

272. The method according to claim 271, the method further comprising: When the user's attention is directed to the second component element and when the second component element is visually emphasized relative to the first component element: Navigate from the first value of the second component element to a corresponding value of the second component element different from the first value of the second component element in the second set of value options of the second component part for the corresponding value based on the user's attention directed to the second component element; and After navigating in the second set of value options of the second component part for the corresponding value based on the user's attention directed to the second component element, initiate a process of updating the corresponding value of the second component element based on the selection of the corresponding value of the second component element.

273. The method according to any one of claims 271 to 272, wherein: According to the parameter for determining the attention directed to the second component element having a first value, the corresponding value of the second component element is a second value of the second component element; and According to the parameter for determining the attention directed to the second component element having a second value different from the first value of the parameter for determining the attention directed to the second component element, the corresponding value of the second component element is a third value of the second component element different from the second value of the second component element.

274. The method according to any one of claims 253 to 273, the method further comprising: When displaying the value selection user interface object including the first component element, the second component element and the third component element, detect the user's attention directed to the third component element via the one or more input devices; And In response to detecting the user's attention directed to the third component element when the third component element corresponds to a first value of the third component element, visually emphasize the third component element relative to the first component element and the second component element.

275. The method according to claim 274, the method further comprising: When the user's attention is directed to the third component element and when the third component element is visually emphasized relative to the first component element and the second component element: navigating, based on the user's attention directed to the third constituent element, from a first value of the third constituent element to a corresponding value of the third constituent element different from the first value among a third set of value options of the third constituent part for the corresponding value; and after navigating, based on the user's attention directed to the third constituent element, among the third set of value options of the third constituent part for the corresponding value, initiating a process to update the corresponding value based on a selection of the corresponding value of the third constituent element.

276. The method according to claim 275, wherein: based on determining that a parameter for the attention directed to the third constituent element has a first value, the corresponding value of the third constituent element is a second value of the third constituent element; and based on determining that the parameter for the attention directed to the third constituent element has a second value different from the first value of the parameter for the attention directed to the second constituent element, the corresponding value of the third constituent element is a third value of the third constituent element different from the second value of the third constituent element.

277. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, a user interface including a value selection user interface object for selecting a corresponding value having a plurality of constituent parts associated with the value selection user interface object, wherein the value selection user interface object includes: a first constituent element corresponding to a first set of value options of a first constituent part for the corresponding value associated with the value selection user interface object; and a second constituent element corresponding to a second set of value options of a second constituent part for the corresponding value associated with the value selection user interface object, wherein the second constituent element is different from the first constituent element and the second constituent element is displayed simultaneously with the first constituent element; when displaying the value selection user interface object including the first constituent element and the second constituent element, detecting, via the one or more input devices, the attention of a user of the computer system directed to the first constituent element; in response to detecting the attention of the user directed to the first constituent element when the first constituent element corresponds to a first value of the first constituent element, visually emphasizing the first constituent element relative to the second constituent element; and when the attention of the user is directed to the first constituent element and when the first constituent element is visually emphasized relative to the second constituent element: Navigate the user's attention directed to the first constituent element from a first value of the first constituent element to a second value of the first constituent element among a first set of value options of the first constituent part for the corresponding value; and After navigating the user's attention directed to the first constituent element among the first set of value options of the first constituent part for the corresponding value, initiate a process of updating the corresponding value based on a selection of the second value of the first constituent element.

278. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: Display, via the display generation component, a user interface including a value selection user interface object for selecting a corresponding value having a plurality of constituent parts associated with the value selection user interface object, wherein the value selection user interface object includes: A first constituent element corresponding to a first set of value options of a first constituent part for the corresponding value associated with the value selection user interface object; And A second constituent element corresponding to a second set of value options of a second constituent part for the corresponding value associated with the value selection user interface object, wherein the second constituent element is different from the first constituent element and the second constituent element is displayed simultaneously with the first constituent element; When displaying the value selection user interface object including the first constituent element and the second constituent element, detect, via the one or more input devices, the attention of a user of the computer system directed to the first constituent element; In response to detecting the attention of the user directed to the first constituent element when the first constituent element corresponds to a first value of the first constituent element, visually emphasize the first constituent element relative to the second constituent element; And When the attention of the user is directed to the first constituent element and when the first constituent element is visually emphasized relative to the second constituent element: Navigate the user's attention directed to the first constituent element from the first value of the first constituent element to the second value of the first constituent element among the first set of value options of the first constituent part for the corresponding value; And After navigating the user's attention directed to the first constituent element among the first set of value options of the first constituent part for the corresponding value, initiate a process of updating the corresponding value based on a selection of the second value of the first constituent element.

279. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; Apparatus for: displaying, via the display generation component, a user interface including a value selection user interface object for selecting a corresponding value having a plurality of components associated with the value selection user interface object, wherein the value selection user interface object includes: A first component element corresponding to a first set of value options for a first component of the corresponding value associated with the value selection user interface object; and A second component element corresponding to a second set of value options for a second component of the corresponding value associated with the value selection user interface object, wherein the second component element is different from the first component element and the second component element is displayed simultaneously with the first component element; Apparatus for: when displaying the value selection user interface object including the first component element and the second component element, detecting, via the one or more input devices, the attention of a user of the computer system directed to the first component element; Apparatus for: in response to detecting the attention of the user directed to the first component element when the first component element corresponds to a first value of the first component element, visually emphasizing the first component element relative to the second component element; and Apparatus for: when the attention of the user is directed to the first component element and when the first component element is visually emphasized relative to the second component element: Navigating, based on the attention of the user directed to the first component element, from the first value of the first component element to a second value of the first component element in the first set of value options for the first component of the corresponding value; and After navigating, based on the attention of the user directed to the first component element, in the first set of value options for the first component of the corresponding value, initiating a process of updating the corresponding value based on the selection of the second value of the first component element.

280. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: One or more processors; Memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods according to claims 253 to 276.

281. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform any of the methods according to claims 253 to 276.

282. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: One or more processors; Memory; and An apparatus for performing any of the methods according to claims 253 to 276.

283. A method, the method comprising: at a computer system in communication with a display generation component and one or more input devices: when a user interface including a first object is displayed via the display generation component, detecting, via the one or more input devices, a first part of a user of the computer system within a threshold distance of a location corresponding to the first object; and in response to detecting the first part of the user of the computer system within the threshold distance of the location corresponding to the first object: displaying the first object via the display generation component with a visual indication of an interaction between the first part of the user and the first object, wherein the visual indication of the interaction between the first part of the user and the first object has a first visual appearance, based on determining that the first object is a first type of object; and displaying the first object with the visual indication of the interaction between the first part of the user and the first object, wherein the visual indication of the interaction between the first part of the user and the first object has a second visual appearance different from the first visual appearance, based on determining that the first object is a second type of object different from the first type of object; wherein the visual indication changes appearance based on a change in position of the first part of the user relative to the location corresponding to the first object.

284. The method according to claim 283, wherein displaying the first object with the visual indication of the interaction between the first part of the user and the first object includes displaying the first object via the display generation component with a virtual highlighting effect.

285. The method according to any one of claims 283 to 284, wherein displaying the first object with the visual indication of the interaction between the first part of the user and the first object includes displaying the first object via the display generation component with a virtual glowing effect.

286. The method according to any one of claims 283 to 285, wherein: the first type of object is a selectable object; and the second type of object is a non-selectable object.

287. The method according to any one of claims 283 to 286, wherein: displaying the first object with the visual indication of the interaction between the first part of the user and the first object having the first visual appearance includes displaying the first object via the display generation component with the visual indication having a first visual saliency; and displaying the first object with the visual indication of the interaction between the first part of the user and the first object having the second visual appearance includes displaying the first object with the visual indication having a second visual saliency less than the first visual saliency.

288. The method according to any one of claims 283 to 287, wherein the first part of the user is the first finger of the user's hand, and in response to detecting the first part of the user within the threshold distance from the position corresponding to the first object, displaying the visual indication of the interaction between the first part of the user and the first object, and determining that the first part of the user is the first finger of the user's hand, the method further comprising: When, in response to detecting the first part of the user within the threshold distance from the position corresponding to the first object, displaying the visual indication of the interaction between the first part of the user and the first object, detecting, via the one or more input devices, a second part of the user within the threshold distance from the position corresponding to the first object; And In response to detecting the second part of the user within the threshold distance from the position corresponding to the first object: Abandoning the display of the visual indication of the interaction between the second part of the user and the first object based on determining that the second part of the user is a second finger of the user's hand other than the first finger. The method according to any one of claims 283 to 288, wherein, Before detecting the first part of the user within the threshold distance from the position corresponding to the first object, displaying the first object at a first distance relative to the user interface, and the threshold distance is a first threshold distance, the method further comprising: In response to detecting the first part of the user within the first threshold distance from the position corresponding to the first object: Based on determining that the first object is the first type of object, displaying the first object at a second distance greater than the first distance relative to the user interface and simultaneously displaying the visual indication of the interaction between the first part of the user and the first object.

290. The method according to claim 289, the method further comprising: When, in response to detecting the first part of the user within the first threshold distance from the position corresponding to the first object, displaying the first object at the second distance relative to the user interface based on determining that the first object is the first type of object, detecting, via the one or more input devices, the first part of the user within a second threshold distance less than the first threshold distance from the position corresponding to the first object; And In response to detecting the first part of the user within the second threshold distance from the position corresponding to the first object: Displaying the first object at a third distance less than the second distance relative to the user interface and simultaneously displaying the visual indication of the interaction between the first part of the user and the first object.

291. The method according to claim 290, the method further comprising: In response to detecting the first portion of the user within the second threshold distance of the location corresponding to the first object: Performing an operation corresponding to the selection of the first object in the user interface.

292. The method according to any one of claims 289 to 291, the method further comprising: In response to detecting the first portion of the user within the first threshold distance of the location corresponding to the first object: Based on determining that the first object is a third type of object different from the first type of object, maintaining the display of the first object at the first distance relative to the user interface via the display generating component and simultaneously displaying the first object with the visual indication of the interaction between the first portion of the user and the first object.

293. The method according to any one of claims 283 to 292, the method further comprising: After displaying the first object with the visual indication of the interaction between the first portion of the user and the first object in response to detecting the first portion of the user within the threshold distance of the location corresponding to the first object, detecting, via the one or more input devices, an input including the user's gaze directed at the first object in the user interface when the first portion of the user is at a distance greater than the threshold distance from the location corresponding to the first object; And When the input is detected, based on determining that one or more first criteria are met, displaying, via the display generating component, a visual feedback indicating the position of the user's gaze relative to the first object, wherein the visual feedback has a third visual appearance different from the first visual appearance and the second visual appearance, and the one or more first criteria include criteria satisfied when the input includes a corresponding type of input from the first portion of the user.

294. The method according to any one of claims 283 to 293, the method further comprising: When displaying the first object with the visual indication of the interaction between the first portion of the user and the first object in response to detecting the first portion of the user within the threshold distance of the location corresponding to the first object, detecting, via the one or more input devices, the movement of the first portion of the user relative to the first object; And In response to detecting the movement of the first portion of the user relative to the first object: Based on the movement of the first portion of the user relative to the first object, changing, via the display generating component, the visual appearance of the visual indication displayed together with the first object.

295. The method according to claim 294, wherein changing the visual appearance of the visual indication displayed together with the first object includes changing the virtual illumination effect of the visual indication.

296. The method according to any one of claims 294 to 295, wherein changing the appearance of the visual indication displayed with the first object includes: displaying the visual indication with a first luminance level based on determining that a first part of the user is at a first distance from the position corresponding to the first object; and displaying the visual indication with a second luminance level greater than the first luminance level based on determining that the first part of the user is at a second distance less than the first distance from the position corresponding to the first object.

297. The method according to any one of claims 294 to 296, wherein changing the appearance of the visual indication displayed with the first object includes: displaying the visual indication with a first size based on determining that a first part of the user is at a first distance from the position corresponding to the first object; and displaying the visual indication with a second size less than the first size based on determining that the first part of the user is at a second distance less than the first distance from the position corresponding to the first object.

298. The method according to any one of claims 283 to 297, wherein: the user interface includes a virtual keyboard; and the first object corresponds to a first key among a plurality of keys of the virtual keyboard.

299. The method according to any one of claims 283 to 298, the method further comprising: when the first object is displayed with the visual indication of the interaction between the first part of the user and the first object, wherein the visual indication has the first visual appearance, detecting, via the one or more input devices, an input corresponding to a selection of the first object provided by the first part of the user based on determining that the first object is a first type of object in response to detecting that the first part of the user is within the threshold distance of the position corresponding to the first object; and in response to detecting the input provided by the first part of the user: performing an operation corresponding to the selection of the first object according to the input; and outputting an audio indicating the selection of the first object.

300. The method according to claim 299, wherein outputting the audio indicating the selection of the first object includes outputting the audio with a corresponding audio characteristic having a first value, the method further comprising: when the first object is displayed with the visual indication of the interaction between the first part of the user and the first object, wherein the visual indication has the second visual appearance, detecting, via the one or more input devices, an input corresponding to a selection of the first object provided by the first part of the user based on determining that the first object is a second type of object in response to detecting that the first part of the user is within the threshold distance of the position corresponding to the first object; and in response to detecting the input provided by the first part of the user: Output audio indicating contact with the first object with the corresponding audio characteristic having a second value different from the first value.

301. The method according to any one of claims 283 to 300, wherein displaying the first object with the visual indication of the interaction between the first part of the user and the first object includes: Based on determining that the boundary of the visual indication extends beyond the boundary of the first object in the user interface, while performing the following operations: Display the first object with the visual indication having the first luminance amount via the display generation component; And Display one or more parts of the user interface surrounding the first object with the visual indication having a second luminance amount less than the first luminance amount.

302. The method according to any one of claims 283 to 301, wherein, Before detecting the first part of the user within the threshold distance at the location corresponding to the first object, display the first object at a first distance relative to the user interface. The method further includes: When displaying the first object with the visual indication of the interaction between the first part of the user and the first object, wherein the visual indication has the first visual appearance, determine that the first object is the first type of object according to detecting the first part of the user within the threshold distance at the location corresponding to the first object, and detect, via the one or more input devices, an input corresponding to the selection of the first object provided by the first part of the user; and In response to detecting the input provided by the first part of the user: Display the first object at a second distance relative to the user interface that is less than the first distance via the display generation component.

303. The method according to any one of claims 283 to 302, wherein displaying the visual indication having the first visual appearance includes displaying the visual indication having a first luminance amount based on the distance between the first part of the user and the location corresponding to the first object. The method further includes: When displaying the first object with the visual indication of the interaction between the first part of the user and the first object, wherein the visual indication has the first luminance amount, determine that the first object is the first type of object according to detecting the first part of the user within the threshold distance at the location corresponding to the first object, and detect, via the one or more input devices, an input corresponding to the selection of the first object provided by the first part of the user; And In response to detecting the input provided by the first part of the user: Display the first object with the visual indication of the interaction between the first part of the user and the first object via the display generation component, wherein the visual indication has a second luminance amount that is at least a corresponding amount greater than the first luminance amount and does not vary based on the distance between the first part of the user and the first object. The method according to any one of claims 283 to 303, wherein The first object has a first size in the user interface, and the visual indication of the interaction between the first part of the user and the first object has a second size. The method further includes: When the first object is displayed with the visual indication of the interaction between the first part of the user and the first object, where the visual indication has the second size, determining that the first object is the first type of object based on detecting the first part of the user within the threshold distance at a position corresponding to the first object, detecting, via the one or more input devices, movement of the first part of the user relative to the user interface; and In response to detecting the movement of the first part of the user, based on determining that the first part of the user of the computer system is within the threshold distance at a position corresponding to a second object different from the first object, where the second object has a third size different from the first size in the user interface: Displaying the second object with a visual indication of the interaction between the first part of the user and the second object via the display generating component, where the visual indication of the interaction between the first part of the user and the second object has the second size.

305. A computer system communicating with a display generating component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: When a user interface including a first object is displayed via the display generating component, detecting, via the one or more input devices, a first part of a user of the computer system within a threshold distance at a position corresponding to the first object; And In response to detecting the first part of the user of the computer system within the threshold distance at the position corresponding to the first object: Based on determining that the first object is a first type of object, displaying the first object with a visual indication of the interaction between the first part of the user and the first object via the display generating component, where the visual indication of the interaction between the first part of the user and the first object has a first visual appearance; And Based on determining that the first object is a second type of object different from the first type of object, displaying the first object with the visual indication of the interaction between the first part of the user and the first object, where the visual indication of the interaction between the first part of the user and the first object has a second visual appearance different from the first visual appearance; Where the visual indication changes appearance based on a change in position of the first part of the user relative to a position corresponding to the first object.

306. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: When a user interface including a first object is displayed via the display generation component, a first portion of a user of the computer system within a threshold distance of a location corresponding to the first object is detected via the one or more input devices; And In response to detecting the first portion of the user of the computer system within the threshold distance of the location corresponding to the first object: Based on determining that the first object is a first type of object, display the first object via the display generation component with a visual indication of an interaction between the first portion of the user and the first object, wherein the visual indication of the interaction between the first portion of the user and the first object has a first visual appearance; And Based on determining that the first object is a second type of object different from the first type of object, display the first object with a visual indication of an interaction between the first portion of the user and the first object, wherein the visual indication of the interaction between the first portion of the user and the first object has a second visual appearance different from the first visual appearance; Wherein the visual indication changes appearance based on a change in position of the first portion of the user relative to the position corresponding to the first object.

307. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; Means for: when displaying a user interface including a first object via the display generation component, detecting a first portion of a user of the computer system within a threshold distance of a location corresponding to the first object via the one or more input devices; And Means for: in response to detecting the first portion of the user of the computer system within the threshold distance of the location corresponding to the first object: Based on determining that the first object is a first type of object, display the first object via the display generation component with a visual indication of an interaction between the first portion of the user and the first object, wherein the visual indication of the interaction between the first portion of the user and the first object has a first visual appearance; And Based on determining that the first object is a second type of object different from the first type of object, display the first object with a visual indication of an interaction between the first portion of the user and the first object, wherein the visual indication of the interaction between the first portion of the user and the first object has a second visual appearance different from the first visual appearance; Wherein the visual indication changes appearance based on a change in position of the first portion of the user relative to the position corresponding to the first object.

308. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any one of the methods according to claims 283 to 304.

309. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 283 to 304.

310. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: One or more processors; Memory; And Means for performing any one of the methods according to claims 283 to 304.

311. A method, the method comprising: At a computer system in communication with a display generation component and one or more input devices: When a user interface including an object is displayed via the display generation component, detecting a first portion of a user of the computer system within a threshold distance of a position corresponding to the object via the one or more input devices; And In response to detecting the first portion of the user within the threshold distance of the position corresponding to the object: Displaying the object with a visual indication of an interaction between the first portion of the user and the object via the display generation component based on determining that the first portion of the user is a first finger of the user's hand; And Based on determining that the first portion of the user is a second finger of the user's hand other than the first finger, forgoing displaying the object with the visual indication of the interaction between the first portion of the user and the object.

312. The method according to claim 311, wherein the visual indication of the first portion of the user changes based on a change in position of the first finger of the user's hand relative to the position corresponding to the object.

313. The method according to any one of claims 311 to 312, the method further comprising: When, in response to detecting the first portion of the user within the threshold distance of the position corresponding to the object, forgoing displaying the object with the visual indication of the interaction between the first portion of the user and the object based on determining that the first portion of the user is a second finger of the user's hand other than the first finger, detecting a corresponding input provided by the second finger of the user's hand pointing to the object via the one or more input devices; And In response to detecting the corresponding input, forgoing performing an operation associated with the object in the user interface.

314. The method according to any one of claims 311 to 312, the method further comprising: When, in response to detecting the first portion of the user within the threshold distance of the location corresponding to the object, the visual indication of the interaction between the first portion of the user and the object is displayed by giving up the interaction between the first portion of the user and the object based on determining that the first portion of the user is the second finger of the user's hand other than the first finger, detecting, via the one or more input devices, a corresponding input provided by the second finger of the user's hand pointing to the object; and In response to detecting the corresponding input, performing an operation associated with the object in the user interface.

315. The method according to any one of claims 311 to 314, wherein the visual indication of the interaction between the first portion of the user and the object is displayed at a first position relative to the object, and the method further comprises: When, in response to detecting the first portion of the user within the threshold distance of the location corresponding to the object, the visual indication of the interaction between the first portion of the user and the object is displayed at the first position relative to the object based on determining that the first portion of the user is the first finger of the user's hand, detecting, via the one or more input devices, a movement of the first finger of the user's hand relative to the object, wherein the movement of the first finger of the hand comprises a lateral movement of the first finger of the hand relative to the object; and In response to detecting the movement of the first finger: Based on determining that the first finger of the user's hand is within the threshold distance of the location corresponding to the object, moving, via the display generating component, the visual indication of the interaction between the first portion of the user and the object to a second position different from the first position relative to the object according to the movement of the first finger.

316. The method according to any one of claims 311 to 315, the method further comprises: When, in response to detecting the first portion of the user within the threshold distance of the location corresponding to the object, the object is displayed with the visual indication of the interaction between the first portion of the user and the object based on determining that the first portion of the user is the first finger of the user's hand, and wherein the visual indication of the interaction has the first visual appearance, detecting, via the one or more input devices, a movement of the first finger of the user's hand relative to the object in the user interface; And In response to detecting the movement of the first finger: Based on determining that when the first finger is within the threshold distance of the position corresponding to the object, the movement of the first finger decreases the distance between the first finger and the position corresponding to the object, display the object with the visual indication of the interaction between the first finger and the object via the display generating component, wherein the visual indication has a second visual appearance different from the first visual appearance; And Based on determining that when the first finger is within the threshold distance of the position corresponding to the object, the movement of the first finger increases the distance between the first finger and the position corresponding to the object, display the object with the visual indication of the interaction between the first finger and the object via the display generating component, wherein the visual indication has a third visual appearance different from the first visual appearance and the second visual appearance.

317. The method according to any one of claims 311 to 316, the method further comprising: When, in response to detecting the first portion of the user within the threshold distance of the position corresponding to the object, based on determining that the first portion of the user is the first finger of the user's hand, the object is displayed with the visual indication of the interaction between the first portion of the user and the object: Detect a second portion of the user of the computer system within the threshold distance of the position corresponding to the corresponding object in the user interface via the one or more input devices; And In response to detecting the second portion of the user within the threshold distance of the position corresponding to the corresponding object, and based on determining that the second portion of the user is the third finger of the user's hand, display the corresponding object with the visual indication of the interaction between the second portion of the user and the corresponding object via the display generating component.

318. The method according to claim 317, wherein: The first portion of the user is a finger from the first hand of the user; and The second portion of the user is a finger from the second hand of the user different from the first hand.

319. The method according to any one of claims 317 to 318, wherein the corresponding object is the object, the method further comprising: When, in response to detecting the first portion and the second portion of the user within the threshold distance of the position corresponding to the object, based on determining that the first portion of the user is the first finger of the user's hand, the object is displayed with the visual indication of the interaction between the first portion of the user and the object, and based on determining that the second portion of the user is the third finger of the user's hand, the object is displayed with the visual indication of the interaction between the second portion of the user and the object: Detect movement of the user's first finger relative to the object and movement of the user's third finger relative to the object via the one or more input devices; And In response to detecting the movement of the user's first finger relative to the object and the movement of the user's third finger relative to the object: Based on determining that the first finger of the user's hand is within the threshold distance corresponding to the position of the object, move the visual indication of the interaction between the first part of the user and the object via the display generation component based on the movement of the first finger, independent of the movement of the third finger; And Based on determining that the third finger of the user's hand is within the threshold distance corresponding to the position of the object, move the visual indication of the interaction between the second part of the user and the object based on the movement of the third finger, independent of the movement of the first finger.

320. The method according to any one of claims 311 to 319, wherein in response to detecting the first part of the user within the threshold distance corresponding to the position of the object, displaying the object with the visual indication of the interaction between the first part of the user and the object based on determining that the first part of the user is the first finger of the user's hand includes: Based on determining that the object is a first type of object, display the visual indication of the interaction between the first part of the user and the first object with a first visual appearance via the display generation component; And Based on determining that the object is a second type of object different from the first type of object, display the visual indication of the interaction between the first part of the user and the first object with a second visual appearance different from the first visual appearance.

321. The method according to any one of claims 311 to 320, wherein: The user interface includes a virtual keyboard; and The object corresponds to a first key among a plurality of keys of the virtual keyboard.

322. The method according to claim 321, the method further comprising: When, in response to detecting the first part of the user within the threshold distance corresponding to the position of the object, displaying the first key with the visual indication of the interaction between the first part of the user and the first key based on determining that the first part of the user is the first finger of the user's hand: Detect a second part of the user of the computer system within the threshold distance corresponding to the position of a second key of the virtual keyboard via the one or more input devices; In response to detecting the second portion of the user within the threshold distance of the position corresponding to the second key, and based on determining that the second portion of the user is the third finger of the user's hand, display the second key via the display generating component with a visual indication of the interaction between the second portion of the user and the second key; When, in response to detecting the second portion of the user within the threshold distance of the position corresponding to the second key, the first key is displayed with the visual indication of the interaction between the first portion of the user and the first key and simultaneously the second key is displayed with the visual indication of the interaction between the second portion of the user and the second key based on determining that the second portion of the user is the third finger of the user's hand: Detect, via the one or more input devices, a first input provided by the first finger of the user's hand pointing to the first key and a second input provided by the third finger of the user's hand pointing to the second key; And In response to detecting the first input and the second input, perform operations associated with the first key and the second key of the virtual keyboard.

323. A computer system communicating with a display generating component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: When a user interface including an object is displayed via the display generating component, detect, via the one or more input devices, a first portion of a user of the computer system within a threshold distance of a position corresponding to the object; And In response to detecting the first portion of the user within the threshold distance of the position corresponding to the object: Based on determining that the first portion of the user is the first finger of the user's hand, display the object via the display generating component with a visual indication of the interaction between the first portion of the user and the object; And Based on determining that the first portion of the user is a second finger of the user's hand other than the first finger, refrain from displaying the object with the visual indication of the interaction between the first portion of the user and the object.

324. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generating component and one or more input devices, cause the computer system to perform a method including the following operations: When a user interface including an object is displayed via the display generation component, a first portion of a user of the computer system within a threshold distance of a location corresponding to the object is detected via the one or more input devices; And In response to detecting the first portion of the user within the threshold distance of the position corresponding to the object: Based on determining that the first part of the user is the first finger of the user's hand, display, via the display generating component, a visual indication of the interaction between the first part of the user and the object; And Based on determining that the first part of the user is a second finger of the user's hand other than the first finger, refrain from displaying, via the display generating component, the visual indication of the interaction between the first part of the user and the object.

325. A computer system communicating with a display generating component and one or more input devices, the computer system comprising: One or more processors; A memory; Means for: when displaying a user interface including an object via the display generating component, detecting, via the one or more input devices, a first part of a user of the computer system within a threshold distance of a location corresponding to the object; And Means for: in response to detecting the first part of the user within the threshold distance of the location corresponding to the object: Based on determining that the first part of the user is the first finger of the user's hand, display, via the display generating component, a visual indication of the interaction between the first part of the user and the object; And Based on determining that the first part of the user is a second finger of the user's hand other than the first finger, refrain from displaying, via the display generating component, the visual indication of the interaction between the first part of the user and the object.

326. A computer system communicating with a display generating component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods according to claims 311 to 322.

327. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generating component and one or more input devices, cause the computer system to perform any of the methods according to claims 311 to 322.

328. A computer system communicating with a display generating component and one or more input devices, the computer system comprising: One or more processors; A memory; And Means for performing any of the methods according to claims 311 to 322.

329. A method, the method comprising: At a computer system communicating with a display generating component and one or more input devices: Display, via the display generating component, a user interface including a virtual object; When displaying the user interface, detect, via the one or more input devices, a first user input including an attention to a first part pointing to the virtual object; And When detecting the first user input: Based on determining that one or more first criteria are met, display a first visual feedback indicating the position of the user's attention in the first portion of the virtual object via the display generating component, the one or more first criteria including criteria that are met when the first user input includes a corresponding type of input from a first portion of the user at a position corresponding to the virtual object; and Based on determining that one or more second criteria are met, display a second visual feedback different from the first visual feedback in the first portion of the virtual object without displaying the first visual feedback, the one or more second criteria including criteria that are met when the first user input includes the user's attention without the corresponding type of input from the user.

330. The method according to claim 329, wherein displaying the second visual feedback via the display generating component includes displaying the second visual feedback having a shape obtained by masking a first shape corresponding to the user's attention with a second shape of the virtual object, and displaying the first visual feedback via the display generating component includes displaying the first visual feedback having a shape that is not obtained by masking the first shape with the second shape.

331. The method according to any one of claims 329 to 330, the method further comprising: When the one or more first criteria or the one or more second criteria are met, detect a change in the position of the first portion of the user relative to the virtual object via the one or more input devices; and In response to detecting the change in the position of the first portion of the user relative to the virtual object: Based on determining that the one or more first criteria were met, change the visual appearance of the first visual feedback by a first amount; and based on determining that the one or more second criteria were met, refrain from changing the visual appearance of the second visual feedback by the first amount.

332. The method according to claim 331, the method further comprising: In response to detecting the change in the position of the first portion of the user relative to the virtual object, and based on determining that the one or more second criteria were met, refrain from changing the visual appearance of the second visual feedback based on the change in the position of the first portion of the user relative to the virtual object.

333. The method according to any one of claims 331 to 332, wherein changing the visual appearance of the first visual feedback includes changing the brightness level of the first visual feedback.

334. The method according to any one of claims 331 to 333, wherein changing the visual appearance of the first visual feedback includes changing the size of the first visual feedback.

335. The method according to any one of claims 331 to 334, wherein changing the visual appearance of the first visual feedback includes changing the amount of blur of the first visual feedback.

336. The method according to any one of claims 329 to 335, wherein the one or more second criteria include criteria that are satisfied when the virtual object is an optional object and not satisfied when the virtual object is not an optional object.

337. The method according to any one of claims 329 to 336, wherein the one or more second criteria include criteria that are satisfied when the first part of the user is in a ready state.

338. The method according to any one of claims 329 to 337, the method further comprising: When displaying the second visual feedback in the first part of the virtual object in response to detecting the first user input: Detecting a second user input corresponding to indirect input from the first part of the user via the one or more input devices; and Based on the detected second user input, interacting with the virtual object in response to detecting the second user input.

339. The method according to any one of claims 329 to 338, the method further comprising: When displaying the first visual feedback in the first part of the virtual object in response to detecting the first user input: Detecting a second user input corresponding to direct input from the first part of the user via the one or more input devices; and Based on the detected second user input, interacting with the virtual object in response to detecting the second user input.

340. The method according to any one of claims 329 to 339, the method further comprising: When the one or more first criteria or the one or more second criteria are satisfied, detecting a second input corresponding to a selection of the virtual object via the one or more input devices, wherein the virtual object is an optional button; And In response to detecting the second input, selecting the optional button, including: According to determining that the one or more first criteria were satisfied when the second input was detected, moving the virtual object away from the user's viewpoint according to the second input; And According to determining that the one or more second criteria were satisfied when the second input was detected, foregoing moving the virtual object away from the user's viewpoint.

341. The method according to any one of claims 329 to 340, the method further comprising: When the one or more first criteria or the one or more second criteria are satisfied, detecting a second input corresponding to a selection of the virtual object via the one or more input devices, wherein the virtual object is an optional button; And In response to detecting the second input: Selecting the optional button; According to determining that the one or more first criteria were satisfied, displaying a visual indication of at least a part of the perimeter around the virtual object indicating that the virtual object has been selected; And According to determining that the one or more second criteria were satisfied, foregoing displaying the visual indication of the part of the perimeter around the virtual object.

342. The method according to any one of claims 329 to 341, the method further comprising: When displaying the second visual feedback at the first portion of the virtual object, detect a second user input corresponding to movement of the first portion of the user toward the virtual object; and in response to detecting the second user input and based on determining that the first portion of the user has moved within a threshold distance of the virtual object, stop displaying the second visual feedback and display the first visual feedback in a corresponding portion of the virtual object.

343. The method according to any one of claims 329 to 342, wherein: displaying the first visual feedback includes: displaying the first visual feedback having a first visual appearance based on determining that the virtual object has a first size; and displaying the first visual feedback having the first visual appearance based on determining that the virtual object has a second size; and displaying the second visual feedback includes: displaying the second visual feedback having a second visual appearance based on determining that the virtual object has the first size; and displaying the second visual feedback having a third visual appearance different from the second visual appearance based on determining that the virtual object has the second size.

344. The method according to any one of claims 329 to 343, the method further comprising: when displaying a corresponding virtual object having a first size relative to the three-dimensional environment in the three-dimensional environment, detecting, via the one or more input devices, a second user input including attention directed to a first portion of the corresponding virtual object; and in response to detecting the second user input, displaying, based on determining that one or more third criteria are met, the corresponding virtual object having a second size greater than the first size relative to the three-dimensional environment.

345. The method according to claim 344, wherein the second user input includes the user's gaze directed to the corresponding virtual object without input from the user other than the user's gaze, and the one or more third criteria include criteria that are met when the second user input includes the user's gaze directed to the corresponding virtual object without input from the user other than the gaze.

346. The method according to claim 345, wherein the one or more third criteria include criteria that are met when the user's gaze is directed to the corresponding virtual object, and criteria that are met when the second user input includes the corresponding type of input from the first portion of the user at a location corresponding to the corresponding virtual object, the method further comprising: when displaying the corresponding virtual object having the second size relative to the three-dimensional environment, detecting, via the one or more input devices, that the user's gaze is no longer directed to the corresponding virtual object and that the first portion of the user is no longer providing the corresponding type of input at the location corresponding to the corresponding virtual object; and In response to detecting that the user's gaze no longer points to the corresponding virtual object and the first part of the user no longer provides the corresponding type of input at the position corresponding to the corresponding virtual object, display in the three-dimensional environment the corresponding virtual object having the first size relative to the three-dimensional environment.

347. The method according to claim 346, the method further comprising: When displaying the corresponding virtual object having the second size relative to the three-dimensional environment, detecting, via the one or more input devices, that the first part of the user no longer provides the corresponding type of input at the position corresponding to the corresponding virtual object but the user's gaze points to the corresponding virtual object; and In response to detecting that the first part of the user no longer provides the corresponding type of input at the position corresponding to the corresponding virtual object but the user's gaze points to the corresponding virtual object, maintain in the three-dimensional environment the display of the corresponding virtual object having the second size relative to the three-dimensional environment.

348. The method according to claim 346, the method further comprising: When displaying the corresponding virtual object having the second size relative to the three-dimensional environment, detecting, via the one or more input devices, that the user's gaze no longer points to the corresponding virtual object but the first part of the user provides the corresponding type of input at the position corresponding to the corresponding virtual object; and In response to detecting that the user's gaze no longer points to the corresponding virtual object but the first part of the user provides the corresponding type of input at the position corresponding to the corresponding virtual object, maintain in the three-dimensional environment the display of the corresponding virtual object having the second size relative to the three-dimensional environment.

349. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: Display, via the display generation component, a user interface including virtual objects; When displaying the user interface, detect, via the one or more input devices, a first user input including the attention of a first part pointing to the virtual object; And When detecting the first user input: According to determining that one or more first criteria are met, display, via the display generation component, in the first part of the virtual object a first visual feedback indicating the position of the user's attention, the one or more first criteria including criteria that are met when the first user input includes a corresponding type of input from the first part of the user at a position corresponding to the virtual object; And Based on determining that one or more second criteria are met, display second visual feedback different from the first visual feedback in the first portion of the virtual object without displaying the first visual feedback, the one or more second criteria including criteria that are met when the first user input includes the user's attention without the corresponding type of input from the user.

350. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: Display, via the display generation component, a user interface including a virtual object; When displaying the user interface, detect a first user input via the one or more input devices that includes attention directed to a first portion of the virtual object; And When the first user input is detected: Based on determining that one or more first criteria are met, display, via the display generation component, first visual feedback indicating the location of the user's attention in the first portion of the virtual object, the one or more first criteria including criteria that are met when the first user input includes the corresponding type of input from a first portion of the user at a location corresponding to the virtual object; And Based on determining that one or more second criteria are met, display second visual feedback different from the first visual feedback in the first portion of the virtual object without displaying the first visual feedback, the one or more second criteria including criteria that are met when the first user input includes the user's attention without the corresponding type of input from the user.

351. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; Means for: displaying, via the display generation component, a user interface including a virtual object; Means for: when the user interface is displayed, detecting, via the one or more input devices, a first user input including attention directed to a first portion of the virtual object; And Means for: when the first user input is detected: Based on determining that one or more first criteria are met, display, via the display generation component, first visual feedback indicating the location of the user's attention in the first portion of the virtual object, the one or more first criteria including criteria that are met when the first user input includes the corresponding type of input from a first portion of the user at a location corresponding to the virtual object; And Based on determining that one or more second criteria are met, display second visual feedback different from the first visual feedback in the first portion of the virtual object without displaying the first visual feedback, the one or more second criteria including criteria that are met when the first user input includes the user's attention without the corresponding type of input from the user.

352. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any of the methods according to claims 329 to 348.

353. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform any of the methods according to claims 329 to 348.

354. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And Means for performing any of the methods according to claims 329 to 348.

355. A method, the method comprising: At a computer system in communication with a display generation component and one or more input devices: Displaying, via the display generation component, a user interface including a virtual object; When the user interface is displayed, detecting, via the one or more input devices, a first user input that points to the virtual object and includes movement of a first portion of a user of the computer system in a corresponding direction toward the virtual object relative to the user's viewing point; And When the first user input is detected: In response to detecting a first portion of the first user input, moving the virtual object away from the user's viewing point according to the movement of the first portion of the user; And In response to detecting a second portion of the first user input after the first portion of the first user input and according to determining that the second portion of the first user input includes movement of the first portion of the user in the corresponding direction away from the user's viewing point by more than a threshold amount, moving the virtual object toward the user's viewing point.

356. The method according to claim 355, the method further comprising: In response to detecting a second portion of the first user input after the first portion of the first user input and according to determining that the second portion of the first user input includes movement of the first portion of the user in the corresponding direction away from the user's viewing point by less than the threshold amount, moving the virtual object away from the user's viewing point according to the second portion of the first user input.

357. The method according to any one of claims 355 to 356, the method further comprising: In response to detecting the second portion of the first user input after the first portion of the first user input and based on determining that the second portion of the first user input includes movement of the first portion of the user in a second corresponding direction relative to the virtual object toward the user's viewpoint, move the virtual object toward the user's viewpoint according to the second portion of the first user input.

358. The method according to claim 357, the method further comprising: In response to detecting the second portion of the first user input after the first portion of the first user input and based on determining that the second portion of the first user input includes movement of the first portion of the user in the corresponding direction away from the user's viewpoint that exceeds the threshold amount, move the virtual object toward the user's viewpoint, regardless of the movement of the first portion of the user.

359. The method according to any one of claims 355 to 358, the method further comprising, in response to detecting the first portion of the first user input, reducing the visual salience of the virtual object relative to the user interface.

360. The method according to claim 359, wherein reducing the visual salience of the virtual object includes gradually reducing the visual salience of the virtual object based on the movement of the first portion of the user in the first portion of the first user input.

361. The method according to any one of claims 359 to 360, the method further comprising: When displaying the virtual object, detect a first portion of a second user input pointing to the virtual object via the one or more input devices; In response to detecting the first portion of the second user input pointing to the virtual object: Perform an operation corresponding to the second user input pointing to the virtual object based on determining that the user input pointing to the virtual object is detected when the virtual object is displayed with a first reduced level of visual salience; And Abandon performing the operation corresponding to the second user input pointing to the virtual object based on determining that the user input pointing to the virtual object is detected when the virtual object is displayed with a second reduced level of visual salience.

362. The method according to any one of claims 359 to 361, the method further comprising: When the virtual object is displayed with a first reduced level of visual salience in response to detecting the first portion of the first user input, detect a second portion of the first user input after the first portion of the first user input via the one or more input devices; And In response to detecting the second portion of the first user input: Increase the visual salience of the virtual object based on determining that the second portion of the first user input includes movement of the first portion of the user in a direction relative to the virtual object toward the user's viewpoint; And Enhancing the visual saliency of the virtual object based on determining that the second portion of the first user input includes movement of the first portion of the user in the respective direction away from the user's viewpoint by more than the threshold amount.

363. The method according to any one of claims 355 to 362, the method further comprising: Before detecting the first portion of the first user input, displaying, via the display generating component, visual feedback indicating a position of an interaction between the first portion of the user and the virtual object at a corresponding portion of the virtual object.

364. The method according to any one of claims 355 to 363, wherein the threshold amount of movement corresponds to a first threshold position behind the virtual object relative to the user's viewpoint before detecting the first user input, the method further comprising when detecting the first user input: Based on determining that the first user input includes movement of the first portion of the user in the respective direction such that the first portion of the user moves past a second threshold position behind the virtual object relative to the user's viewpoint before detecting the first user input, moving the virtual object away from the user's viewpoint and reducing the visual saliency of the virtual object.

365. The method according to any one of claims 355 to 364, the method further comprising: Before detecting the first portion of the first user input, displaying visual feedback of an interaction between the first portion of the user and the virtual object with a first luminance amount based on a distance between the first portion of the user and the position corresponding to the virtual object; When the virtual object is displayed via the display generating component with the first luminance amount, detecting, via the one or more input devices, an input corresponding to a selection of the virtual object provided by the first portion of the user; And In response to detecting the input provided by the first portion of the user, displaying, via the display generating component, the visual feedback at a corresponding portion of the virtual object with a second luminance amount that is at least a corresponding amount greater than the first luminance amount and does not vary based on a distance between the first portion of the user and the first object.

366. The method according to any one of claims 355 to 365, the method further comprising: When the virtual object is displayed via the display generating component, detecting, via the one or more input devices, an input corresponding to a selection of the virtual object provided by the first portion of the user; And In response to detecting the input provided by the first portion of the user, outputting an audio indicating a selection of the virtual object.

367. The method according to any one of claims 355 to 366, the method further comprising: When the virtual object is displayed via the display generating component, an input provided by the first part of the user and pointing to the virtual object and not corresponding to a selection of the virtual object is detected via the one or more input devices, including detecting the first part of the user at a position corresponding to the virtual object; And In response to detecting the input provided by the first part of the user, output an audio indicating a direct interaction between the first part of the user and the virtual object.

368. The method according to any one of claims 355 to 367, the method further comprising: When the virtual object is displayed at a first position in the user interface, where the first position is a first distance from the user's viewpoint, detect a second user input for displaying a virtual keyboard via the one or more input devices; In response to detecting the second user input, display the virtual keyboard at a second position in the user interface, including moving the virtual object from the first position to a third position in the user interface according to a determination that the virtual object will intersect the virtual keyboard when the virtual keyboard is displayed at the second position, where the third position is a second distance greater than the first distance from the user's viewpoint; When the virtual keyboard is displayed at the second position, detect a third user input for stopping the display of the virtual keyboard via the one or more input devices; And In response to detecting the third user input, stop the display of the virtual keyboard and move the virtual object from the third position to the first position in the user interface.

369. The method according to any one of claims 355 to 368, the method further comprising: When the first part of the first user input is detected, display the virtual object behind the first part of the user via the display generating component relative to the user's viewpoint; And In response to detecting the second part of the first user input and according to a determination that the second part of the first user input includes a movement of the first part of the user in the corresponding direction away from the user's viewpoint that exceeds the threshold amount, display the virtual object in front of the first part of the user relative to the user's viewpoint.

370. The method according to any one of claims 355 to 369, the method further comprising when the first user input is detected: In response to detecting a third portion of the first user input before the first portion of the first user input and based on determining that the third portion of the first user input includes a movement of the first portion of the user in the corresponding direction that is no more than a threshold distance after the first portion of the user reaches a position corresponding to the virtual object, refrain from moving the virtual object away from the user's viewpoint, where the first portion of the first user input includes a movement of the first portion of the user in the corresponding direction that is more than the threshold distance after the first portion of the user reaches the position corresponding to the virtual object.

371. The method according to any one of claims 355 to 370, the method further comprising: Simultaneously display, via the display generation component, the virtual object and a second virtual object associated with the virtual object, where the position of the second virtual object is based on an analog physical relationship with the position of the virtual object.

372. The method according to claim 371, the method further comprising: Detect, via the one or more input devices, a movement of the virtual object away from the user's viewpoint; And In response to detecting the movement of the virtual object away from the user's viewpoint, move the second virtual object away from the user's viewpoint, where the movement of the second virtual object away from the user's viewpoint is based on the analog physical relationship with the movement of the virtual object.

373. The method according to any one of claims 371 to 372, the method further comprising: Detect, via the one or more input devices, a movement of the virtual object towards the user's viewpoint; And In response to detecting the movement of the virtual object towards the user's viewpoint, move the second virtual object towards the user's viewpoint, where the movement of the second virtual object towards the user's viewpoint is based on the analog physical relationship with the movement of the virtual object.

374. The method according to any one of claims 371 to 373, where simultaneously displaying the virtual object and the second virtual object via the display generation component includes displaying the virtual object at a first distance relative to the user's viewpoint and displaying the second virtual object at a second distance different from the first distance relative to the user's viewpoint.

375. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: Displaying, via the display generation component, a user interface including a virtual object; When displaying the user interface, detect a first user input via the one or more input devices, the first user input pointing to the virtual object and including movement of a first part of the user of the computer system in a corresponding direction towards the virtual object relative to the user's viewpoint; and When the first user input is detected: In response to detecting a first part of the first user input, move the virtual object away from the user's viewpoint according to the movement of the first part of the user; and In response to detecting a second part of the first user input after the first part of the first user input and based on determining that the second part of the first user input includes movement of the first part of the user in the corresponding direction away from the user's viewpoint that exceeds a threshold amount, move the virtual object towards the user's viewpoint.

376. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: Display a user interface including a virtual object via the display generation component; When displaying the user interface, a first user input is detected via the one or more input devices, the first user input being directed to the virtual object and including movement of a first portion of the user of the computer system in a respective direction towards the virtual object relative to the user's viewpoint; and When the first user input is detected: In response to detecting a first part of the first user input, move the virtual object away from the user's viewpoint according to the movement of the first part of the user; and In response to detecting a second part of the first user input after the first part of the first user input and based on determining that the second part of the first user input includes movement of the first part of the user in the corresponding direction away from the user's viewpoint that exceeds a threshold amount, move the virtual object towards the user's viewpoint.

377. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; Means for: displaying a user interface including a virtual object via the display generation component; Means for: when displaying the user interface, detecting a first user input via the one or more input devices, the first user input pointing to the virtual object and including movement of a first part of the user of the computer system in a corresponding direction towards the virtual object relative to the user's viewpoint; and Means for: when the first user input is detected: In response to detecting a first part of the first user input, move the virtual object away from the user's viewpoint according to the movement of the first part of the user; and In response to detecting a second portion of the first user input after the first portion of the first user input and based on determining that the second portion of the first user input includes a movement of the first portion of the user in a corresponding direction away from the user's viewpoint that exceeds a threshold amount, move the virtual object toward the user's viewpoint.

378. A computer system communicatively coupled to a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods according to claims 355 to 374.

379. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicatively coupled to a display generation component and one or more input devices, cause the computer system to perform any of the methods according to claims 355 to 374.

380. A computer system communicatively coupled to a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and means for performing any of the methods according to claims 355 to 374.

381. A method, the method comprising: at a computer system communicatively coupled to a display generation component and one or more input devices: when a user interface is displayed via the display generation component, detect a first air gesture performed by a first portion of the user via the one or more input devices; in response to detecting the first air gesture: select a first object in the user interface based on determining that the first air gesture meets a first set of interaction criteria and that the user's gaze is directed at the first object in the user interface when the first air gesture is detected; and display a select user interface object in the user interface without selecting an object in the user interface based on determining that the first air gesture meets a second set of interaction criteria different from the first set of interaction criteria; when, in response to detecting the first input, the select user interface object is displayed in the user interface based on determining that the first air gesture meets the second set of interaction criteria, detect a movement of the first portion of the user via the one or more input devices; in response to detecting the movement of the first portion of the user, move the select user interface object in the user interface based on the movement of the first portion of the user; after moving the select user interface object in the user interface based on the movement of the first portion of the user, detect a second air gesture corresponding to a selection input performed by the first portion of the user via the one or more input devices; and In response to detecting the second air gesture: Select the first object in the user interface based on determining that the selected user interface object is located at a position corresponding to the first object when the second air gesture is detected.

382. The method according to claim 381, wherein the second set of interaction criteria includes criteria satisfied based on determining that the first air gesture performed by the first portion of the user is a first gesture of the user's hand relative to a corresponding portion of the user's body.

383. The method according to claim 382, wherein the movement of the first portion of the user includes movement of the user's hand.

384. The method according to any one of claims 382 to 383, wherein the second air gesture performed by the first portion of the user is a second gesture different from the first gesture performed using at least a first finger of the user's hand. The method according to any one of claims 381 to 384, wherein, In response to detecting the first air gesture, display the selected user interface object at a position corresponding to the position of the user's gaze in the user interface based on determining that the first air gesture satisfies the second set of interaction criteria.

386. The method according to any one of claims 381 to 385, wherein the first set of interaction criteria includes criteria satisfied based on determining that when an air gesture is detected, the user's gaze points to a corresponding object having a size greater than a threshold size, and the method further includes: When, in response to detecting the first input, the selected user interface object is displayed in the user interface based on determining that the first air gesture satisfies the second set of interaction criteria, detect a third air gesture corresponding to a selection input performed by the first portion of the user via the one or more input devices; And In response to detecting the third air gesture: Select the second object in the user interface based on determining that the selected user interface object is located at a position corresponding to the second object when the third air gesture is detected, wherein the second object has a size within the threshold size.

387. The method according to claim 386, wherein the first object has a size within the threshold size, and the method further includes: In response to detecting the first air gesture corresponding to a selection input, based on determining that the first air gesture does not satisfy the first set of interaction criteria and that the user's gaze points to the first object in the user interface when the first air gesture is detected: Abandon selecting the first object in the user interface; And Select a third object in the user interface within a threshold distance from the first object, wherein the third object has a size greater than the threshold size.

388. The method according to any one of claims 381 to 387, the method further includes: When, in response to detecting the first input, the selection user interface object is displayed in the user interface based on determining that the first air gesture meets the second set of interaction criteria, detect an air gesture corresponding to a selection input via the one or more input devices; And In response to detecting the air gesture: Select the optional text based on determining that the position of the selection user interface object corresponds to the position of the optional text in the user interface when the air gesture is detected.

389. The method according to any one of claims 381 to 388, the method further comprising: When, in response to detecting the first input, the selection user interface object is displayed in the user interface based on determining that the first air gesture meets the second set of interaction criteria, detect a corresponding event that meets one or more criteria; And In response to detecting the corresponding event: Stop displaying the selection user interface object in the user interface.

390. The method according to claim 389, wherein detecting the corresponding event that meets the one or more criteria includes determining that a corresponding object is selected in response to detecting a corresponding air gesture corresponding to a selection input performed by the first part of the user when the selection user interface object is displayed.

391. The method according to any one of claims 389 to 390, wherein detecting the corresponding event that meets the one or more criteria includes determining that the user interface is scrolled in response to detecting a corresponding input corresponding to a scroll input performed by the first part of the user when the selection user interface object is displayed.

392. The method according to any one of claims 389 to 391, wherein detecting the corresponding event that meets the one or more criteria includes detecting that the gaze of the user deviates from the selection user interface object in the user interface by more than a threshold distance.

393. The method according to any one of claims 389 to 392, wherein detecting the corresponding event that meets the one or more criteria includes detecting that the first part of the user is lowered to a height lower than a threshold height relative to the viewpoint of the user.

394. The method according to any one of claims 389 to 393, wherein stopping displaying the selection user interface object in the user interface includes fading out the selection user interface object.

395. The method according to claim 394, the method further comprising: When fading out the selection user interface object and before detecting the completion of the corresponding event, detect a second corresponding event; And In response to detecting the second corresponding event: Restore the display of the selection user interface object in the user interface based on determining that the second corresponding event meets one or more second criteria.

396. The method according to any one of claims 381 to 395, the method further comprising: When, in response to detecting the first input, the selection user interface object is displayed based on determining that the first air gesture meets the second set of interaction criteria, detect, via the one or more input devices, a corresponding input corresponding to a request to stop the display of the selection user interface object; and In response to detecting the corresponding input: Stop the display of the selection user interface object in the user interface.

397. The method according to any one of claims 381 to 396, the method further comprising: When the selection user interface object is displayed after detecting the second air gesture, detect, via the one or more input devices, a corresponding input pointing to a second object different from the first object; and In response to detecting the corresponding input pointing to the second object: Stop the display of the selection user interface object in the user interface.

398. The method according to any one of claims 381 to 397, the method further comprising: When the selection user interface object is located at the position corresponding to the first object and before detecting the second air gesture: Based on determining that the first object includes first type content, display the selection user interface object in a first visual appearance via the display generation component; and Based on determining that the first object includes second type content different from the first type, display the selection user interface object in a second visual appearance different from the first visual appearance.

399. The method according to any one of claims 381 to 398, the method further comprising: In response to detecting the second air gesture: Based on determining that the selection user interface is located at the position corresponding to the first object and that the user's gaze points to the first object when the second air gesture is detected, calibrate the determined position of the gaze relative to the user interface based on the corresponding part of the first object.

400. The method according to claim 399, wherein calibrating the determined position of the gaze relative to the user interface based on the corresponding part of the first object comprises: Based on determining that the first object has a size within a threshold size and that a horizontal component of the position of the gaze is offset by a first amount from the corresponding part of the first object, and a vertical component of the position of the gaze is offset by a second amount from the corresponding part of the first object: Change the offset of the horizontal component by a third amount based on the first amount; And Change the offset of the vertical component by a fourth amount based on the second amount.

401. The method according to any one of claims 399 to 400, wherein calibrating the determined position of the gaze relative to the user interface based on the corresponding part of the first object comprises: Based on determining that the first object has a size greater than the threshold size and that a first component of the position of the gaze is offset by a first amount from the corresponding part of the first object, and a second component of the position of the gaze is offset by a second amount greater than the first amount from the corresponding part of the first object: Change the offset of the first component by a third amount based on the first amount without changing the offset of the second component.

402. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: When a user interface is displayed via the display generation component, detecting a first air gesture performed by a first part of the user via the one or more input devices; In response to detecting the first air gesture: Selecting the first object in the user interface based on determining that the first air gesture meets a first set of interaction criteria and that the user's gaze points to the first object in the user interface when the first air gesture is detected; And Based on determining that the first air gesture meets a second set of interaction criteria different from the first set of interaction criteria, displaying a select user interface object in the user interface via the display generation component without selecting an object in the user interface; When the select user interface object is displayed in the user interface in response to detecting the first input based on determining that the first air gesture meets the second set of interaction criteria, detecting a movement of the first part of the user via the one or more input devices; In response to detecting the movement of the first part of the user, moving the select user interface object in the user interface according to the movement of the first part of the user; After moving the select user interface object in the user interface according to the movement of the first part of the user, detecting a second air gesture corresponding to a selection input performed by the first part of the user via the one or more input devices; And In response to detecting the second air gesture: Selecting the first object in the user interface based on determining that the select user interface object is located at a position corresponding to the first object when the second air gesture is detected.

403. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, cause the computer system to perform a method including: When a user interface is displayed via the display generation component, detecting a first air gesture performed by a first part of the user via the one or more input devices; In response to detecting the first air gesture: Selecting the first object in the user interface based on determining that the first air gesture meets a first set of interaction criteria and that the user's gaze points to the first object in the user interface when the first air gesture is detected; And Based on determining that the first in-air gesture meets a second set of interaction criteria different from the first set of interaction criteria, display a selection user interface object in the user interface without selecting an object in the user interface via the display generation component; When, in response to detecting the first input, the selection user interface object is displayed in the user interface based on determining that the first in-air gesture meets the second set of interaction criteria, detect a movement of the first part of the user via the one or more input devices; In response to detecting the movement of the first part of the user, move the selection user interface object in the user interface based on the movement of the first part of the user; After moving the selection user interface object in the user interface based on the movement of the first part of the user, detect a second in-air gesture corresponding to a selection input performed by the first part of the user via the one or more input devices; And In response to detecting the second in-air gesture: Select the first object in the user interface based on determining that the selection user interface object is located at a position corresponding to the first object when the second in-air gesture is detected.

404. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; Means for: when a user interface is displayed via the display generation component, detecting a first in-air gesture performed by a first part of a user via the one or more input devices; Means for: in response to detecting the first in-air gesture: Selecting the first object in the user interface based on determining that the first in-air gesture meets a first set of interaction criteria and that the user's gaze is directed at the first object in the user interface when the first in-air gesture is detected; And Based on determining that the first in-air gesture meets a second set of interaction criteria different from the first set of interaction criteria, display a selection user interface object in the user interface without selecting an object in the user interface via the display generation component; Means for: when, in response to detecting the first input, the selection user interface object is displayed in the user interface based on determining that the first in-air gesture meets the second set of interaction criteria, detecting a movement of the first part of the user via the one or more input devices; Means for: in response to detecting the movement of the first part of the user, moving the selection user interface object in the user interface based on the movement of the first part of the user; Means for: after moving the selection user interface object in the user interface based on the movement of the first part of the user, detecting a second in-air gesture corresponding to a selection input performed by the first part of the user via the one or more input devices; And Means for: in response to detecting the second in-air gesture: Select the first object in the user interface based on determining that the selected user interface object is located at a position corresponding to the first object when the second air gesture is detected.

405. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one of the methods according to claims 381 to 401.

406. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 381 to 401.

407. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And Means for performing any one of the methods according to claims 381 to 401.

408. A method, the method comprising: At a computer system communicating with a display generation component and one or more input devices: Display, via the display generation component, a user interface including selectable virtual objects; When the user interface including the selectable virtual objects is displayed: Based on determining that the attention of a user of the computer system pointing to the selectable virtual object meets a first set of one or more criteria, display, via the display generation component, a visual indicator indicating progress towards selecting the selectable virtual object based on the user's attention, wherein the first set of one or more criteria includes that the movement of the user's attention is below a movement threshold to meet the requirements of the first set of one or more criteria; And After displaying the visual indicator: Based on determining that the user's attention has met a second set of one or more criteria relative to the visual indicator when the visual indicator is displayed, perform a selection operation associated with the selectable virtual object, wherein the second set of one or more criteria includes that the user's attention has persisted within a threshold distance of the visual indicator for at least a first time threshold to meet the requirements of the second set of one or more criteria; And Based on determining that the user's attention does not meet the second set of one or more criteria when the visual indicator is displayed, abandon performing the selection operation associated with the selectable virtual object.

409. The method according to claim 408, the method further comprising when the user interface including the selectable virtual objects is displayed: Based on determining that the attention of the user directed to the optional virtual object does not meet one or more of the first set of criteria, abandon displaying the visual indicator indicating the progress of selecting the optional virtual object based on the attention of the user.

410. The method according to any one of claims 408 to 409, wherein the one or more of the first set of criteria includes that the movement of the attention of the user is below the movement threshold for a threshold time to meet a second requirement of the one or more of the first set of criteria.

411. The method according to any one of claims 408 to 410, the method further comprising, after displaying the visual indicator: Based on determining that the attention of the user is outside the threshold distance of the visual indicator, pause the progress of selecting the optional virtual object as indicated by the visual indicator.

412. The method according to any one of claims 408 to 411, wherein the one or more of the second set of criteria includes that the movement of the attention of the user is below the movement threshold to meet a second requirement of the one or more of the second set of criteria, and the method further comprising, after displaying the visual indicator: Based on determining that the attention of the user does not meet the one or more of the second set of criteria because the movement of the attention of the user is not below the movement threshold when the visual indicator is displayed, pause the progress of selecting the optional virtual object as indicated by the visual indicator.

413. The method according to any one of claims 408 to 412, wherein the one or more of the second set of criteria includes that the movement of the attention of the user is below the movement threshold to meet a second requirement of the one or more of the second set of criteria, and the method further comprising, after displaying the visual indicator: Based on determining that the attention of the user does not meet the one or more of the second set of criteria because the movement of the attention of the user is not below the movement threshold when the visual indicator is displayed, display an indication that the progress of selecting the optional virtual object as indicated by the visual indicator has decreased.

414. The method according to any one of claims 408 to 413, wherein displaying the visual indicator indicating the progress of selecting the optional virtual object based on the attention of the user via the display generating component includes: Based on determining that the position of the attention of the user when the one or more of the first set of criteria was met was a first position, display the visual indicator at the first position of the visual indicator; And based on determining that the position of the attention of the user when the one or more of the first set of criteria was met was a second position different from the first position, display the visual indicator at the second position of the visual indicator.

415. The method according to any one of claims 408 to 414, wherein displaying the visual indicator via the display generation component includes displaying the visual indicator at a position corresponding to the position of the user's attention relative to the optional virtual object.

416. The method according to any one of claims 408 to 415, the method further comprising, after displaying the visual indicator: Stopping the display of the visual indicator based on determining that the user's attention is outside the threshold distance of the visual indicator.

417. The method according to any one of claims 408 to 416, the method further comprising when performing the selection operation associated with the optional virtual object: Based on determining that the user's attention has moved outside the threshold distance of the visual indicator: Maintaining the execution of the selection operation associated with the optional virtual object; Maintaining the display of the visual indicator at a first position relative to the optional virtual object; and displaying visual feedback at a second position in the user interface where the user's attention is outside the optional virtual object in the user interface.

418. The method according to any one of claims 408 to 417, wherein the one or more first criteria include criteria that are satisfied when a first user-specified option is enabled and not satisfied when the first user-specified option is disabled.

419. The method according to any one of claims 408 to 418, the method further comprising, before displaying the visual indicator: Displaying, via the display generation component, visual feedback indicating the position of the user's attention relative to the optional virtual object, wherein the visual feedback indicating the position of the user's attention relative to the optional virtual object is different from the visual indicator indicating the progress towards selecting the optional virtual object.

420. The method according to claim 419, the method further comprising when displaying the visual indicator: Maintaining the display of the visual feedback indicating the position of the user's attention relative to the optional virtual object via the display generation component.

421. The method according to any one of claims 408 to 420, the method further comprising: Based on determining that the user-specified option has a first value, the first time threshold is a first corresponding time threshold; Based on determining that the user-specified option has a second value different from the first value, the first time threshold is a second corresponding time threshold different from the first corresponding time threshold.

422. The method according to any one of claims 408 to 421, wherein the visual indicator includes a boundary surrounding an internal area, and the method further comprises: When displaying the visual indicator, before performing the selection operation associated with the optional virtual object and when the user's attention is within the threshold distance of the visual indicator, change the visual appearance of the inner region of the visual indicator according to the elapsed time since the user's attention moved within the threshold distance of the visual indicator.

423. The method according to claim 422, wherein changing the visual appearance of the inner region of the visual indicator according to the elapsed time since the user's attention moved within the threshold distance of the visual indicator includes changing the visual appearance of the visual indicator from the center to the boundary.

424. The method according to any one of claims 422 to 423, wherein displaying the visual indicator via the display generating component includes gradually increasing the visual saliency of the visual indicator while changing the visual appearance of the inner region of the visual indicator according to the elapsed time that the user's attention remains within the threshold distance.

425. The method according to any one of claims 422 to 424, wherein changing the visual appearance of the inner region of the visual indicator includes: changing the visual appearance of the inner region of the visual indicator at a first rate according to determining that the elapsed time since having met the one or more first criteria is a first amount of time; and changing the visual appearance of the inner region of the visual indicator at a second rate different from the first rate according to determining that the elapsed time since having met the one or more first criteria is a second amount of time different from the first amount of time.

426. A computer system in communication with a display generating component and one or more input devices, the computer system comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: displaying a user interface including an optional virtual object via the display generating component; when displaying the user interface including the optional virtual object: displaying, via the display generating component, a visual indicator indicating progress towards selecting the optional virtual object based on the user's attention according to determining that the attention of a user of the computer system pointing to the optional virtual object meets a first set of one or more criteria, wherein the first set of one or more criteria includes that the movement of the user's attention is below a movement threshold to meet the requirements of the first set of one or more criteria; and after displaying the visual indicator: Performing a selection operation associated with the optional virtual object based on determining that the user's attention satisfies a second set of one or more criteria with respect to the visual indicator when the visual indicator is displayed, wherein the second set of one or more criteria includes that the user's attention has been within a threshold distance of the visual indicator for at least a first time threshold to meet the requirements of the second set of one or more criteria; And Aborting the execution of the selection operation associated with the optional virtual object based on determining that the user's attention does not satisfy the second set of one or more criteria when the visual indicator is displayed.

427. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: Displaying, via the display generation component, a user interface including an optional virtual object; When the user interface including the optional virtual object is displayed: Based on determining that the attention of a user of the computer system directed to the selectable virtual object meets one or more criteria of a first set, display, via the display generating component, a visual indicator indicative of progress towards selecting the selectable virtual object based on the user's attention, wherein the one or more criteria of the first set include that movement of the user's attention is below a movement threshold to meet the requirements of the one or more criteria of the first set; And After displaying the visual indicator: Performing a selection operation associated with the optional virtual object based on determining that the user's attention satisfies a second set of one or more criteria with respect to the visual indicator when the visual indicator is displayed, wherein the second set of one or more criteria includes that the user's attention has been within a threshold distance of the visual indicator for at least a first time threshold to meet the requirements of the second set of one or more criteria; And Aborting the execution of the selection operation associated with the optional virtual object based on determining that the user's attention does not satisfy the second set of one or more criteria when the visual indicator is displayed.

428. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; Means for: displaying, via the display generation component, a user interface including an optional virtual object; Means for: when the user interface including the optional virtual object is displayed: Displaying, via the display generation component, a visual indicator indicating progress towards selecting the optional virtual object based on the user's attention, based on determining that the attention of the user of the computer system pointing to the optional virtual object satisfies a first set of one or more criteria, wherein the first set of one or more criteria includes that the movement of the user's attention is below a movement threshold to meet the requirements of the first set of one or more criteria; And Means for: after displaying the visual indicator: Performing a selection operation associated with the optional virtual object based on determining that the user's attention has met one or more criteria of a second set with respect to the visual indicator when the visual indicator is displayed, wherein the one or more criteria of the second set include that the user's attention has been within a threshold distance of the visual indicator for at least a first time threshold to meet the requirements of the one or more criteria of the second set; and Abandoning the execution of the selection operation associated with the optional virtual object based on determining that the user's attention does not meet the one or more criteria of the second set when the visual indicator is displayed.

429. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one of the methods according to claims 408 to 425.

430. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 408 to 425.

431. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; and Means for performing any one of the methods according to claims 408 to 425.