Device and method for interacting with graphical user interface

Through the combination of gaze-tracking sensors and display generation components, a more efficient and intuitive graphical user interface interaction is achieved, solving the problems of low interaction efficiency and waste of resources in the prior art, especially power consumption savings in battery-driven devices.

CN120447739APending Publication Date: 2025-08-08APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510544677.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-20
Filing Date
2023-09-22
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing graphical user interface interaction methods are inefficient, cumbersome user inputs and error-prone, resulting in wasted computer system resources and user cognitive burden, especially in battery-driven devices, which are relatively high in power consumption.

Method used

Through the computer system's communication with the gaze tracking sensor and the display generation component, user interface operations are performed using gaze pointing detection, reducing the number and complexity of user input, such as automatically performing display and hidden operations of interface objects by detecting the user's gaze direction.

Benefits of technology

Improves the efficiency and intuitiveness of user interface interaction, reduces user input, saves resource consumption of computer systems, and extends battery life in battery-driven devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120447739A_ABST
    Figure CN120447739A_ABST
Patent Text Reader

Abstract

The present disclosure generally relates to devices, methods for interacting with a graphical user interface. In some embodiments, the present disclosure includes techniques and user interfaces for interacting with graphical user interfaces using gaze. In some embodiments, the present disclosure includes techniques and user interfaces for relocating virtual objects. In some embodiments, the present disclosure includes techniques and user interfaces for converting modes in which a camera captures a user interface.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application with international application number PCT / US2023 / 033549, international application date September 22, 2023, entry into the Chinese national phase date March 21, 2025, national application number 202380068154.1, and invention name “Device and method for interacting with a graphical user interface”. Technical Field

[0002] This application claims priority to U.S. Patent Application No. 18 / 370,847, entitled “DEVICES, METHODS, FOR INTERACTING WITH GRAPHICAL USER INTERFACES,” filed on September 20, 2023, and U.S. Provisional Patent Application No. 63 / 409,744, entitled “DEVICES, METHODS, FOR INTERACTING WITH GRAPHICAL USER INTERFACES,” filed on September 24, 2022. The contents of each of these patent applications are incorporated herein by reference in their entirety.

[0003] The present disclosure generally relates to computer systems that provide computer-generated experiences in communication with display generation components and, optionally, one or more cameras and / or gaze tracking sensors, including but not limited to electronic devices that provide virtual reality and mixed reality experiences via displays. Background Art

[0004] In recent years, the development of computer systems for augmented reality has increased significantly. Example augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices for computer systems and other electronic computing devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays) are used to interact with virtual / augmented reality environments. Example virtual elements include virtual objects such as digital images, videos, text, icons, and control elements (such as buttons and other graphics). Summary of the Invention

[0005] Some methods and interfaces for interacting with graphical user interfaces (e.g., interacting with virtual objects, applications, augmented reality environments, mixed reality environments, and virtual reality environments via graphical user interfaces) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve desired results in an augmented reality environment, and systems where virtual object manipulation is complex, cumbersome, and error-prone can place a significant cognitive burden on users and detract from the experience of the virtual / augmented reality environment. Furthermore, these methods take longer than necessary, wasting the computer system's energy. This latter consideration is particularly important in battery-powered devices.

[0006] Therefore, there is a need for computer systems having improved methods and interfaces for interacting with graphical user interfaces, making the interactions more efficient and intuitive for users. Such methods and interfaces optionally supplement or replace conventional methods for interacting with graphical user interfaces. Such methods and interfaces reduce the amount, extent, and / or nature of inputs from users by helping users understand the connection between inputs provided and the device's responses to those inputs, thereby creating a more efficient human-computer interface.

[0007] The above-mentioned defects and other problems associated with the user interface of the computer system are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a laptop, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device, such as a watch or a head-mounted device). In some embodiments, the computer system has a touch pad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touch screen" or "touch screen display"). In some embodiments, the computer system has one or more eye tracking components. In some embodiments, the computer system has one or more hand tracking components. In some embodiments, in addition to the display generation component, the computer system also has one or more output devices, which include one or more tactile output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, a memory, and one or more modules, a program or instruction set stored in the memory for performing multiple functions. In some embodiments, the user interacts with the GUI through contacts and gestures of a stylus and / or fingers on a touch-sensitive surface, movement of the user's eyes and hands in space relative to the GUI (and / or computer system) or the user's body (as captured by a camera and other motion sensors), and / or voice input (as captured by one or more audio input devices). In some embodiments, the functions performed by interaction optionally include image editing, drawing, presentations, word processing, spreadsheet creation, playing games, making and receiving calls, video conferencing, sending and receiving emails, instant messaging, test support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. Executable instructions for performing these functions are optionally included in a transient and / or non-transient computer-readable storage medium or other computer program product configured for execution by one or more processors.

[0008] There is a need for electronic devices having improved methods and interfaces for interacting with graphical user interfaces. Such methods and interfaces can supplement or replace conventional methods for interacting with graphical user interfaces. Such methods and interfaces reduce the amount, extent, and / or nature of input from a user and produce a more efficient human-computer interface. For battery-powered computing devices, such methods and interfaces conserve power and increase the time between battery charges.

[0009] In some embodiments, a method is described, performed at a computer system in communication with one or more gaze tracking sensors and a display generation component. The method includes: displaying, via the display generation component, a corresponding user interface, wherein displaying the corresponding user interface includes displaying: a plurality of edges of the corresponding user interface including a first edge and a second edge different from the first edge; a first user interface object positioned along the first edge corresponding to a first operation; and a second user interface object positioned along the second edge corresponding to a second operation different from the first operation; while displaying the first user interface object and the second user interface object, detecting, via the one or more gaze tracking sensors, a gaze of a user of the computer system directed toward a corresponding portion of the corresponding user interface; and in response to detecting that the gaze of the user of the computer system is directed toward the corresponding portion of the corresponding user interface: based on determining that the corresponding portion of the corresponding user interface corresponds to the first user interface object: performing the first operation; and continuing to display the first user interface object while ceasing to display the second user interface object; based on determining that the corresponding portion of the corresponding user interface corresponds to the second user interface object: performing the second operation; and continuing to display the second user interface object while ceasing to display the first user interface object.

[0010] In some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more gaze tracking sensors and a display generation component, the one or more programs including instructions for: displaying a corresponding user interface via the display generation component, wherein displaying the corresponding user interface includes displaying: a plurality of edges of the corresponding user interface including a first edge and a second edge different from the first edge; a first user interface object positioned along the first edge corresponding to a first operation; and a second user interface object positioned along the second edge corresponding to a second operation different from the first operation; while displaying the first user interface object and the second user interface object, detecting via the one or more gaze tracking sensors that a gaze of a user of the computer system is directed toward a corresponding portion of the corresponding user interface; and in response to detecting that the gaze of the user of the computer system is directed toward the corresponding portion of the corresponding user interface: based on determining that the corresponding portion of the corresponding user interface corresponds to the first user interface object: performing the first operation; and continuing to display the first user interface object while ceasing to display the second user interface object; based on determining that the corresponding portion of the corresponding user interface corresponds to the second user interface object: performing the second operation; and continuing to display the second user interface object while ceasing to display the first user interface object.

[0011] In some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more gaze tracking sensors and a display generation component, the one or more programs including instructions for: displaying, via the display generation component, a corresponding user interface, wherein displaying the corresponding user interface comprises displaying: a plurality of edges of the corresponding user interface including a first edge and a second edge different from the first edge; a first user interface object positioned along the first edge corresponding to a first operation; and a second user interface object positioned along the second edge corresponding to a second operation different from the first operation; while displaying the first user interface object and the second user interface object, detecting, via the one or more gaze tracking sensors, a gaze of a user of the computer system directed toward a corresponding portion of the corresponding user interface; and in response to detecting the gaze of the user of the computer system directed toward the corresponding portion of the corresponding user interface: based on determining that the corresponding portion of the corresponding user interface corresponds to the first user interface object: performing the first operation; and continuing to display the first user interface object while ceasing to display the second user interface object; and based on determining that the corresponding portion of the corresponding user interface corresponds to the second user interface object: performing the second operation; and continuing to display the second user interface object while ceasing to display the first user interface object.

[0012] In some embodiments, a computer system configured to communicate with one or more gaze tracking sensors and a display generation component is described. The computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for the following operations: displaying a corresponding user interface via the display generation component, wherein displaying the corresponding user interface includes displaying: multiple edges of the corresponding user interface including a first edge and a second edge different from the first edge; a first user interface object corresponding to a first operation positioned along the first edge; and a second user interface object corresponding to a second operation different from the first operation positioned along the second edge; when displaying the first user interface object and the second user interface object, detecting via the one or more gaze tracking sensors that the gaze of the user of the computer system is directed to the corresponding portion of the corresponding user interface; and in response to detecting that the gaze of the user of the computer system is directed to the corresponding portion of the corresponding user interface: based on determining that the corresponding portion of the corresponding user interface corresponds to the first user interface object: performing the first operation; and continuing to display the first user interface object while stopping displaying the second user interface object; based on determining that the corresponding portion of the corresponding user interface corresponds to the second user interface object: performing the second operation; and continuing to display the second user interface object while stopping displaying the first user interface object.

[0013] In some embodiments, a computer system is described. The computer system is configured to communicate with one or more gaze tracking sensors and a display generation component and includes: means for displaying, via the display generation component, a corresponding user interface, wherein displaying the corresponding user interface includes displaying: a plurality of edges of the corresponding user interface including a first edge and a second edge different from the first edge; a first user interface object positioned along the first edge corresponding to a first operation; and a second user interface object positioned along the second edge corresponding to a second operation different from the first operation; means for detecting, via the one or more gaze tracking sensors, a gaze of a user of the computer system directed toward a corresponding portion of the corresponding user interface while displaying the first and second user interface objects; and means for, in response to detecting that the gaze of the user of the computer system is directed toward the corresponding portion of the corresponding user interface, performing the first operation based on a determination that the corresponding portion of the corresponding user interface corresponds to the first user interface object; and continuing to display the first user interface object while ceasing to display the second user interface object; and performing the second operation based on a determination that the corresponding portion of the corresponding user interface corresponds to the second user interface object; and continuing to display the second user interface object while ceasing to display the first user interface object.

[0014] In some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with one or more gaze tracking sensors and a display generation component. The one or more programs include instructions for: displaying a corresponding user interface via the display generation component, wherein displaying the corresponding user interface includes displaying: a plurality of edges of the corresponding user interface including a first edge and a second edge different from the first edge; a first user interface object positioned along the first edge corresponding to a first operation; and a second user interface object positioned along the second edge corresponding to a second operation different from the first operation; while displaying the first user interface object and the second user interface object, detecting via the one or more gaze tracking sensors that a gaze of a user of the computer system is directed toward a corresponding portion of the corresponding user interface; and in response to detecting that the gaze of the user of the computer system is directed toward the corresponding portion of the corresponding user interface: based on determining that the corresponding portion of the corresponding user interface corresponds to the first user interface object: performing the first operation; and continuing to display the first user interface object while ceasing to display the second user interface object; based on determining that the corresponding portion of the corresponding user interface corresponds to the second user interface object: performing the second operation; and continuing to display the second user interface object while ceasing to display the first user interface object.

[0015] In some embodiments, a method performed at a computer system in communication with a display generation component is described. The method includes: displaying a corresponding user interface via the display generation component, wherein displaying the corresponding user interface includes displaying: a first user interface object, wherein at least a first portion of the first user interface object is at least partially translucent and includes first content; and a second user interface object, wherein: at least a first portion of the second user interface object includes second content different from the first content; and the first user interface object is displayed in front of the second user interface object such that the first portion of the first user interface object covers the first portion of the second user interface object; while the first user interface object is displayed in front of the second user interface object, receiving a request to move the second user interface object in front of the first user interface object; and in response to receiving the request to move the second user interface object in front of the first user interface object: initiating a process to move the second user interface object in front of the first user interface object, the process including: modifying the visual appearance of the first portion of the first user interface object to include third content based on a first combination of the first content and the second content.

[0016] In some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for: displaying a corresponding user interface via the display generation component, wherein displaying the corresponding user interface includes displaying: a first user interface object, wherein at least a first portion of the first user interface object is at least partially translucent and includes first content; and a second user interface object, wherein: at least a first portion of the second user interface object includes second content different from the first content; and the first user interface object is displayed in front of the second user interface object such that the first portion of the first user interface object covers the first portion of the second user interface object; while the first user interface object is displayed in front of the second user interface object, receiving a request to move the second user interface object in front of the first user interface object; and in response to receiving the request to move the second user interface object in front of the first user interface object: initiating a process to move the second user interface object in front of the first user interface object, the process comprising: modifying the visual appearance of the first portion of the first user interface object to include third content based on a first combination of the first content and the second content.

[0017] In some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for: displaying a corresponding user interface via the display generation component, wherein displaying the corresponding user interface includes displaying: a first user interface object, wherein at least a first portion of the first user interface object is at least partially translucent and includes first content; and a second user interface object, wherein: at least a first portion of the second user interface object includes second content different from the first content; and the first user interface object is displayed in front of the second user interface object such that the first portion of the first user interface object covers the first portion of the second user interface object; while the first user interface object is displayed in front of the second user interface object, receiving a request to move the second user interface object in front of the first user interface object; and in response to receiving the request to move the second user interface object in front of the first user interface object: initiating a process to move the second user interface object in front of the first user interface object, the process comprising: modifying the visual appearance of the first portion of the first user interface object to include third content based on a first combination of the first content and the second content.

[0018] In some embodiments, a computer system configured to communicate with a display generation component is described. The computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for: displaying a corresponding user interface via the display generation component, wherein displaying the corresponding user interface includes displaying: a first user interface object, wherein at least a first portion of the first user interface object is at least partially translucent and includes first content; and a second user interface object, wherein: at least a first portion of the second user interface object includes second content different from the first content; and the first user interface object is displayed in front of the second user interface object such that the first portion of the first user interface object covers the first portion of the second user interface object; while the first user interface object is displayed in front of the second user interface object, receiving a request to move the second user interface object in front of the first user interface object; and in response to receiving the request to move the second user interface object in front of the first user interface object: initiating a process to move the second user interface object in front of the first user interface object, the process comprising: modifying the visual appearance of the first portion of the first user interface object to include third content based on a first combination of the first content and the second content.

[0019] In some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and includes: components for displaying a corresponding user interface via the display generation component, wherein displaying the corresponding user interface includes displaying: a first user interface object, wherein at least a first portion of the first user interface object is at least partially translucent and includes first content; and a second user interface object, wherein: at least a first portion of the second user interface object includes second content different from the first content; and the first user interface object is displayed in front of the second user interface object such that the first portion of the first user interface object overlays the first portion of the second user interface object; components for receiving a request to move the second user interface object in front of the first user interface object while the first user interface object is displayed in front of the second user interface object; and components for, in response to receiving the request to move the second user interface object in front of the first user interface object, initiating a process to move the second user interface object in front of the first user interface object, the process comprising: modifying the visual appearance of the first portion of the first user interface object to include third content based on a first combination of the first content and the second content.

[0020] In some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component. The one or more programs include instructions for: displaying a corresponding user interface via the display generation component, wherein displaying the corresponding user interface includes displaying: a first user interface object, wherein at least a first portion of the first user interface object is at least partially translucent and includes first content; and a second user interface object, wherein: at least a first portion of the second user interface object includes second content different from the first content; and the first user interface object is displayed in front of the second user interface object such that the first portion of the first user interface object covers the first portion of the second user interface object; while the first user interface object is displayed in front of the second user interface object, receiving a request to move the second user interface object in front of the first user interface object; and in response to receiving the request to move the second user interface object in front of the first user interface object: initiating a process to move the second user interface object in front of the first user interface object, the process comprising: modifying the visual appearance of the first portion of the first user interface object to include third content based on a first combination of the first content and the second content.

[0021] In some embodiments, a method is described for execution at a computer system in communication with a display generation component and one or more cameras. The method includes: displaying, via the display generation component and in a mixed reality environment, a camera capture user interface overlaid on a portion of a physical environment visible to a user of the computer system, wherein: the camera capture user interface is in a first mode; and when in the first mode, the camera capture user interface includes a set of one or more framed virtual objects, the set of one or more framed virtual objects being viewpoint-locked and indicating a first sub-portion of the physical environment to be captured by the one or more cameras upon receiving a first media capture request; while displaying the camera capture user interface in the first mode, receiving a request to transition the camera capture user interface to a second mode different from the first mode; and in response to receiving the request to transition the camera capture user interface to the second mode, displaying the camera capture user interface in the second mode, wherein: when in the second mode, the camera capture user interface includes a first representation of a field of view of at least a first camera of the one or more cameras; and the first representation is overlaid on a second sub-portion of the physical environment to be captured by the one or more cameras upon receiving a second media capture request.

[0022] In some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs including instructions for: displaying, via the display generation component and in a mixed reality environment, a camera capture user interface overlaid on a portion of a physical environment visible to a user of the computer system, wherein: the camera capture user interface is in a first mode; and when in the first mode, the camera capture user interface includes a set of one or more framed virtual objects, the set of one or more framed virtual objects being viewpoint-locked and indicating a first sub-portion of the physical environment to be captured by the one or more cameras upon receiving a first media capture request; while displaying the camera capture user interface in the first mode, receiving a request to transition the camera capture user interface to a second mode different from the first mode; and in response to receiving the request to transition the camera capture user interface to the second mode, displaying the camera capture user interface in the second mode, wherein: when in the second mode, the camera capture user interface includes a first representation of a field of view of at least a first camera of the one or more cameras; and the first representation is overlaid on a second sub-portion of the physical environment to be captured by the one or more cameras upon receiving a second media capture request.

[0023] In some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs including instructions for: displaying, via the display generation component and in a mixed reality environment, a camera capture user interface overlaid on a portion of a physical environment visible to a user of the computer system, wherein: the camera capture user interface is in a first mode; and when in the first mode, the camera capture user interface includes a set of one or more framed virtual objects, the set of one or more framed virtual objects being viewpoint-locked and indicating a first sub-portion of the physical environment to be captured by the one or more cameras upon receiving a first media capture request; while displaying the camera capture user interface in the first mode, receiving a request to transition the camera capture user interface to a second mode different from the first mode; and in response to receiving the request to transition the camera capture user interface to the second mode, displaying the camera capture user interface in the second mode, wherein: when in the second mode, the camera capture user interface includes a first representation of a field of view of at least a first camera of the one or more cameras; and the first representation is overlaid on a second sub-portion of the physical environment to be captured by the one or more cameras upon receiving a second media capture request.

[0024] In some embodiments, a computer system configured to communicate with a display generation component and one or more cameras is described. The computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for performing the following operations: displaying a camera capture user interface via the display generation component and in a mixed reality environment, the camera capture user interface overlaid on a portion of a physical environment visible to a user of the computer system, wherein: the camera capture user interface is in a first mode; and when in the first mode, the camera capture user interface includes a set of one or more framed virtual objects, the set of one or more framed virtual objects being viewpoint-locked and indicating a first sub-portion of the physical environment to be captured by the one or more cameras upon receiving a first media capture request; while displaying the camera capture user interface in the first mode, receiving a request to transition the camera capture user interface to a second mode different from the first mode; and in response to receiving the request to transition the camera capture user interface to the second mode, displaying the camera capture user interface in the second mode, wherein: when in the second mode, the camera capture user interface includes a first representation of a field of view of at least a first camera of the one or more cameras; and the first representation is overlaid on a second sub-portion of the physical environment to be captured by the one or more cameras upon receiving a second media capture request.

[0025] In some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and one or more cameras and includes: means for displaying, via the display generation component and in a mixed reality environment, a camera capture user interface overlaid on a portion of a physical environment visible to a user of the computer system, wherein: the camera capture user interface is in a first mode; and when in the first mode, the camera capture user interface includes a set of one or more framed virtual objects that are viewpoint-locked and indicate a first subportion of the physical environment to be captured by the one or more cameras upon receiving a first media capture request; means for receiving a request to transition the camera capture user interface to a second mode different from the first mode while displaying the camera capture user interface in the first mode; and means for displaying the camera capture user interface in the second mode in response to receiving the request to transition the camera capture user interface to the second mode, wherein: when in the second mode, the camera capture user interface includes a first representation of a field of view of at least a first camera of the one or more cameras; and the first representation is overlaid on a second subportion of the physical environment to be captured by the one or more cameras upon receiving a second media capture request.

[0026] In some embodiments, a computer program product is described that includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras. The one or more programs include instructions for performing the following operations: displaying a camera capture user interface via the display generation component and in a mixed reality environment, the camera capture user interface overlaid on a portion of a physical environment visible to a user of the computer system, wherein: the camera capture user interface is in a first mode; and when in the first mode, the camera capture user interface includes a set of one or more framed virtual objects, the set of one or more framed virtual objects being viewpoint-locked and indicating a first sub-portion of the physical environment to be captured by the one or more cameras upon receiving a first media capture request; while displaying the camera capture user interface in the first mode, receiving a request to transition the camera capture user interface to a second mode different from the first mode; and in response to receiving the request to transition the camera capture user interface to the second mode, displaying the camera capture user interface in the second mode, wherein: when in the second mode, the camera capture user interface includes a first representation of a field of view of at least a first camera of the one or more cameras; and the first representation is overlaid on a second sub-portion of the physical environment to be captured by the one or more cameras upon receiving a second media capture request.

[0027] In some embodiments, a method performed at a computer system in communication with one or more gaze tracking sensors and a display generation component is described. The method includes: displaying a corresponding user interface via the display generation component, wherein displaying the corresponding user interface includes: displaying a group of one or more virtual objects, the group of one or more virtual objects including a first virtual object displayed at a first position within a displayable area in which the display generation component is capable of displaying content; while displaying the first virtual object at the first position within the displayable area, detecting via the one or more gaze tracking sensors that a gaze of a user of the computer system is directed toward the first virtual object; in response to detecting that the gaze of the user of the computer system is directed toward the first virtual object, moving the first virtual object from the first position within the displayable area toward a second position within the displayable area that is different from the first position; while moving the first virtual object toward the second position within the displayable area and before the first virtual object reaches the second position, detecting movement of the gaze via the one or more gaze tracking sensors; and in response to detecting the movement of the gaze: continuing to move the first virtual object toward the second position based on determining that the gaze of the user of the computer system continues to be directed toward the first virtual object; and stopping moving the first virtual object toward the second position within the displayable area based on determining that the gaze of the user of the computer system has stopped being directed toward the first virtual object.

[0028] In some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more gaze tracking sensors and a display generation component, the one or more programs including instructions for the following operations: displaying a corresponding user interface via the display generation component, wherein displaying the corresponding user interface includes: displaying a group of one or more virtual objects, the group of one or more virtual objects including a first virtual object displayed at a first position within a displayable area in which the display generation component can display content; when displaying the first virtual object at the first position within the displayable area, detecting via the one or more gaze tracking sensors that a user of the computer system is looking at the first virtual object; in response to detecting that the computer system The user's gaze is directed toward the first virtual object, moving the first virtual object from the first position within the displayable area toward a second position within the displayable area that is different from the first position; when the first virtual object is moved toward the second position within the displayable area and before the first virtual object reaches the second position, detecting the movement of the gaze via the one or more gaze tracking sensors; and in response to detecting the movement of the gaze: based on determining that the gaze of the user of the computer system continues to be directed toward the first virtual object, continue moving the first virtual object toward the second position; and based on determining that the gaze of the user of the computer system has stopped being directed toward the first virtual object, stop moving the first virtual object toward the second position within the displayable area.

[0029] In some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more gaze tracking sensors and a display generation component, the one or more programs including instructions for the following operations: displaying a corresponding user interface via the display generation component, wherein displaying the corresponding user interface includes: displaying a group of one or more virtual objects, the group of one or more virtual objects including a first virtual object displayed at a first position within a displayable area in which the display generation component can display content; when displaying the first virtual object at the first position within the displayable area, detecting via the one or more gaze tracking sensors that a user of the computer system is looking at the first virtual object; in response to detecting that the computer system is looking at the first virtual object, The user's gaze is directed toward the first virtual object, and the first virtual object is moved from the first position within the displayable area toward a second position within the displayable area that is different from the first position; when the first virtual object is moved toward the second position within the displayable area and before the first virtual object reaches the second position, the one or more gaze tracking sensors detect the movement of the gaze; and in response to detecting the movement of the gaze: based on determining that the gaze of the user of the computer system continues to be directed toward the first virtual object, continue to move the first virtual object toward the second position; and based on determining that the gaze of the user of the computer system has stopped being directed toward the first virtual object, stop moving the first virtual object toward the second position within the displayable area.

[0030] In some embodiments, a computer system configured to communicate with one or more gaze tracking sensors and a display generation component is described. The computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for performing the following operations: displaying, via the display generation component, a corresponding user interface, wherein displaying the corresponding user interface includes: displaying a group of one or more virtual objects, the group of one or more virtual objects including a first virtual object displayed at a first position within a displayable area within which the display generation component is capable of displaying content; while displaying the first virtual object at the first position within the displayable area, detecting, via the one or more gaze tracking sensors, that a gaze of a user of the computer system is directed toward the first virtual object; in response to detecting that the gaze of the user of the computer system is directed toward the first virtual object, moving the first virtual object from the first position within the displayable area toward a second position within the displayable area that is different from the first position; while moving the first virtual object toward the second position within the displayable area and before the first virtual object reaches the second position, detecting, via the one or more gaze tracking sensors, movement of the gaze; and in response to detecting the movement of the gaze: continuing to move the first virtual object toward the second position based on determining that the gaze of the user of the computer system continues to be directed toward the first virtual object; and ceasing to move the first virtual object toward the second position within the displayable area based on determining that the gaze of the user of the computer system has stopped being directed toward the first virtual object.

[0031] In some embodiments, a computer system is described. The computer system is configured to communicate with one or more gaze tracking sensors and a display generation component and includes: a component for displaying a corresponding user interface via the display generation component, wherein displaying the corresponding user interface includes: displaying a group of one or more virtual objects, the group of one or more virtual objects including a first virtual object displayed at a first position within a displayable area within which the display generation component is capable of displaying content; a component for detecting, via the one or more gaze tracking sensors, that a gaze of a user of the computer system is directed toward the first virtual object when the first virtual object is displayed at the first position within the displayable area; and a component for moving the first virtual object from the displayable area to the first virtual object in response to detecting that the gaze of the user of the computer system is directed toward the first virtual object. The computer system further comprises a component for moving the first virtual object from a first position within the display area toward a second position within the display area that is different from the first position; a component for detecting the movement of the gaze via the one or more gaze tracking sensors when the first virtual object is moved toward the second position within the display area and before the first virtual object reaches the second position; and a component for performing the following operations in response to detecting the movement of the gaze: continuing to move the first virtual object toward the second position based on determining that the gaze of the user of the computer system continues to be directed toward the first virtual object; and stopping moving the first virtual object toward the second position within the display area based on determining that the gaze of the user of the computer system has stopped being directed toward the first virtual object.

[0032] In some embodiments, a computer program product is described that includes one or more programs configured to be executed by one or more processors of a computer system in communication with one or more gaze tracking sensors and a display generation component. The one or more programs include instructions for performing the following operations: displaying, via the display generation component, a corresponding user interface, wherein displaying the corresponding user interface includes: displaying a group of one or more virtual objects, the group of one or more virtual objects including a first virtual object displayed at a first position within a displayable area within which the display generation component is capable of displaying content; while displaying the first virtual object at the first position within the displayable area, detecting, via the one or more gaze tracking sensors, that a gaze of a user of the computer system is directed toward the first virtual object; in response to detecting that the gaze of the user of the computer system is directed toward the first virtual object, moving the first virtual object from the first position within the displayable area toward a second position within the displayable area that is different from the first position; while moving the first virtual object toward the second position within the displayable area and before the first virtual object reaches the second position, detecting, via the one or more gaze tracking sensors, movement of the gaze; and in response to detecting the movement of the gaze: continuing to move the first virtual object toward the second position based on determining that the gaze of the user of the computer system continues to be directed toward the first virtual object; and ceasing to move the first virtual object toward the second position within the displayable area based on determining that the gaze of the user of the computer system has stopped being directed toward the first virtual object.

[0033] It should be noted that the various embodiments described above can be combined with any other embodiment described herein. The features and advantages described in this specification are not comprehensive. In particular, many additional features and advantages will be apparent to those skilled in the art from the drawings, the description, and the claims. In addition, it should be noted that the language used in this specification has been selected in principle for readability and instructional purposes, and may not be selected to describe or define the subject matter of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] For a better understanding of the various described embodiments, reference should be made to the following detailed description taken in conjunction with the following drawings, wherein like reference numerals designate corresponding parts throughout the several views.

[0035] Figure 1 is a block diagram illustrating an operating environment for a computer system for providing an XR experience according to some embodiments.

[0036] Figure 2 is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's XR experience according to some embodiments.

[0037] Figure 3 is a block diagram illustrating display generation components of a computer system configured to provide a visual component of an XR experience to a user according to some embodiments.

[0038] Figure 4 is a block diagram illustrating a hand tracking unit of a computer system configured to capture gesture input from a user according to some embodiments.

[0039] Figure 5 is a block diagram illustrating an eye tracking unit of a computer system configured to capture gaze input from a user according to some embodiments.

[0040] Figure 6 is a flowchart illustrating a flash-assisted gaze tracking pipeline according to some embodiments.

[0041] Figures 7A to 7K Example techniques for interacting with a graphical user interface using gaze are shown in accordance with some embodiments.

[0042] Figures 8A to 8B is a flow chart of a method for interacting with a graphical user interface using gaze, according to various embodiments.

[0043] Figures 9A to 9E Example techniques for repositioning virtual objects are shown according to some embodiments.

[0044] Figure 10 is a flow chart of a method for repositioning a virtual object according to various embodiments.

[0045] Figures 11A to 11I Example techniques for transitioning modes of a camera capture user interface are shown according to some embodiments.

[0046] Figure 12 is a flow chart of a method for transitioning a mode of a camera capture user interface according to various embodiments.

[0047] Figures 13A to 13K Example techniques for interacting with a graphical user interface using gaze are shown in accordance with some embodiments.

[0048] Figure 14 is a flow chart of a method for interacting with a graphical user interface using gaze, according to various embodiments. DETAILED DESCRIPTION

[0049] According to some embodiments, the present disclosure relates to a user interface for providing an extended reality (XR) experience to a user.

[0050] Figures 1 to 6A description of an example computer system for providing an XR experience to a user is provided. Figures 7A to 7K Example techniques for interacting with a graphical user interface using gaze are shown in accordance with some embodiments. Figures 8A to 8B is a flow chart of a method for interacting with a graphical user interface using gaze, according to various embodiments. Figures 7A to 7K The user interface in Figures 8A to 8B in the process. Figures 9A to 9E Example techniques for repositioning virtual objects are shown according to some embodiments. Figure 10 is a flow chart of a method for repositioning a virtual object according to various embodiments. Figures 9A to 9E The user interface in Figure 10 in the process. Figures 11A to 11I Example techniques for transitioning modes of a camera capture user interface are shown according to some embodiments. Figure 12 is a flow chart of a method for transitioning a mode of a camera capture user interface according to various embodiments. Figures 11A to 11I The user interface is used to show Figure 12 in the process. Figures 13A to 13K Example techniques for interacting with a graphical user interface using gaze are shown in accordance with some embodiments. Figure 14 is a flow chart of a method for interacting with a graphical user interface using gaze, according to various embodiments. Figures 13A to 13K The user interface is used to show Figure 14 in the process.

[0051] The processes described below enhance the operability of the device and make the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device) through various techniques, including by providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional display controls, performing an operation without further user input when a set of conditions have been met, improving privacy and / or security, providing a richer, more detailed, and / or more realistic user experience while saving storage space, and / or additional technologies. These techniques also reduce power usage and extend the battery life of the device by enabling the user to use the device faster and more efficiently. Saving battery power, and therefore weight, improves the ergonomics of the device. These techniques also enable real-time communication, allow the use of fewer and / or less accurate sensors, thereby resulting in a more compact, lighter, and less expensive device, and enable the device to be used in a variety of lighting conditions. These techniques reduce energy usage and thereby reduce the heat emitted by the device, which is particularly important for wearable devices where if the device generates too much heat well within the operating parameters of the device components, it may become uncomfortable for the user to wear the device.

[0052] In addition, in the method described herein where one or more steps depend on having met one or more conditions, it should be understood that the method can be repeated in multiple repetitions so that in the process of repetition, all conditions of the steps in the method of determining the method have been met in different repetitions of the method. For example, if the method needs to perform the first step (if the condition is met), and perform the second step (if the condition is not met), then those of ordinary skill will know that the steps stated are repeated until both the condition is met and the condition is not met (in no particular order). Therefore, the method described as having one or more steps depending on having met one or more conditions can be rewritten as a method of repeating until each condition described in the method is met. However, this does not require a system or computer-readable medium to declare that the system or computer-readable medium includes instructions for performing a contingent operation based on the satisfaction of the corresponding one or more conditions, and is therefore able to determine whether a possible situation has been met without explicitly repeating the steps of the method until all conditions of the steps in the method of determining the method have been met. Those of ordinary skill in the art will also understand that, similar to the method with a contingent step, a system or computer-readable storage medium can repeat the steps of the method as needed multiple times to ensure that all contingent steps have been performed.

[0053] In some embodiments, as Figure 1As shown in FIG, an XR experience is provided to a user via an operating environment 100 including a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a position sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., a home appliance, a wearable device, etc.). In some embodiments, one or more of the input device 125, the output device 155, the sensor 190, and the peripheral device 195 are integrated with the display generation component 120 (e.g., in a head-mounted device or a handheld device).

[0054] When describing an XR experience, various terms are used to distinctly refer to several related but distinct environments that a user can sense and / or interact with (e.g., using inputs detected by the computer system 101 generating the XR experience, which inputs cause the computer system generating the XR experience to generate audio, visual, and / or haptic feedback corresponding to the various inputs provided to the computer system 101). The following is a subset of these terms:

[0055] Physical Environment: The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. A physical environment, such as a physical park, includes physical objects, such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment, such as through sight, touch, hearing, taste, and smell.

[0056] Extended Reality: In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic systems. In XR, a subset of a person's physical movements, or representations thereof, is tracked, and in response, one or more properties of one or more virtual objects simulated in the XR environment are adjusted in a manner consistent with at least one law of physics. For example, an XR system may detect a person's head rotation and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds change in a physical environment. In some cases (e.g., for accessibility reasons), adjustments to the properties of virtual objects in the XR environment may be made in response to representations of physical movement (e.g., voice commands). People can sense and / or interact with XR objects using any of their senses, including vision, hearing, touch, taste, and smell. For example, people can sense and / or interact with audio objects, which create a 3D or spatial audio environment that provides the perception of point audio sources in 3D space. As another example, audio objects can enable audio transparency, which selectively introduces ambient sounds from the physical environment with or without computer-generated audio. In some XR environments, people can sense and / or interact only with audio objects.

[0057] Examples of XR include virtual reality and mixed reality.

[0058] Virtual Reality: A virtual reality (VR) environment is a simulated environment designed to be based entirely on computer-generated sensory input to one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, trees, buildings, and computer-generated images representing human avatars are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment through the simulation of their presence within the computer-generated environment and / or through the simulation of a subset of their physical movement within the computer-generated environment.

[0059] Mixed Reality: In contrast to VR environments, which are designed to be based entirely on computer-generated sensory input, a mixed reality (MR) environment refers to a simulated environment that is designed to include sensory input from the physical environment, or representations thereof, in addition to computer-generated sensory input (e.g., virtual objects). On the virtuality continuum, a mixed reality environment is anything between, but not including, a fully physical environment at one end and a virtual reality environment at the other. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Additionally, some electronic systems used to render MR environments may track position and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical items from the physical environment, or representations thereof). For example, the system may cause motion so that virtual trees appear stationary relative to the physical ground.

[0060] Examples of mixed reality include augmented reality and augmented virtuality.

[0061] Augmented Reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on a physical environment or a representation of a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on a transparent or translucent display so that a person uses the system to perceive the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system combines the images or videos with the virtual objects and presents the combination on the opaque display. A person uses the system to indirectly view the physical environment via the images or videos of the physical environment and perceives the virtual objects superimposed on the physical environment. As used herein, a video of the physical environment displayed on an opaque display is referred to as "transparent video," meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Further alternatively, the system may have a projection system that projects virtual objects into a physical environment, for example as holograms or on a physical surface, so that a person using the system can perceive the virtual objects superimposed on the physical environment. An augmented reality environment also refers to a simulated environment in which a representation of a physical environment is transformed by computer-generated sensory information. For example, in providing a pass-through video, the system may transform one or more sensor images to apply a selected perspective (e.g., a viewpoint) that is different from the perspective captured by the imaging sensor. For another example, the representation of the physical environment may be transformed by graphically modifying (e.g., enlarging) a portion thereof so that the modified portion may be a representative but not real version of the original captured image. For another example, the representation of the physical environment may be transformed by graphically eliminating a portion thereof or blurring a portion thereof.

[0062] Augmented Virtual: An augmented virtual (AV) environment is a simulated environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from the physical environment. The sensory inputs can be representations of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, but human faces are realistically reproduced from images of physical people. In another example, a virtual object may adopt the shape or color of a physical object imaged by one or more imaging sensors. In another example, a virtual object may adopt a shadow that matches the sun's position in the physical environment.

[0063] Viewpoint-locked virtual objects: When a computer system displays a virtual object at the same position and / or location in a user's viewpoint, even if the user's viewpoint shifts (e.g., changes), the virtual object is viewpoint-locked. In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the forward direction of the user's head (e.g., when the user is looking straight ahead, the user's viewpoint is at least a portion of the user's field of view); thus, without moving the user's head, the user's viewpoint remains fixed even when the user's gaze shifts. In embodiments where the computer system has a display generation component (e.g., a display screen) that is repositionable relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the display generation component of the computer system. For example, a viewpoint-locked virtual object that is displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's viewpoint even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the position and / or location of the viewpoint-locked virtual object displayed in the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, such that the virtual object is also referred to as a "head-locked virtual object."

[0064] Environment-locked visual objects: A virtual object is environment-locked (alternatively, "world-locked") when a computer system displays it at a location and / or position in a user's viewpoint that is based on (e.g., selected with reference to and / or anchored to) a location and / or object in a three-dimensional environment (e.g., a physical environment or a virtual environment). As the user's viewpoint shifts, the location and / or objects in the environment change relative to the user's viewpoint, which causes the environment-locked virtual object to be displayed at a different location and / or position in the user's viewpoint. For example, an environment-locked virtual object locked to a tree immediately in front of the user is displayed at the center of the user's viewpoint. When the user's viewpoint shifts to the right (e.g., the user's head turns to the right) such that the tree is now to the left of center in the user's viewpoint (e.g., the tree's position in the user's viewpoint shifts), the environment-locked virtual object locked to the tree is displayed to the left of center in the user's viewpoint. In other words, the position and / or location at which an environment-locked virtual object is displayed in the user's viewpoint depends on the position and / or orientation of the object in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system anchored to fixed locations and / or objects in the physical environment) to determine the location at which an environment-locked virtual object is displayed in the user's viewpoint. An environment-locked virtual object can be locked to a stationary portion of the environment (e.g., a floor, wall, table, or other stationary object), or can be locked to a movable portion of the environment (e.g., a vehicle, animal, person, or even a representation of a part of the user's body that moves independently of the user's viewpoint, such as a hand, wrist, arm, or foot of the user) so that the virtual object moves as the viewpoint or that portion of the environment moves to maintain a fixed relationship between the virtual object and that portion of the environment.

[0065] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits an inertial following behavior that reduces or delays the movement of the environment-locked or viewpoint-locked virtual object relative to the movement of a reference point that the virtual object follows. In some embodiments, when exhibiting inertial following behavior, the computer system intentionally delays the movement of the virtual object when movement of a reference point (e.g., a portion of the environment, a viewpoint, or a point fixed relative to the viewpoint, such as a point between 5 cm and 300 cm from the viewpoint) that the virtual object is following is detected. For example, when the reference point (e.g., a portion of the environment or a viewpoint) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point, but at a second speed that is slower than the first speed (e.g., until the reference point stops moving or slows down, at which point the virtual object begins to catch up with the reference point). In some embodiments, when the virtual object exhibits inertial following behavior, the device ignores small amounts of movement of the reference point (e.g., ignoring movements of the reference point below a threshold movement amount, such as movement of 0 to 5 degrees or movement of 0 to 50 cm). For example, when a reference point (e.g., a portion or viewpoint of an environment to which a virtual object is locked) moves a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed so as to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and when the reference point (e.g., the portion or viewpoint of the environment to which the virtual object is locked) moves a second amount greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is being displayed so as to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and then decreases when the amount of movement of the reference point increases above a threshold (e.g., a “lazy follow” threshold) because the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point. In some embodiments, maintaining a substantially fixed position of the virtual object relative to the reference point includes displaying the virtual object within a threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 15 cm, 20 cm, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward relative to the position of the reference point).

[0066] Hardware: There are many different types of electronic systems that enable people to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablet devices, and desktop / laptop computers. A head-mounted system may include speakers and / or other audio output devices integrated into the head-mounted system for providing audio output. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system may have a transparent or translucent display instead of an opaque display. A transparent or translucent display may have a medium through which light representing an image is directed to a person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light sources, or any combination of these technologies. The medium may be an optical waveguide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to selectively become opaque. Projection-based systems may employ retinal projection technology that projects graphic images onto a person's retina. The projection system may also be configured to project virtual objects into a physical environment, such as as a hologram or on a physical surface. In some embodiments, the controller 110 is configured to manage and coordinate the user's XR experience. In some embodiments, the controller 110 includes a suitable combination of software, firmware, and / or hardware. Figure 2Controller 110 is described in more detail. In some embodiments, controller 110 is a computing device that is located locally or remotely relative to scene 105 (e.g., physical environment). For example, controller 110 is a local server located within scene 105. In another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside of scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., HMD, display, projector, touch screen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE802.3x, etc.). In another example, the controller 110 is included within a housing (e.g., a physical housing) of the display generation component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors, etc.), one or more input devices of the input devices 125, one or more output devices of the output devices 155, one or more sensors of the sensors 190, and / or one or more peripheral devices 195, or shares the same physical housing or support structure with one or more of the above devices.

[0067] In some embodiments, the display generation component 120 is configured to provide an XR experience (e.g., at least the visual component of the XR experience) to the user. In some embodiments, the display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Figure 3 Display generation component 120 is described in further detail. In some embodiments, the functionality of controller 110 is provided by and / or combined with display generation component 120.

[0068] According to some embodiments, the display generation component 120 provides an XR experience to the user while the user is virtually and / or physically present within the scene 105.

[0069] In some embodiments, the display generation component is worn on a part of the user's body (e.g., on his / her head, on his / her hand, etc.). In this way, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smart phone or tablet device) configured to present XR content, and the user holds a device with a display facing the user's field of view and a camera facing the scene 105. In some embodiments, the handheld device is optionally placed in a housing worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is an XR room, housing, or room configured to present XR content, wherein the user does not wear or hold the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface illustrating interactions with XR content that are triggered based on interactions occurring in the space in front of a handheld device or a tripod-mounted device may similarly be implemented with an HMD, where the interactions occur in the space in front of the HMD and responses to the XR content are displayed via the HMD. Similarly, a user interface illustrating interactions with XR content that are triggered based on movement of a handheld device or a tripod-mounted device relative to a physical environment (e.g., scene 105 or a part of a user's body (e.g., the user's eyes, head, or hands)) may similarly be implemented with an HMD, where the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of a user's body (e.g., the user's eyes, head, or hands)).

[0070] Despite Figure 1 Relevant features of the operating environment 100 are shown, but those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the example embodiments disclosed herein.

[0071] Figure 2is a block diagram of an example of a controller 110 according to some embodiments. While some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the embodiments disclosed herein. To this end, as a non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a central processing unit (CPU), a processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., a universal serial bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), infrared (IR), Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, a memory 220, and one or more communication buses 204 for interconnecting these and various other components.

[0072] In some embodiments, the one or more communication buses 204 include circuits that interconnect and control communications between system components. In some embodiments, the one or more I / O devices 206 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, and the like.

[0073] Memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located away from one or more processing units 202. Memory 220 includes non-transitory computer-readable storage media. In some embodiments, memory 220 or a non-transitory computer-readable storage medium of memory 220 stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 230 and an XR experience module 240.

[0074] The operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate single or multiple XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for corresponding groups of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.

[0075] In some embodiments, the data acquisition unit 241 is configured to Figure 1 1 and / or peripherals 195. The display generation component 120 of the embodiment of the present invention may be configured to generate a display image of the display image, and optionally acquire data (e.g., presentation data, interaction data, sensor data, position data, etc.) from one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for instructions, as well as heuristics and metadata for the heuristics.

[0076] In some embodiments, the tracking unit 242 is configured to map the scene 105 and track at least the display generation component 120 relative to the scene 105. Figure 1 105, and optionally tracks the position / location of one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the position / location of one or more parts of the user's hand and / or the position of one or more parts of the user's hand relative to the user's hand. Figure 1 The movement of the scene 105 relative to the display generation component 120 and / or relative to the coordinate system (the coordinate system is defined relative to the user's hand). Figure 4 The hand tracking unit 244 is described in more detail. In some embodiments, the eye tracking unit 243 is configured to track the position or movement of the user's gaze (or more broadly, the user's eyes, face, or head) relative to the scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hands)) or relative to the XR content displayed via the display generation component 120. Figure 5 The eye tracking unit 243 is described in more detail.

[0077] In some embodiments, the coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by the display generation component 120, and optionally by one or more of the output device 155 and / or peripheral devices 195. To this end, in various embodiments, the coordination unit 246 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0078] In some embodiments, the data sending unit 248 is configured to send data (e.g., presentation data, position data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data sending unit 248 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0079] Although the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 are shown as residing on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 may be located in separate computing devices.

[0080] also, Figure 2 It serves more as a functional description of various features that may be present in a particular implementation, rather than as a structural diagram of the embodiments described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 2 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the division of specific functions and how features are distributed among them will vary depending on the specific implementation and, in some embodiments, will depend in part on the specific combination of hardware, software, and / or firmware selected for a particular implementation.

[0081] Figure 3is a block diagram of an example of a display generation component 120 according to some embodiments. While some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the embodiments disclosed herein. For this purpose, as a non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., a microprocessor, an ASIC, an FPGA, a GPU, a CPU, a processing core, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional internal-facing and / or external-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.

[0082] In some embodiments, one or more communication buses 304 include circuits for interconnecting and controlling communications between various system components. In some embodiments, one or more I / O devices and sensors 306 include an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, and / or one or more depth sensors (e.g., structured light, time of flight, etc.), etc.

[0083] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to the user. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field effect transistor (OLET), organic light-emitting diode (OLED), surface conduction electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical system (MEMS) and / or similar display types. In some embodiments, one or more XR displays 312 correspond to diffraction, reflection, polarization, holographic and other waveguide displays. For example, the display generation component 120 (e.g., HMD) includes a single XR display. In another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, one or more XR displays 312 are capable of presenting MR and VR content. In some embodiments, one or more XR displays 312 are capable of presenting MR or VR content.

[0084] In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to the user's hands and, optionally, at least a portion of the user's arms (and may be referred to as a hand-tracking camera). In some embodiments, the one or more image sensors 314 are configured to face forward so as to acquire image data corresponding to the scene that the user would see in the absence of the display generation component 120 (e.g., an HMD) (and may be referred to as a scene camera). The one or more optional image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal oxide semiconductor (CMOS) image sensor or a charge coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, among others.

[0085] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located away from one or more processing units 302. Memory 320 includes non-transitory computer-readable storage media. In some embodiments, memory 320 or a non-transitory computer-readable storage medium of memory 320 stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 330 and an XR rendering module 340.

[0086] The operating system 330 includes processes for handling various basic system services and for performing hardware-related tasks. In some embodiments, the XR rendering module 340 is configured to present XR content to the user via one or more XR displays 312. To this end, in various embodiments, the XR rendering module 340 includes a data acquisition unit 342, an XR rendering unit 344, an XR map generation unit 346, and a data transmission unit 348.

[0087] In some embodiments, the data acquisition unit 342 is configured to at least Figure 1 The controller 110 acquires data (e.g., presentation data, interaction data, sensor data, location data, etc.). For this purpose, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for instructions and heuristics and metadata for the heuristics.

[0088] In some embodiments, the XR rendering unit 344 is configured to render XR content via one or more XR displays 312. For such purposes, in various embodiments, the XR rendering unit 344 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.

[0089] In some embodiments, the XR map generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate an extended reality) based on the media content data. For this purpose, in various embodiments, the XR map generation unit 346 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.

[0090] In some embodiments, the data sending unit 348 is configured to send data (e.g., presentation data, position data, etc.) to at least the controller 110, and optionally one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For such purposes, in various embodiments, the data sending unit 348 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.

[0091] Although the data acquisition unit 342, the XR rendering unit 344, the XR map generation unit 346, and the data transmission unit 348 are shown as residing on a single device (e.g., Figure 1 , but it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR rendering unit 344, the XR map generation unit 346, and the data sending unit 348 may be located in a separate computing device.

[0092] also, Figure 3 It serves more as a functional description of various features that may be present in a particular implementation, rather than as a structural diagram of the embodiments described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 3 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the division of specific functions and how features are distributed among them will vary depending on the specific implementation and, in some embodiments, will depend in part on the specific combination of hardware, software, and / or firmware selected for a particular implementation.

[0093] Figure 4 is a schematic illustration of an example embodiment of a hand tracking device 140. In some embodiments, the hand tracking device 140 ( Figure 1 ) is controlled by the hand tracking unit 244 ( Figure 2 ) to track the position / location of one or more parts of the user's hand and / or one or more parts of the user's hand relative to Figure 1The hand tracking device 140 is configured to monitor movement of the scene 105 relative to the user's surroundings (e.g., relative to a portion of the physical environment surrounding the user, relative to the display generation component 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system defined relative to the user's hands). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0094] In some embodiments, hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures three-dimensional scene information, including at least a human user's hand 406. Image sensor 404 captures hand images at a sufficient resolution to enable the fingers and their respective positioning to be distinguished. Image sensor 404 typically captures images of other parts of the user's body, or may also capture images of all parts of the body, and may have zoom capabilities or specialized sensors with increased magnification to capture images of the hand at a desired resolution. In some embodiments, image sensor 404 also captures 2D color video images of hand 406 and other elements of the scene. In some embodiments, image sensor 404 is used in conjunction with other image sensors to capture the physical environment of scene 105, or as an image sensor to capture the physical environment of scene 105. In some embodiments, the image sensor is positioned relative to the user or the user's environment in such a way that the field of view of image sensor 404, or a portion thereof, is used to define an interaction space in which hand movements captured by the image sensor are treated as input to controller 110.

[0095] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D image data (and possibly color image data) to the controller 110, which extracts high-level information from the image data. This high-level information is typically provided via an application program interface (API) to an application running on the controller, which in turn drives the display generation component 120. For example, a user can interact with the software running on the controller 110 by moving his hand 406 and changing his hand pose.

[0096] In some embodiments, the image sensor 404 projects a speckled pattern onto a scene containing the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral offsets of the spots in the pattern. This approach is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. The method gives the depth coordinates of points in the scene relative to a predetermined reference plane at a specific distance from the image sensor 404. In the present disclosure, it is assumed that the image sensor 404 defines an orthogonal set of x-axis, y-axis, and z-axis such that the depth coordinates of points in the scene correspond to the z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) may use other 3D mapping methods, such as stereo imaging or time-of-flight measurement, based on a single or multiple cameras or other types of sensors.

[0097] In some embodiments, the hand tracking device 140 captures and processes a time series of depth maps containing the user's hand as the user moves his hand (e.g., the entire hand or one or more fingers). Software running on the image sensor 404 and / or the processor in the controller 110 processes the 3D map data to extract image patch descriptors of the hand in these depth maps. The software can match these descriptors with image patch descriptors stored in the database 408 based on a previous learning process to estimate the pose of the hand in each frame. The pose typically includes the 3D positions of the user's hand joints and fingertips.

[0098] The software can also analyze the trajectory of the hand and / or finger over multiple frames in the sequence to identify gestures. The pose estimation functionality described herein can be interleaved with the motion tracking functionality so that the patch-based pose estimation is performed only once every two (or more) frames, and tracking is used to find changes in pose that occur over the remaining frames. The pose, motion, and gesture information is provided to the application running on the controller 110 via the above-mentioned API. The application can, for example, move and modify the image presented on the display generation component 120 in response to the pose and / or gesture information, or perform other functions.

[0099] In some embodiments, gestures include air gestures. An air gesture is a gesture that is detected without the user touching an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140) (or independent of an input element that is part of the device) and is based on detected movement of a part of the user's body (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs) through air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of a user's finger relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture in which the hand moves a predetermined amount and / or speed in a predetermined posture, or a shake gesture including a predetermined speed or amount of rotation of a part of the user's body)).

[0100] In some embodiments, according to some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures for interacting with an XR environment (e.g., a virtual or mixed reality environment) performed by movement of a user's fingers relative to other fingers (or parts of the user's hands). In some embodiments, an air gesture is a gesture detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and based on detected movement of a part of the user's body through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of a user's finger relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture in which the hand moves a predetermined amount and / or speed in a predetermined posture, or a shake gesture in which a part of the user's body is rotated at a predetermined speed or amount)).

[0101] In some embodiments where the input gesture is an in-air gesture (e.g., in the absence of physical contact with an input device that provides information to the computer system about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touch screen, or contact with a mouse or trackpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Thus, in specific implementations involving in-air gestures, for example, the input gesture is combined (e.g., simultaneously) with movement of the user's fingers and / or hand to detect attention (e.g., gaze) toward a user interface element to perform a pinch and / or tap input, as described below.

[0102] In some embodiments, an input gesture directed to a user interface object is performed directly or indirectly with reference to the user interface object. For example, user input is performed directly on the user interface object based on performing input with the user's hand at a location corresponding to the location of the user interface object in the three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, while detecting the user's attention (e.g., gaze) to the user interface object, an input gesture is performed indirectly on the user interface object based on the user's hand being located not at the location corresponding to the location of the user interface object in the three-dimensional environment while the user performs the input gesture. For example, for a direct input gesture, the user can direct the user's input to the user interface object by initiating a gesture at or near a location corresponding to the displayed location of the user interface object (e.g., within 0.5 cm, 1 cm, 5 cm, or a distance between 0 and 5 cm measured from the outer edge of the option or the center portion of the option). For an indirect input gesture, the user can direct the user's input to the user interface object by focusing on the user interface object (e.g., by gazing at the user interface object), and while focusing on the option, the user initiates an input gesture (e.g., at any location detectable by the computer system) (e.g., at a location that does not correspond to the displayed location of the user interface object).

[0103] In some embodiments, according to some embodiments, input gestures (e.g., air gestures) used in various examples and embodiments described herein include pinch input and tap input for interacting with a virtual or mixed reality environment. For example, the pinch input and tap input described below are performed as air gestures.

[0104] In some embodiments, a pinch input is part of an air gesture that includes one or more of: a pinch gesture, a long pinch gesture, a pinch and drag gesture, or a double pinch gesture. For example, a pinch gesture as an air gesture includes movement of two or more fingers of a hand to contact each other, i.e., optionally followed by a break in contact with each other immediately (e.g., within 0 seconds to 1 second). A long pinch gesture as an air gesture includes movement of two or more fingers of a hand in contact with each other for at least a threshold amount of time (e.g., at least 1 second) before a break in contact with each other is detected. For example, a long pinch gesture includes the user maintaining a pinch gesture (e.g., in which the two or more fingers are in contact), and the long pinch gesture continues until a break in contact between the two or more fingers is detected. In some embodiments, a double pinch gesture as an air gesture includes two (e.g., more) pinch inputs (e.g., performed by the same hand) that are detected consecutively immediately (e.g., within a predefined time period) with respect to each other. For example, the user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., interrupts contact between two or more fingers), and performs a second pinch input within a predefined time period (e.g., within 1 second or within 2 seconds) after releasing the first pinch input.

[0105] In some embodiments, a pinch and drag gesture as an air gesture includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in conjunction with (e.g., following) a drag input that changes the position of the user's hand from a first position (e.g., the starting position of the drag) to a second position (e.g., the ending position of the drag). In some embodiments, the user maintains the pinch gesture while performing the drag input, and releases the pinch gesture (e.g., opens their two or more fingers) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., the user pinches two or more fingers to contact each other and moves the same hand to a second position in the air using a drag gesture). In some embodiments, the pinch input is performed by the user's first hand, and the drag input is performed by the user's second hand (e.g., the user's second hand moves in the air from the first position to the second position while the user continues the pinch input with the user's first hand). In some embodiments, an input gesture as an air gesture includes input performed using both hands of the user (e.g., a pinch and / or tap input). For example, the input gesture includes two (e.g., more) pinch inputs performed in conjunction with each other (e.g., simultaneously or within a predefined time period). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch and drag input) is performed using a first hand of a user, and a second pinch input is performed using another hand (e.g., a second hand of the user) in conjunction with the pinch input performed using the first hand. In some embodiments, movement between the user's two hands (e.g., increasing and / or decreasing the distance or relative orientation between the user's two hands) occurs.

[0106] In some embodiments, a tap input performed as an air gesture (e.g., pointing to a user interface element) includes movement of a user's finger toward the user interface element, movement of the user's hand toward the user interface element (optionally, extension of the user's finger toward the user interface element), a downward motion of the user's finger (e.g., mimicking a mouse click motion or a tap on a touch screen), or other predefined movement of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on movement characteristics of the finger or hand performing the tap gesture movement of the finger or hand, which is a movement of the finger or hand away from the user's viewpoint and / or toward an object that is the target of the tap input, followed by an end of the movement. In some embodiments, the end of the movement is detected based on a change in movement characteristics of the finger or hand performing the tap gesture (e.g., an end of movement away from the user's viewpoint and / or toward an object that is the target of the tap input, a reversal of the direction of movement of the finger or hand, and / or a reversal of the acceleration direction of the movement of the finger or hand).

[0107] In some embodiments, the user's attention is determined to be directed toward a portion of the three-dimensional environment based on detection of a gaze directed toward the portion of the three-dimensional environment (optionally, no other conditions are required). In some embodiments, the user's attention is determined to be directed toward a portion of the three-dimensional environment based on detection of a gaze directed toward the portion of the three-dimensional environment using one or more additional conditions, such as requiring the gaze to be directed toward the portion of the three-dimensional environment for at least a threshold duration (e.g., a dwell duration) and / or requiring the gaze to be directed toward the portion of the three-dimensional environment when the user's viewpoint is within a distance threshold from the portion of the three-dimensional environment, so that the device determines that the user's attention is directed toward the portion of the three-dimensional environment, wherein if one of these additional conditions is not met, the device determines that the attention is not directed toward the portion of the three-dimensional environment to which the gaze is directed (e.g., until the one or more additional conditions are met).

[0108] In some embodiments, the detection of a ready state configuration of a user or a portion of a user is detected by a computer system. The detection of a ready state configuration of a hand is used by the computer system as an indication that the user may be preparing to interact with the computer system using one or more air gesture inputs performed by the hand (e.g., a pinch, a tap, a pinch and drag, a double pinch, a long pinch, or other air gestures described herein). For example, the ready state of a hand is determined based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape with the thumb and one or more fingers extended and spaced apart in preparation for a pinch or grab gesture, or a pre-tap with one or more fingers extended and the palm facing away from the user), based on whether the hand is in a predetermined position relative to the user's viewpoint (e.g., below the user's head and above the user's waist and extending at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or based on whether the hand has moved in a particular manner (e.g., toward an area in front of the user above the user's waist and below the user's head, or away from the user's body or legs). In some embodiments, the ready state is used to determine whether an interactive element of a user interface responds to attention (e.g., gaze) input.

[0109] In scenarios where input is described with reference to in-air gestures, it should be understood that similar gestures can be detected using a hardware input device attached to or held by one or more hands of a user, where the positioning of the hardware input device in space can be tracked using optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units, and the positioning and / or movement of the hardware input device is used instead of the positioning and / or movement of the one or more hands in the corresponding in-air gesture. In scenarios where input is described with reference to in-air gestures, it should be understood that similar gestures can be detected using a hardware input device attached to or held by one or more hands of a user, and user input can be detected using controls contained in the hardware input device, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, one or more hand or finger overlays that can detect the position or change in position of parts of a hand and / or finger relative to each other, relative to the user's body, and / or relative to the user's physical environment, and / or other hardware input device controls, wherein user input performed using controls contained in the hardware input device is used in place of hand and / or finger gestures such as an air tap or air pinch in the corresponding in-air gesture. For example, a selection input described as being performed using an air tap or air pinch input may alternatively be detected using a button press, a tap on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input. As another example, movement input described as being performed using a mid-air pinch and drag may alternatively be detected based on interaction with a hardware input control, such as a button press and hold, a touch on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input following movement of a hardware input device (e.g., along with a hand associated with the hardware input device) through space. Similarly, two-handed input involving movement of hands relative to each other may be performed using one mid-air gesture and one hardware input device in the hand that is not performing the mid-air gesture, two hardware input devices held in different hands, or two mid-air gestures performed by different hands using various combinations of mid-air gestures and / or input detected by one or more of the aforementioned hardware input devices.

[0110] In some embodiments, the software may be downloaded to the controller 110 in electronic form, for example, over a network, or may alternatively be provided on tangible, non-transitory media such as optical, magnetic, or electronic memory media. In some embodiments, the database 408 is also stored in memory associated with the controller 110. Alternatively or in addition, some or all of the described functions of the computer may be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). Although in Figure 4, but some or all of the processing functions of the controller may be performed by a suitable microprocessor and software or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand tracking device) or other device associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television receiver, handheld device, or head-mounted device) or integrated with any other suitable computerized device (such as a game console or media player). The sensing functions of the image sensor 404 may also be integrated into a computer or other computerized device to be controlled by the sensor output.

[0111] Figure 4 Also included is a schematic diagram of a depth map 410 captured by the image sensor 404 according to some embodiments. As described above, the depth map includes a matrix of pixels with corresponding depth values. Pixels 412 corresponding to the hand 406 have been segmented from the background and wrist in the figure. The brightness of each pixel within the depth map 410 is inversely proportional to its depth value (i.e., the measured z distance from the image sensor 404), where shades of gray become darker with increasing depth. The controller 110 processes these depth values in order to identify and segment components of the image (i.e., a group of adjacent pixels) that have characteristics of a human hand. These characteristics may include, for example, overall size, shape, and motion from frame to frame in the depth map sequence.

[0112] Figure 4 Also schematically shown is a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406 according to some embodiments. Figure 4 , a hand skeleton 414 is superimposed on a hand background 416 that has been segmented from the original depth map. In some embodiments, key feature points of the hand and, optionally, on the wrist or arm connected to the hand (e.g., points corresponding to knuckles, finger tips, the center of the palm, the end of the hand connected to the wrist, etc.) are identified and located on the hand skeleton 414. In some embodiments, the controller 110 uses the position and movement of these key feature points over multiple image frames to determine a gesture performed by the hand or the current state of the hand according to some embodiments.

[0113] Figure 5 An eye tracking device 130 ( Figure 1 ). In some embodiments, the eye tracking device 130 is composed of an eye tracking unit 243 ( Figure 2) controls to track the position and movement of the user's gaze relative to the scene 105 or relative to the XR content displayed via the display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as headphones, helmet, goggles, or glasses) or a handheld device placed in a wearable frame, the head-mounted device includes both components for generating XR content for the user to view and components for tracking the user's gaze relative to the XR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a handheld device or an XR room, the eye tracking device 130 is optionally a device separate from the handheld device or the XR room. In some embodiments, the eye tracking device 130 is a head-mounted device or a part of the head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used in conjunction with a display generation component that is also head-mounted or a display generation component that is not head-mounted. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally used in conjunction with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally part of a non-head-mounted display generation component.

[0114] In some embodiments, the display generation component 120 uses a display mechanism (e.g., a left near-eye display panel and a right near-eye display panel) to display a frame including a left image and a right image in front of the user's eyes, thereby providing a 3D virtual view to the user. For example, the head-mounted display generation component may include a left optical lens and a right optical lens (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display and display virtual objects on the transparent or translucent display, through which the user can directly view the physical environment. In some embodiments, the display generation component projects the virtual objects into the physical environment. The virtual objects may, for example, be projected onto a physical surface or projected as a hologram, so that an individual using the system observes the virtual objects superimposed on the physical environment. In this case, separate display panels and image frames for the left and right eyes may not be required.

[0115] like Figure 5As shown in , in some embodiments, the eye tracking device 130 (e.g., a gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) or near infrared (NIR) camera), and an illumination source (e.g., an IR or NIR light source, such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera can be pointed at the user's eyes to receive IR or NIR light that the light source reflects directly from the eyes, or alternatively can be pointed at "hot" mirrors located between the user's eyes and the display panel, which reflect IR or NIR light from the eyes toward the eye tracking camera while allowing visible light to pass through. The eye tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes these images to generate gaze tracking information, and transmits the gaze tracking information to the controller 110. In some embodiments, both eyes of the user are tracked separately by corresponding eye tracking cameras and illumination sources. In some embodiments, only one eye of the user is tracked by corresponding eye tracking camera and illumination source.

[0116] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye tracking device for a specific operating environment 100, such as the 3D geometry and parameters of the LED, camera, thermal mirror (if present), eye lens, and display screen. The device-specific calibration process may be performed at a factory or another facility before the AR / VR equipment is delivered to the end user. The device-specific calibration process may be an automatic calibration process or a manual calibration process. According to some embodiments, the user-specific calibration process may include an estimation of eye parameters of a specific user, such as pupil position, fovea position, optical axis, visual axis, eye spacing, etc. According to some embodiments, once the device-specific parameters and user-specific parameters are determined for the eye tracking device 130, a flash-assisted method may be used to process the images captured by the eye tracking camera to determine the current visual axis and the user's gaze point relative to the display.

[0117] like Figure 5As shown in FIG, an eye tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system that includes at least one eye tracking camera 540 (e.g., an infrared (IR) or near infrared (NIR) camera) positioned on the side of the user's face on which eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source, such as an array or ring of NIR light emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye 592. The eye tracking camera 540 may be directed toward a mirror 550 (which reflects the IR or NIR light from the eye 592 while allowing visible light to pass) located between the user's eye 592 and a display 510 (e.g., a left display panel or a right display panel of a head-mounted display, or a display of a handheld device, a projector, etc.). Figure 5 ), or alternatively may be directed toward the user's eye 592 to receive reflected IR or NIR light from the eye 592 (e.g., as shown in the top portion of Figure 5 (as shown in the bottom portion of the ).

[0118] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames for left and right display panels) and provides the frames 562 to the display 510. The controller 110 uses the gaze tracking input 542 from the eye tracking camera 540 for various purposes, such as for processing the frames 562 for display. The controller 110 optionally estimates the user's gaze point on the display 510 based on the gaze tracking input 542 obtained from the eye tracking camera 540 using a flash-assisted method or other suitable method. The gaze point estimated from the gaze tracking input 542 is optionally used to determine the direction the user is currently looking.

[0119] The following describes several possible use cases for the user's current gaze direction and is not intended to be limiting. As an example use case, the controller 110 may render virtual content differently based on the determined direction of the user's gaze. For example, the controller 110 may generate virtual content at a higher resolution in the foveal region determined based on the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content within the view based at least in part on the user's current gaze direction. As another example, the controller may display specific virtual content within the view based at least in part on the user's current gaze direction. As another example use case in an AR application, the controller 110 may direct an external camera used to capture the physical environment of an XR experience to focus in the determined direction. The external camera's autofocus mechanism may then focus on an object or surface in the environment on the display 510 that the user is currently looking at. As another example use case, the eye lens 520 may be a focusable lens, and the controller may use gaze tracking information to adjust the focus of the eye lens 520 so that the virtual object the user is currently looking at has the appropriate vergence to match the convergence of the user's eye 592. The controller 110 may utilize the gaze tracking information to guide the eye lenses 520 to adjust focus so that nearby objects that the user is looking at appear at the correct distance.

[0120] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eye lenses (e.g., eye lenses 520), an eye tracking camera (e.g., eye tracking camera 540), and a light source (e.g., light source 530 (e.g., IR or NIR LED)). The light source emits light (e.g., IR or NIR light) toward the user's eyes 592. In some embodiments, the light sources can be arranged in a ring or circle around each of the lenses, such as Figure 5 In some embodiments, for example, eight light sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer light sources 530 can be used, and other arrangements and positions of the light sources 530 can be used.

[0121] In some embodiments, the display 510 emits light in the visible range and does not emit light in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. It should be noted that the positions and angles of the eye tracking cameras 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.

[0122] like Figure 5 The illustrated embodiments of the gaze tracking system may be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a user with a computer-generated reality, virtual reality, augmented reality, and / or enhanced virtual experience.

[0123] Figure 6 FIGURE 1 illustrates a flash-assisted gaze tracking pipeline according to some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a flash-assisted gaze tracking system (e.g., Figure 1 and Figure 5 The flash-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no." When in the tracking state, the flash-assisted gaze tracking system uses previous information from previous frames when analyzing the current frame to track the pupil outline and glint in the current frame. When not in the tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in the tracking state.

[0124] like Figure 6 As shown in , the gaze tracking camera can capture left and right images of the user's left and right eyes. The captured images are then input to the gaze tracking pipeline for processing starting at 610. As indicated by the arrow returning to element 600, the gaze tracking system can continue to capture images of the user's eyes at a rate of, for example, 60 to 120 frames per second. In some embodiments, each set of captured images can be input to the pipeline for processing. However, in some embodiments or under some conditions, not all captured frames are processed by the pipeline.

[0125] At 610, for the currently captured image, if the tracking status is yes, the method proceeds to element 640. At 610, if the tracking status is no, the image is analyzed to detect the user's pupil and glint in the image, as indicated at 620. At 630, if the pupil and glint are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eye.

[0126] At 640, if proceeding from element 610, the current frame is analyzed to track the pupil and glint based in part on previous information from the previous frame. At 640, if proceeding from element 630, the tracking state is initialized based on the pupil and glint detected in the current frame. The processing result at element 640 is checked to verify that the tracking or detection result can be trusted. For example, the result can be checked to determine whether the pupil and a sufficient number of glints were successfully tracked or detected in the current frame to perform gaze estimation. At 650, if the result is not likely to be trusted, at element 660, the tracking state is set to no, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is trustworthy, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's gaze point.

[0127] Figure 6 This is intended to be used as an example of an eye tracking technology that may be used for a particular implementation. As one of ordinary skill in the art will appreciate, according to various embodiments, other eye tracking technologies currently existing or developed in the future may be used in place of or in combination with the flash-assisted eye tracking technology described herein in the computer system 101 for providing an XR experience to a user.

[0128] In this disclosure, various input methods are described with respect to interaction with a computer system. When an example is provided using one input device or input method, and another example is provided using another input device or input method, it should be understood that each example is compatible with and optionally utilizes the input device or input method described with respect to the other example. Similarly, various output methods are described with respect to interaction with a computer system. When an example is provided using one output device or output method, and another example is provided using another output device or output method, it should be understood that each example is compatible with and optionally utilizes the output device or output method described with respect to the other example. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment through a computer system. When an example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, it should be understood that each example is compatible with and optionally utilizes the methods described with respect to the other example. Therefore, this disclosure discloses embodiments that are combinations of features from multiple examples, without necessarily listing all features of the embodiments in detail in the description of each example embodiment.

[0129] User interface and associated processes

[0130] Attention is now focused on embodiments of a user interface ("UI") and associated processes that may be implemented on a computer system, such as a portable multifunction device or a head-mounted device, in communication with a display generation component, one or more input devices, and optionally one or more cameras.

[0131] Figures 7A to 7K An example of a technique for interacting with a graphical user interface using gaze is shown. Figures 8A to 8B is a flow diagram of an exemplary method 800 for interacting with a graphical user interface using gaze. Figures 7A to 7K The user interface in the diagram is used to illustrate the process described below, including Figures 8A to 8B in the process.

[0132] exist Figure 7AAt , device 700 (a tablet) displays a user interface 702 overlaid on a representation of a physical environment 701 including a tree 701a. Device 700 includes a display 700a and buttons (e.g., mechanical buttons, solid-state buttons, and / or touch-sensitive buttons) 700b, 700c, and 700d. In some embodiments, device 700 includes one or more features of computer system 101, such as eye tracking device 130 and sensor 306, which may include an accelerometer for detecting movement of device 700. In some embodiments, device 700 is a head-mounted system or display (e.g., an HMD), and when operating device 700 as an HMD, device 700 detects and / or tracks the direction and / or location of the user's gaze. In such embodiments, the position of display 700a typically occupies a majority of the user's field of view and has a fixed orientation relative to the user's head. In such embodiments, the user can control certain operations of device 700 via the location of the user's gaze without having to manually contact device 700 (e.g., provide touch input). Control via the user's gaze can be particularly useful for an HMD because hardware elements of the HMD (e.g., buttons or touch-sensitive surfaces) may not be visible to the user while the user is wearing the HMD and / or may be difficult to operate due to their location and / or lack of visibility. In addition, doing so can allow the user to use his or her hands for other purposes (e.g., interacting with the physical environment and / or operating other devices). In some embodiments, when operating the HMD, the user's gaze is often near the center area of his or her field of view (e.g., the user generally looks forward). Placing control elements along edges and / or corners can reduce the occurrence of false positives because the user is looking at elements in the physical environment that are generally near the center of the user's field of view.

[0133] exist Figure 7A , the user interface 702 includes several virtual objects (e.g., user interactive virtual objects), including corner virtual objects 704a, 704b, 704c, and 704d, a top virtual object 706 including top sub-portion virtual objects 706a, 706b, and 706c, center virtual objects 708a, 708b, and 708c, and time 702a. In some embodiments, the representation of the physical environment 701 is a pass-through representation of the physical environment (e.g., optical or camera-based), and the user interface 702 is an XR interface. In some embodiments, the corner virtual objects 704a, 704b, 704c, and 704d correspond to various system controls, including but not limited to a settings menu virtual object, a gaze control mode virtual object, a display sleep mode virtual object, and a center virtual object mode control (e.g., the center virtual object mode control determines how to arrange the center virtual object (e.g., as Figure 7A The carousel shown in Figure 7H)). In some embodiments, the top sub-portion virtual objects 706a, 706b, and 706c of the top virtual object 706 correspond to expandable system controls, such as output volume, display brightness, and / or display contrast. In some embodiments, the top sub-portion virtual objects 706a, 706b, and 706c of the top virtual object 706 correspond to system information, such as alarm / notification status (e.g., whether do not disturb mode is enabled and / or counts of various notification types), battery status, and / or connection status. In some embodiments, the center virtual objects 708a, 708b, and 708c correspond to one or more applications, such as a camera application, a navigation application, or an exercise / fitness application. In some embodiments, the center virtual objects 708a, 708b, and 708c correspond to one or more XR experiences in which virtual objects are presented to the user along with a representation of the physical environment, such as a camera XR experience for capturing media while interacting with the physical environment, a navigation XR experience for navigating within a physical environment, or an exercise XR experience for performing exercises while being presented with a representation of the physical environment. In some embodiments, the appearance (e.g., color, boldness, transparency level, and / or orientation) of a virtual object of the user interface 702 indicates the state of an associated control (e.g., whether audio output is muted or whether gaze-based controls are enabled). In some embodiments, one or more of the virtual objects of the user interface 702 are viewpoint-locked (e.g., 706 and 704a-704c). In some embodiments, one or more of the virtual objects of the user interface 702 are viewpoint-locked (e.g., 708a-708c). In some embodiments, the user interface 702 includes a plurality of virtual objects. Figure 7A In some embodiments, additional virtual objects are displayed along one or more other edges of the display (just outside the top).

[0134] exist Figure 7AAt , the device 700 detects (e.g., via the eye tracking device 130) that the user's gaze is alternately or sequentially directed toward the corner virtual object 704b, the corner virtual object 704d, and the top sub-portion virtual object 706a, as indicated by gaze indications 710a, 710b, and 710, respectively. In the present disclosure, the gaze indication can or cannot be visually presented as part of the user interface (e.g., the device 700 may or may not present a visual indication of the location of the currently detected user's gaze). In some embodiments, detecting that the user's gaze is directed toward a virtual object includes determining that the user's gaze has remained on the virtual object for a predetermined period of time (e.g., 0.1 seconds, 0.25 seconds, 0.5 seconds, or 1 second). In some embodiments, the user can configure (e.g., via a settings menu) whether a visual indication of the user's gaze is displayed by the device 700.

[0135] exist Figure 7B At, in response to detecting that the user's gaze is directed toward the corner virtual object 704b (as indicated by Figure 7A ), the device 700 stops displaying the corner virtual objects 704a, 704c, and 704d and the top virtual object 706. Figure 7B , the center virtual objects 708a to 708c and time 702a continue to be displayed, but in some embodiments, those virtual objects are also stopped from being displayed. The device 700 also bolds the corner virtual object 704b to indicate that the device 700 has detected the user's gaze directed at the corner virtual object 704b. In some embodiments, the corner virtual object 704b is emphasized in other ways, such as changing color, changing size, or changing shape. Figures 7A to 7K In an embodiment of the present invention, the corner virtual object 704b is a system control for toggling whether gaze-based control is enabled for most of the user interface 702 (e.g., enabled for only gaze-based control), as discussed in more detail below. Figure 7A and Figure 7B At , gaze-based control is enabled for most of the user interface 702. Figure 7B At , the device 700 continues to detect that the user's gaze is directed toward the corner virtual object 704b, as indicated by the gaze indication 710d. In some embodiments, continuing to detect that the user's gaze is directed toward the virtual object includes: after the initial detection (e.g., after the user's gaze is directed toward the virtual object by Figure 7A In some embodiments, if the device 700 is in the state where it is in the state of Figure 7B When the user's gaze is detected to move away from the corner virtual object 704b, the device 700 returns the user interface 702 to Figure 7B In some embodiments, the device 700 does this only after the gaze has moved away from the corner virtual object 704b for a predetermined time (e.g., 0.1 seconds, 0.25 seconds, 0.5 seconds, or 1 second), without returning to the virtual object.

[0136] exist Figure 7C In response to detecting that the user's gaze continues to be directed toward the corner virtual object 704b, the device 700 disables gaze-based control for a majority of the user interface 702, modifies the appearance of the corner virtual object 702b to indicate that gaze-based control is now disabled for a majority of the user interface 702, and redisplays the corner virtual object 704b. Figure 7B Stop displaying other virtual objects at . Figures 7A to 7K In an embodiment, when gaze-based control for a majority of the user interface 702 is disabled, gaze alone cannot be used to activate functionality associated with virtual objects of the user interface 702, except for the corner virtual object 704b that can be gazed at to re-enable gaze-based control for a majority (e.g., the remainder) of the user interface 702.

[0137] exist Figure 7C At , the device 700 detects that the user's gaze is alternately or sequentially directed toward the corner virtual object 704d and the top sub-portion virtual object 706a, as indicated by gaze indications 710e and 710f, respectively. Figure 7C Because gaze-based control is now disabled for most of the user interface 702 (including the corner virtual object 704d and the top sub-portion virtual object 706a), when gaze-based control for most of the user interface 702 is enabled, the device 700 does not perform the operations that would be performed if the user's gaze was directed toward those virtual objects (e.g., for the corner virtual object 704d reference image). Figure 7G and Figure 7H The operations discussed and / or referenced with respect to the top sub-portion virtual object 706a Figure 7D and Figure 7E 706a). In some embodiments, when gaze-based control is disabled for a majority of the user interface 702, the user can still activate functionality associated with virtual objects (e.g., corner virtual object 704d and top sub-portion virtual object 706a) by directing his or her gaze at the virtual objects in combination with another input (e.g., gaze in combination with actuation of buttons 700b, 700c, and / or 700d, or in combination with performance of an air gesture (e.g., air pinch). In some embodiments, disabling gaze-based interaction for a majority of the user interface provides the user with a mode in which the user can gaze at most virtual objects for extended periods of time without activating undesired functionality.

[0138] exist Figure 7D In response to detecting that the user's gaze is directed toward the top sub-portion virtual object 706a, as indicated by Figure 7A As indicated by the gaze indication 710c in FIG, the device 700 stops displaying the corner virtual objects 704a to 704d. Figure 7C , the center virtual objects 708a to 708c and time 702a continue to be displayed, but in some embodiments, those virtual objects are also stopped from being displayed. The device 700 also bolds the top sub-portion virtual object 706a to indicate that the device 700 has detected the user's gaze directed at the top sub-portion virtual object 706a. Figures 7A to 7K In some embodiments, the top sub-portion virtual object 706a is associated with the battery capacity / level of the device 700. In some embodiments, the top sub-portion virtual object 706a is associated with the battery capacity / level of the device 700. 7A to 7B The state shown in FIG (e.g., the unexpanded state) includes a first set of battery-related information and / or controls. Figure 7D At , device 700 continues to detect that the user's gaze is directed toward top sub-portion virtual object 706a, as indicated by gaze indication 710g.

[0139] exist Figure 7E In response to detecting that the user's gaze continues to be directed toward the sub-portion virtual object 706a, the device 700 transitions the top sub-portion virtual object 706a to an expanded state (e.g., via a predetermined animation) and reduces the size of the top sub-portion virtual objects 706b and 706c so that the overall size of the top virtual object 706 remains the same. In some embodiments, the top sub-portion virtual object 706a is expanded without reducing the size of the top sub-portion virtual objects 706b and 706c. In some embodiments, reducing the size of the top sub-portion virtual object includes: ceasing to display one or more controls or information in the top sub-portion, or ceasing to display a particular top sub-portion entirely. In the expanded state, the top sub-portion virtual object 706a includes the top sub-portion virtual object 706a that would have been displayed when the top sub-portion virtual object 706a was in its unexpanded state (e.g., as Figure 7D In some embodiments, when the top sub-portion virtual object 706a is expanded, if the device 700 detects that the user's gaze is no longer directed at the top sub-portion virtual object 706a, the object transitions back to the unexpanded state, such as Figure 7D As shown in . Figure 7E, while the top sub-portion virtual object 706c is in the unexpanded state, the device 700 detects that the user's gaze is directed toward the object, as indicated by the gaze indication 710h. In some embodiments, while the top sub-portion virtual object 706a is in the expanded state, if the user's gaze remains directed toward the object while the top sub-portion virtual object 706a is in the expanded state, the device 700 expands the object to a further expanded state to display even more controls or information. Figure 7A exist Figure 7K In an embodiment, the top sub-portion virtual object 706c is associated with the audio output volume and includes a set of information related to the system volume (e.g., a numerical value of the current volume level).

[0140] exist Figure 7F In response to detecting that the user's gaze is directed toward the top sub-portion virtual object 706c, the device 700 transitions the top sub-portion virtual object 706c to an expanded state and reduces the size of the top sub-portion virtual objects 706a and 706b. In the expanded state, the top sub-portion virtual object 706c includes additional information (e.g., a graphical depiction of a volume level and / or a mute control relative to a maximum level) and / or controls (e.g., such as a mouse click) that are not present when the top sub-portion virtual object 706c is in its unexpanded state. Figure 7D and Figure 7E ). Therefore, in Figures 7A to 7K In an embodiment, a user can access additional information or controls in a sub-portion of the top virtual object 706 without consuming additional display area and without obscuring additional portions of the representation of the physical environment 701.

[0141] exist Figure 7G At, in response to detecting that the user's gaze is directed toward the corner virtual object 704d (as indicated by Figure 7A ), the device 700 stops displaying the corner virtual objects 704a, 704b, and 704c and the top virtual object 706. Figure 7G , the center virtual objects 708a to 708c and time 702a continue to be displayed, but in some embodiments, those virtual objects are also stopped from being displayed. The device 700 also bolds the corner virtual object 704b to indicate that the device 700 has detected the user's gaze directed at the corner virtual object 704b. In some embodiments, the corner virtual object 704b is emphasized in other ways, such as changing color, changing size, or changing shape. Figure 7G At , device 700 continues to detect that the user's gaze is directed toward corner virtual object 704d, as indicated by gaze indication 710d.

[0142] exist Figure 7HIn response to detecting that the user's gaze continues to be directed toward the corner virtual object 704d, the device 700 changes the layout of the displayed center virtual objects 708a to 708c from the carousel layout to the grid layout (including displaying the additional center virtual object 708d), modifies the appearance of the corner virtual object 702d to indicate that the grid layout is active, and redisplays the display at Figure 7G Stop displaying other virtual objects at that location.

[0143] exist Figure 7I At , the device 700 displays a user interface 712 overlaid on a representation of the physical environment 701 on the display 702. The user interface 712 includes central virtual objects 708a to 708c, as discussed with reference to the user interface 702. In some embodiments, the user interface 712 includes one or more other features and / or virtual objects of the user interface 702 discussed above. The user interface 712 includes a picture-in-picture ("PiP") virtual object 714, which includes a main portion 714a and a control virtual object 714b, the main portion including a representation of the user interface generated by the first application. In some embodiments, the PiP virtual object 714 corresponds to a video conferencing application, and the main portion 714a includes video content (e.g., a video of a user of the external device) sent from an external device that is currently in a video conferencing session with the device 700. In some embodiments, before displaying the video content as shown in FIG. Figure 7I Before the user interface 712 shown in FIG, the device 700 receives an indication of an incoming video conference request and displays a selectable notification on the display 700a. In such an embodiment, in response to selecting the notification, a user interface 712 shown in FIG. Figure 7I . In some embodiments, the PiP virtual object 714 corresponds to a media player application, and the main portion 714a includes media (e.g., a movie, a show, or music) being played back at the device 700. In some embodiments, while displaying the PiP virtual object 714, the device 700 (e.g., via the user interface 712) presents an XR experience, such as a camera XR experience for capturing media while interacting with a physical environment, a navigation XR experience for navigating within a physical environment, or an exercise XR experience for performing exercises while presented with a representation of the physical environment. In such embodiments, the PiP virtual object 714 may be displayed at a different depth than one or more virtual objects of the XR experience (e.g., perceived as being at a different distance from the user). In some embodiments, the XR experience is generated by an application that is different from the application that generated the PiP virtual object 714. Figure 7IAt 710 j, the device 700 detects input 716 a as an actuation of button 700 d while the device 700 is detecting that the user's gaze is directed toward the control virtual object 714 b of the PiP virtual object 714, as indicated by gaze indication 710 j. In some embodiments, the input 716 a is an air gesture, a touch on a touch-sensitive surface, or a speech input.

[0144] exist Figure 7J At , in response to detecting input 716 b while the device 700 is detecting that the user's gaze is directed toward the control virtual object 714 b of the PiP virtual object 714, the device 700 moves the PiP virtual object 714 from the lower right corner of the display 700 a to a predefined position at the upper right corner of the display 700. In some embodiments, when the device 700 detects that the user's gaze is directed toward the control virtual object 714 b without detecting an actuation of the button 700 d (in some embodiments, without detecting another additional input such as an air gesture, a touch on a touch-sensitive surface, or a speech input), the device 700 does not move the PiP virtual object 714 (e.g., because the user's gaze is typically directed toward the vicinity of the control virtual object 714 b due to the user viewing content within the PiP virtual object 714, resulting in a high false positive rate). In some embodiments, after moving the PiP virtual object 714 to the upper right corner, if the device 700 detects another actuation of the button 700d while the device 700 is detecting that the user's gaze is directed toward the control virtual object 714b of the PiP virtual object 714, the PiP virtual object 714 is moved to a predefined position at the upper left corner of the display 700a (and then, in some embodiments, to the lower left corner upon further actuation). In such embodiments, the user can position the PiP virtual object 714 at a desired corner position by cycling through the corners. Figure 7J At 710 k, the device 700 detects input 716 b as an actuation of button 700 d while the device 700 is detecting that the user's gaze is directed toward the main portion 714 a of the PiP virtual object 714, as indicated by gaze indication 710 k. In some embodiments, the input 716 b is an air gesture, a touch on a touch-sensitive surface, or a speech input.

[0145] exist Figure 7KIn response to detecting input 716b while the device 700 is detecting that the user's gaze is directed toward the main portion 714a of the PiP virtual object 714, the device 700 expands the PiP virtual object 714, displays the PiP virtual object 714 in the center of the display 700a, and displays additional PiP control virtual objects 714c to 714e. In some embodiments, the PiP control virtual objects 714c to 714e, when selected (e.g., via gaze, via gaze and hardware input, or gaze and air gesture), cause the device 700 to perform one or more functions associated with the PiP virtual object 714. For example, when the PiP virtual object 714 is associated with a video conferencing application, the device 700 may terminate the conference (e.g., and stop displaying the PiP virtual object 714), add another party to the conference, send content corresponding to a representation of the physical environment 701 to one or more other participants, stop sending video and / or audio from the device 700, or switch the content of the main portion 714a from another participant's view to the user's self-view of the device 700.

[0146] about Figures 7A to 7K See below for additional description of Figures 7A to 7K Method 800 is described.

[0147] Figures 8A to 8B is a flow chart of an exemplary method 800 for interacting with a graphical user interface using gaze, according to some embodiments. In some embodiments, the method 800 is performed on a computer system (e.g., Figure 1 In some embodiments, method 800 is performed at a computer system 101 in a computer system; a head-mounted display; an optical head-mounted display; a personal computer; a smart phone; and / or a tablet computer) that communicates with one or more gaze tracking sensors (e.g., an optical and / or IR camera configured to track the direction of gaze of a user of the computer system; an eye tracking device 130; and / or a sensor 306) and a display generation component (e.g., a display generation component 120; a display controller; a touch-sensitive display system; a pass-through display (e.g., integrated and / or connected), a 3D display, a transparent display, a projector, a head-up display, and / or a head-mounted display). In some embodiments, method 800 is performed by storing in a non-transitory (or transient) computer-readable storage medium and by one or more processors of a computer system (such as one or more processors 202 of computer system 101) (e.g., Figure 1 Some operations in method 800 may be optionally combined, and / or the order of some operations may be optionally changed.

[0148] A computer system (e.g., 700) displays (802) a corresponding user interface (e.g., 702) via a display generation component (e.g., 700a). In some embodiments, the corresponding user interface is a group of one or more virtual objects displayed in an extended reality environment. In some embodiments, at least one virtual object in the group of one or more virtual objects is a viewpoint-locked virtual object. Displaying the corresponding user interface includes displaying: a plurality of edges (804) of the corresponding user interface including a first edge (e.g., a first outer edge) and a second edge (e.g., a second outer edge) different from the first edge (e.g., the second outer edge) (e.g., outer edges and / or edges that intersect with one or more other edges (e.g., to form a corner of the corresponding user interface)) (in some embodiments, the plurality of edges define an outer boundary of the corresponding user interface); a first user interface object (806) (e.g., 706 and / or 704a) positioned along the first edge (e.g., the top edge of 702) (e.g., displayed adjacent to the first outer edge; displayed adjacent to a first corner formed by the first outer edge and another outer edge of the plurality of outer edges). to 704d) (e.g., gazing at a selectable object or other enabling representation); the first user interface object corresponds to a first operation (e.g., an operation associated with 706 and / or 704a to 704d) (e.g., an action to be performed by the computer system and / or affect (e.g., system control) the computer system; an operation that affects and / or modifies the corresponding user interface); and a second user interface object (808) (e.g., 706 and / or 704a to 704d) positioned along the second edge (gazing at a selectable object or other enabling representation), the second user interface object corresponding to a second operation different from the first operation (e.g., an operation associated with 706 and / or 704a to 704d).

[0149] While displaying the first user interface object and the second user interface object, the computer system detects (810) via the one or more gaze tracking sensors that the gaze of the user of the computer system is directed toward the corresponding portion (e.g., 710a to 710c) of the corresponding user interface (e.g., directed toward a direction corresponding to the gaze of the user that intersects with the corresponding portion) (in some embodiments, directed toward the first virtual object for at least a predetermined period of time (e.g., 0.25 seconds, 0.5 seconds, or 1 second)).

[0150] In response to detecting (812) that the gaze of the user of the computer system is directed toward the corresponding portion of the corresponding user interface and based on determining (814) that the corresponding portion of the corresponding user interface corresponds to (e.g., includes and / or overlaps with) (in some embodiments, determining that the gaze is directed toward a first gaze-selectable control object and / or toward the first outer edge of) the first user interface object (e.g., 704b), the computer system: performs (816) the first operation (e.g., as Figure 7B and continuing (818) displaying the first user interface object while ceasing to display the second user interface object (e.g., 706 and / or 704a, 704c, and 704d); in some embodiments, ceasing to display the second gaze-selectable control object while the gaze continues to be directed toward the respective portion. In some embodiments, re-displaying the second gaze-selectable control object once the gaze is no longer directed toward the respective portion. In some embodiments, ceasing to display a plurality of gaze-selectable control objects (including the second gaze-selectable object) that are each associated with a different outer edge and / or corner of the respective user interface.

[0151] In response to detecting (812) that the gaze of the user of the computer system is directed toward the corresponding portion of the corresponding user interface and based on determining (820) that the corresponding portion of the corresponding user interface corresponds to (e.g., includes and / or overlaps with) (in some embodiments, determining that the gaze is directed toward the second gaze-selectable control object and / or toward the second outer edge) the second user interface object (e.g., 704d), the computer system: performs (822) the second operation (e.g., Figure 7G); and continuing (824) displaying the second user interface object while ceasing to display the first user interface object (e.g., 706 and / or 704a to 704c); providing control objects at the edge of the user interface that can be activated by the user's gaze provides additional control options without cluttering the central area of the UI with additional display controls; for gaze-based interactions, doing so also reduces the risk of false positives because the central area of the UI presents a higher likelihood of false positives because the user's gaze tends to naturally remain near the center. Reducing the risk of false positives for user input schemes enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors in operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently. When the computer system detects that the user's gaze is directed to the selected user interface object, continuing to display the selected user interface object while ceasing to display another user interface object provides improved visual feedback regarding the detected direction of the user's gaze.

[0152] In some embodiments, the corresponding user interface is an extended reality user interface (e.g., as shown in FIG. Figure 7A and displaying the corresponding user interface includes displaying: a representation of the physical environment (e.g., 701) (e.g., an optical or video pass-through representation); and a third user interface object (e.g., a focus and / or face detection indicator, or a point of interest indicator), wherein the third user interface object is an environment-locked virtual object (in some embodiments, 708a to 708c are environment-locked, as in Figure 7A ) (e.g., locked to locations and / or objects in this representation of the physical environment).

[0153] In some embodiments, continuing to display the first user interface object while ceasing to display the second user interface object includes visually emphasizing the first user interface object (e.g., Figure 7B 704b in (e.g., bolding, enlarging, highlighting, animating, and / or brightening the first user interface object); and continuing to display the second user interface object while ceasing to display the first user interface object includes: visually emphasizing the second user interface object (e.g., Figure 7G 704d in). When the computer system detects that the user's gaze is directed toward the selected user interface object, visually emphasizing the selected user interface object while ceasing to display another user interface object provides improved visual feedback regarding the detected direction of the user's gaze.

[0154] In some embodiments, performing the first operation includes: expanding the first user interface object (e.g., 706) from an unexpanded state (e.g., a contracted state) to a first expanded state by expanding at least a first portion (e.g., 706a) of the first user interface object (e.g., a first sub-portion, or a portion corresponding to a control in a set of controls associated with the first user interface object), wherein the first expanded state of the first user interface object includes a first control object that is not included in the unexpanded state of the first user interface object (e.g., a selectable object corresponding to a control option (e.g., a volume control, a brightness control, or another control with a range of possible values)) and first information (e.g., textual information and / or graphical information associated with the first control object (e.g., a current volume level, a current brightness level, or another value in the range of possible values)) (e.g., as described in reference to Figure 7E Expanding the first user interface object to include the previously non-existent first control object and first information only when requested provides additional control options, conserves display area, and makes it easier to continue interacting with the now-expanded first user interface object via gaze (e.g., because the gaze target is now larger) without cluttering the UI with additional display controls. Doing so also provides improved visual feedback about the detected gaze direction of the user.

[0155] In some embodiments, the first control object (e.g., 706 and / or 706a), when selected, causes the computer system to perform operations related to display brightness (e.g., adjusting the value of display brightness, and / or causing additional brightness-related options to be displayed); and the first information corresponds to brightness information.

[0156] In some embodiments, the first control object (e.g., 706 and / or 706b), when selected (e.g., by gazing, by performing an air gesture while the object is in focus, by pressing a hardware button while the object is in focus), causes the computer system to perform an operation related to the audio output volume (e.g., adjusting the value of the output volume); and the first information corresponds to volume information (e.g., current volume level, mute status, audio output source, and / or input source).

[0157] In some embodiments, the first control object (e.g., 706 and / or 706c), when selected (e.g., by gazing, by performing an air gesture while the object is in focus, by pressing a hardware button while the object is in focus), causes the computer system to perform an operation related to an energy storage component (e.g., a battery) of the computer system (e.g., transitioning to a low power mode and / or causing additional battery information to be displayed); and the first information corresponds to energy storage (e.g., battery charge, power mode, and / or estimated remaining usage time) information.

[0158] In some embodiments, the computer system detects, via the one or more gaze tracking sensors, that the gaze of the user of the computer system is directed toward the first control object (e.g., 706 and / or 706a) (e.g., directed toward a direction corresponding to the gaze of the user that intersects the first control object) (in some embodiments, directed toward the first control object for at least a predetermined period of time (e.g., 0.25 seconds, 0.5 seconds, or 1 second)). In response to detecting that the gaze of the user of the computer system is directed toward the first control object, the computer system expands the first control object from a first control object unexpanded state to a first control object expanded state (e.g., as Figure 7D and Figure 7E ) (e.g., larger than the first control object unexpanded state and / or includes information and / or controls not included in the first control object unexpanded state) (in some embodiments, expanding the first control object from the first control object unexpanded state to the first control object expanded state includes: stopping displaying different control objects included in the first user interface (e.g., objects displayed at the position into which the first control object is expanded) before detecting that the gaze of the user of the computer system is directed toward the first control object). Expanding the first control object only when requested saves display area and makes it easier to continue interacting with the now expanded first control object via gaze (e.g., because the gaze target is now larger) without cluttering the UI with additional display controls. Doing so also provides improved visual feedback about the detected direction of the user's gaze.

[0159] In some embodiments, the first user interface object (e.g., 706) is in a non-expanded state of the first control object (e.g., Figure 7E 706c) in the first user interface object, the first user interface object includes a second control object ( Figure 7E706a in ), and expanding the first control object from the first control object unexpanded state to the first control object expanded state includes: contracting the second control object from the second control object expanded state to the second control object unexpanded state (e.g., Figures 7E to 7F Collapsing the second control object when the first control object is expanded reduces the risk of false positives by reducing the area of the second control object that may trigger a gaze interaction. Doing so also provides improved visual feedback about the detected gaze direction of the user.

[0160] In some embodiments, in the first control object expanded state (e.g., Figure 7E 706a), the first control object (e.g., 706) includes a third control object (e.g., a selectable object corresponding to a control option associated with the first control object (e.g., the first control object includes information about volume, and the third control object, when selected, causes a volume-related operation to be performed (e.g., the third control object is a volume slider)), which was not included in the first control object when the first control object was in the first control object's unexpanded state. Expanding the first control object to include the previously non-existent third control object when requested provides additional control options and conserves display area without cluttering the UI with additional displayed controls when not needed. Doing so also provides improved visual feedback about the detected gaze direction of the user.

[0161] In some embodiments, when in the expanded state of the first control object, the first control object (e.g., 706) includes second information (e.g., Figure 7E In some embodiments, the second information is not included in the first control object when the first control object is in an unexpanded state (e.g., text information and / or graphical information associated with the first control object and / or the third control object (e.g., current volume level or brightness level and / or battery information)). Expanding the first control object to include the second information that was not previously present when requested provides additional information and conserves display area without cluttering the UI with additional display information when it is not needed. Doing so also provides improved visual feedback regarding the detected gaze direction of the user.

[0162] In some embodiments, the first control object (e.g., Figure 7A704a) described above, when selected (e.g., by gazing, by performing an air gesture while the object is in focus, and / or by pressing a hardware button while the object is in focus), causes the display generation component to transition from a first mode (e.g., an active mode) to a second mode (e.g., a sleep mode; an inactive mode; an off mode; and / or a mode in which the display generation stops displaying content until the second mode is exited). In some embodiments, the first control object is located in a corner of the display area of the display generation component. Providing a control object that can cause the display generation component to transition from the first mode to the second mode provides an gaze-based option for transitioning display modes. In the case where the second mode is a low power mode, doing so also reduces power usage and extends the battery life of the device.

[0163] The first control object (e.g., 704b), when selected (e.g., by gazing, by performing an air gesture while the object is in focus, and / or by pressing a hardware button while the object is in focus), causes the computer system to disable (e.g., deactivate) a set of one or more functions that are activated by detecting the gaze of the user of the computer system. In some embodiments, the one or more functions may still be activated by non-gaze-based input (e.g., by an air gesture and / or by a button press). Providing a control object that causes the computer system to disable (e.g., deactivate) a set of one or more functions that are activated by detecting the gaze of the user of the computer system can reduce false positives because gaze-based interaction schemes can have a higher likelihood of false positives due to the natural flow of gaze when viewing displayed content.

[0164] In some embodiments, causing the computer system to disable the set of one or more functions that are activated by detecting the gaze of the user of the computer system includes: disabling a first function (e.g., function 704d), which is activated when the computer system detects that the gaze of the user of the computer system is directed to a first position of the corresponding user interface (e.g., a position at or near the center of the corresponding user interface); and maintaining a second function (e.g., function 704b) (e.g., a function for re-enabling the one or more functions activated by detecting the gaze of the user of the computer system) available for activation (e.g., available for activation via gaze), which is activated when the computer system detects that the gaze of the user of the computer system is directed to a second position of the corresponding user interface (e.g., a corner position, or a position corresponding to the first user interface object) (e.g., as shown in reference to FIG). Figure 7CMaintaining a particular function available for activation while disabling other gaze-based functions provides an gaze-based mechanism for activating certain functions while still reducing false positives for other functions. When the function is for reactivating a disabled gaze-based function, this allows the user to avoid having to resort to non-gaze-based control schemes to reactivate the gaze-based function.

[0165] In some embodiments, based on determining that the set of one or more functions activated by detecting a gaze of the user of the computer system is available for activation, the computer system displays an indication (e.g., as an example) that the set of one or more functions activated by detecting a gaze of the user of the computer system is available for activation. Figure 7A 704b) (e.g., a textual indication and / or a graphical indication); and based on determining that the set of one or more functions activated by detecting a gaze of the user of the computer system is disabled (e.g., not available for activation), the computer system displays an indication that the set of one or more functions activated by detecting a gaze of the user of the computer system is disabled (e.g., as shown in 704b); Figure 7C Displaying an indication of whether the set of one or more functions is available for gaze-based activation provides improved visual feedback and may also reduce user frustration that may result from attempting to use a disabled function.

[0166] In some embodiments, the corresponding user interface also includes a current time indicator (e.g., 702a) (e.g., an indication of the current time at the location of the computer system and / or an indicator positioned along a bottom edge of the user interface). In some embodiments, the first user interface object includes an indication of the current time and / or the first user interface object is positioned along a bottom edge of the user interface. Displaying the current time indicator provides the user with improved visual feedback about the current time.

[0167] In some embodiments, the corresponding user interface includes multiple application user interface objects (e.g., 708a to 708c) displayed in a first spatial arrangement (e.g., a row arrangement or a three-dimensional carousel arrangement); and the first control object (e.g., 704d) when selected (e.g., by gazing, by performing an air gesture while the object is in focus, and / or by pressing a hardware button while the object is in focus) causes the multiple application user interface objects to transition from being displayed in the first spatial arrangement to being displayed in a second spatial arrangement (e.g., a grid arrangement or a column arrangement) that is different from the first spatial arrangement.

[0168] In some embodiments, the respective user interface also includes a first representation (e.g., 714) of an application user interface of a first application (e.g., a teleconferencing application, a video application) (e.g., a picture-in-picture ("PiP") representation of a dynamic application) (in some embodiments, the representation is located in a corner of the respective user interface and occupies less than 50%, 40%, 30%, 25%, 20%, or 10% of the area of the respective user interface). Displaying the first representation of the application user interface of the first application provides improved visual feedback about the status and / or properties of the application user interface.

[0169] In some embodiments, the first representation (e.g., 714) is displayed at a first location (e.g., as shown in FIG. Figure 7I ) (e.g., positioned), the computer system detects a first input (e.g., 716a) corresponding to the first representation (e.g., an air gesture (e.g., performed when the first representation is in focus), a gaze-based input, or a hardware button press). In response to the first input, the computer system moves the first representation to a second position in the corresponding user interface that is different from the first position (e.g., as shown in Figure 7J ). In some embodiments, the first representation is displayed in a first corner of the corresponding user interface and is moved to a second corner of the corresponding user interface. Moving the first representation to the second position in the corresponding user interface, different from the first position, in response to input allows the user to free up space for making other content (e.g., pass-through content and / or other virtual objects) visible (e.g., at the first position) and provides the user with greater control over the display position and content. Providing greater control over the display position and the content displayed therein enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0170] In some embodiments, the first position is predefined and the second position is predefined (e.g., the first representation is moved from the first predefined position to the second predefined position). Moving the first representation to the predefined position reduces the number of inputs required to perform the move operation because no input is required to identify the target position.

[0171] In some embodiments, the computer system detects a second input corresponding to the first representation (e.g., an air gesture (e.g., performed while the first representation is in focus), a gaze-based input, or a hardware button press) while the first representation is in the first-representation unexpanded state. In response to detecting the first input, the computer system expands the first representation to the first-representation expanded state (e.g., as Figure 7K 714 shown in ), wherein the first representation is larger (e.g., 25%, 30%, 40%, 50%, 60%, 100%, or 200% larger) in the first representation expanded state than in the first representation unexpanded state. In some embodiments, when in the expanded state, the first representation includes additional content and / or control objects that are not included in the unexpanded state. Expanding the first representation to the first representation expanded state improves the user's ability to interact with the representation (e.g., via gaze input), thereby improving ease of use. Improving ease of use enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0172] In some embodiments, aspects / operations of methods 800, 1000, 1200, and 1400 may be interchanged, replaced, and / or added between these methods. For example, the techniques for using gaze to interact with a graphical user interface of method 800 may be used to interact with a virtual object repositioned via method 1000. As another example, a virtual object repositioned via method 1400 may initially be displayed via the interaction techniques of method 800. For the sake of brevity, these details will not be repeated here.

[0173] 9A to 9D An example of a technique for repositioning a virtual object is shown. Figure 10 is a flow chart of an exemplary method 800 for repositioning a virtual object. 9A to 9D The user interface in the diagram is used to illustrate the process described below, including Figure 10 in the process.

[0174] exist Figure 9AAt , the device 700 is displaying a user interface 702 via a display 700a that includes various virtual objects displayed concurrently with a representation of a physical environment 701. The user interface 702 includes central virtual objects 708a, 708b, and 708c that are displayed at various depths (e.g., displayed in a manner such that a user of the device 700 perceives the objects as being displayed at various depths). In some embodiments, the central virtual objects are arranged in sequence (e.g., the central virtual objects 708a through 708c are part of an ordered sequence) as a rotating carousel of objects. In some embodiments, objects at different depths are displayed with different sizes, different levels of sharpness (e.g., with different levels of bokeh), different degrees of overlap, and / or different levels of blur to convey the depth of the object (e.g., relative to other objects). In some embodiments, the device 700 is an HMD capable of presenting different visual information to each eye of a user, and the virtual objects are stereoscopic virtual objects whose depth is conveyed by the difference in how the object is presented to each eye of the user. Figure 9A , the center virtual object 708b is displayed to cover the center virtual objects 708a and 708c (e.g., the center virtual object 708b is displayed at a depth closer to the user's perspective than the depth at which the center virtual objects 708a and 708c are displayed). Figure 9A At , the center virtual object 708b is perceived by the user as being displayed at a distance of 1 foot from the user, while the center virtual objects 708a and 708c are displayed at a distance of 2 feet from the user. In addition to overlapping the center virtual objects 708a and 708c, the center virtual object 708b is also displayed at a larger size than the center virtual objects 708a and 708c due to the difference in display depth. Figure 9A , each of these center virtual objects has different content, as depicted by the fill pattern of each object (e.g., center virtual object 708a has vertical stripes, center virtual object 708a is stippled, and center virtual object 708a has horizontal stripes). In some embodiments, the difference in content is also a difference in color (e.g., background color or foreground color). For example, center virtual objects 708a to 708c are yellow, blue, and red, respectively. In some embodiments, each center virtual object is associated with a different application or a different XR experience, and the different content is based on the corresponding application or XR experience. Figure 9A , the center virtual object 708b (and in some embodiments, the center virtual objects 708a and 708c) are partially translucent (e.g., as indicated by the increased density of the stippled pattern of the center virtual object 708b at the portion where it overlaps with the other center virtual objects); however, the center virtual object 708b is partially translucent at Figure 9A, so the details of the underlying central virtual object are not visible in the overlapping portion (e.g., dashed lines indicate the boundaries of central virtual objects 708a and 708c in the overlapping portion, but the exact outlines of these boundaries will not be visible). In some embodiments, where each central virtual object has a corresponding color, the area where central virtual object 708b overlaps with central virtual object 708a can be displayed with a greater color saturation (e.g., the overlapping area is a darker blue than the non-overlapping portion of central virtual object 708b), but the hue of the overlapping area is not a combination of blue and red (e.g., not purple) in the overlapping area, and as shown in FIG. Figure 9A These objects are shown in the . Figure 9A At , the device 700 detects input 902a as an actuation of button 700b, which is processed by the device 700 as a request to reposition the center virtual objects 708a to 708c to bring the center virtual object 708c to the front. In some embodiments, the input 902a is an air gesture, a touch on a touch-sensitive surface, or a speech input.

[0175] exist Figure 9B At , in response to detecting input 902a, the device initiates a process for repositioning the center virtual objects 708a to 708c to bring the center virtual object 708c to the front. Figures 9B to 9E In Figure 9B At this point, the process involves shifting each of the central virtual objects to the left in the plane of the display, as can be seen by comparing them with Figure 9A and Figure 9B The relative position of the tree 901a can be seen. Figure 9B The process also includes changing the depth of the central virtual objects 708a to 708c: moving the central virtual objects 708a and 708b further away from the user's perspective (and thus Figure 9A sized than shown in a smaller size), while moving the center virtual object 708c closer to the user's viewing angle (and thus relative to its Figure 9A is shown in a larger size). Figure 9BAt , from the user's perspective, the center virtual object 708a is now displayed at 2.25 feet, the center virtual object 708b is displayed at 1.25 feet, and the center virtual object 708c is displayed at 1.75 feet. The device 700 also displays a center virtual object 708d, which is displayed at a depth farthest from the user's perspective (of the multiple objects) compared to the other center virtual objects (e.g., at 2.75 feet). In some embodiments, these center virtual objects are arranged in sequence (e.g., center virtual objects 708a to 708d are part of an ordered sequence) as a carousel of objects, and as the carousel rotates, they are displayed or stop being displayed when the objects cross a threshold distance (e.g., 3 feet) from the user's perspective. Figure 9B , the area where the center virtual object 708b overlaps the center virtual object 708a remains substantially unchanged (e.g., the center virtual object 708b remains at the same level of translucency in the overlapping area) because the relative distance between the objects remains at 1 foot, as shown in FIG. Figure 9A In contrast, the area where the center virtual object 708b overlaps with the center virtual object 708c has begun to change because the relative distance between the two objects has now decreased to 0.5 feet. Figure 9B , the area where the center virtual object 708b overlaps the center virtual object 708c is now semi-transparent because some details of the center virtual object 708c (e.g., underlying shape content and / or color) are now visible through the center virtual object 708b, as indicated by the portion of the horizontal line pattern of the center virtual object 708c being visible at the overlapping area. Figure 9B , the appearance of the overlapping area is still mainly the appearance of the central virtual object 708b (for example, the background color is indigo with blue as the main color).

[0176] exist Figure 9C , the process for repositioning center virtual objects 708a through 708c to bring center virtual object 708c to the front has progressed further, with all center virtual objects now shifted further to the left. Center virtual objects 708b and 708c are now both displayed at a depth of 1.5 feet from the user's perspective, and both objects are at the same distance from the center of display 700a. From the user's perspective, center virtual object 708a is now 2.5 feet further back, which is the same depth at which center virtual object 708d is now displayed. Figure 9B Likewise, the area where the center virtual object 708b overlaps with the center virtual object 708a remains substantially unchanged (e.g., the center virtual object 708b remains at the same level of translucency in this overlapping area with the center virtual object 708a) because the relative distance between the objects remains at 1 foot, as shown in FIG. Figure 9A and9B In contrast, the appearance of the overlapping area between the center virtual objects 708b and 708c is now a combination (in some embodiments, an equal combination) of the visual characteristics (e.g., content and / or color) of the two objects. For example, the color of the overlapping portion will be purple.

[0177] exist Figure 9D At , the process for repositioning the center virtual objects 708a to 708c to bring the center virtual object 708c to the front has progressed further. The center virtual object 708b is now displayed at a distance of 1.75 feet (e.g., at Figure 9B ft. away from the center virtual object 708c), while the center virtual object 708c is now displayed at a distance of 1.25 feet (e.g., at Figure 9B 8b). Center virtual object 708c now covers center virtual object 708b, and the overlapping area is now primarily based on the appearance of center virtual object 708c, which is semi-transparent in the overlapping area (e.g., the color will be a dark red with a predominantly red color). Center virtual object 708a is now at 2.75 feet, and center virtual object 708d is now at 2.25 feet.

[0178] exist Figure 9E , the process for repositioning center virtual objects 708a through 708c to bring center virtual object 708c to the front is now complete. Center virtual object 708c is now in the center of the display at a distance of 1 foot from the user's viewing angle, and center virtual objects 708b and 708d are both at a distance of 2 feet. Center virtual object 708a is no longer displayed because it transitioned beyond the threshold distance of 3 feet. In areas where center virtual object 708c overlaps with other objects, center virtual object 708c is partially translucent (e.g., as indicated by an increased density of the horizontal line pattern of center virtual object 708c at the portion where it overlaps with the other center virtual objects); however, center virtual object 708c is not displayed in the center of the display. Figure 9E 708c is not transparent at the overlapped portion, so the details of the underlying central virtual object are not visible at the overlapped portion (e.g., dashed lines indicate the boundaries of central virtual objects 708b and 708d at the overlapped portion, but the exact outlines of these boundaries will not be visible). In some embodiments, where each central virtual object has a corresponding color, the area where central virtual object 708c overlaps with central virtual object 708b can be displayed with a greater color saturation (e.g., the overlapped area is a darker red than the non-overlapped portion of central virtual object 708c), but the hue of the overlapped area is not a combination of red and blue (e.g., not purple) at the overlapped area, and as shown in FIG. Figure 9EIn some embodiments, if the device 700 detects another actuation of button 700b, the device 700 will initiate a process of repositioning the center virtual object to bring the center virtual object 708d to the front. Conversely, if the device 700 detects an actuation of button 700c, the device 700 will initiate a process of repositioning the center virtual object to bring the center virtual object 708b to the front again (e.g., Figures 9B to 9E ). Therefore, Figures 9A to 9E A process for repositioning virtual objects is shown, which includes progressively blending / combining the visual characteristics of objects as they shift in relative position and depth. In some embodiments, device 700 is an HMD that presents an XR experience to a user. In such embodiments, virtual objects can be presented to the user at different depths (e.g., via stereoscopic display technology), and Figures 9A to 9E The techniques shown in

[0066] can be used to provide feedback about the relative depth of a virtual object as it is being repositioned. In some embodiments, when the device 700 is an HMD, the user can conveniently provide a user interface model for performing a function on a currently selected and / or currently focused virtual object by positioning the virtual object at a certain location in the user interface. Figures 9A to 9E The techniques shown in can be used to reposition an object and provide visual feedback about the repositioning so that a target virtual object can be focused.

[0179] about Figures 9A to 9E See below for additional description of Figures 9A to 9E Method 1000 is described.

[0180] Figure 10 is a flow chart of an exemplary method 1000 for repositioning a virtual object according to some embodiments. In some embodiments, the method 1000 is performed on a computer system (e.g., Figure 1 The present invention is a method of executing a computer system 101 in a computer system; a head-mounted display; an optical head-mounted display; a personal computer; a smart phone; and / or a tablet computer) that is connected to a display generation component (e.g., a display controller; a touch-sensitive display system; a pass-through display (e.g., integrated and / or connected), a 3D display, a transparent display, a projector, a head-up display, and / or a head-mounted display) (e.g., Figure 1 、 Figure 3 and Figure 4 In some embodiments, the method 1000 communicates with the display generation component 120 in the computer system (e.g., a head-up display, a display, a touch screen, a projector, etc.). In some embodiments, the method 1000 is performed by storing in a non-transitory (or transient) computer-readable storage medium and by one or more processors of a computer system (such as one or more processors 202 of the computer system 101) (e.g., Figure 1 Some operations in method 1000 may be optionally combined, and / or the order of some operations may be optionally changed.

[0181] The computer system (e.g., 700) displays (1002) a corresponding user interface (e.g., 702) via the display generation component. In some embodiments, the corresponding user interface is a set of one or more virtual objects displayed in an extended reality environment. In some embodiments, at least one virtual object in the set of one or more virtual objects is a viewpoint-locked virtual object. Displaying the corresponding user interface includes displaying: a first user interface object (1004) (e.g., 708b), wherein at least a first portion of the first user interface object (e.g., a portion covering 708c) is at least partially translucent (e.g., allowing light (e.g., including the color of the light) from an object behind the object, but not precise details, to at least partially pass through and / or be displayed through the object) and includes first content (e.g., graphical and / or textual content); and in some embodiments, the entire first user interface object is transparent. In some embodiments, in addition to being located in the first portion, the first content also extends into additional portions of the first user interface object. Displaying the corresponding user interface also includes displaying: a second user interface object (1006) (e.g., 708c), wherein: at least a first portion of the second user interface object (e.g., the portion covered by 708b) includes second content that is different from the first content; in some embodiments, the entire second user interface object is transparent (in some embodiments, in addition to being located in the first portion, the second content also extends into additional portions of the second user interface object; in some embodiments, the second content is not visible or is completely visible while the first user interface object remains in front of the second user interface object), and the first user interface object is displayed in front of the second user interface object (e.g., closer to the z-axis along a plane perpendicular to the display generating component from the perspective of a user of the computer system) so that the first portion of the first user interface object covers the first portion of the second user interface object (e.g., displayed in front of the first portion of the second user interface object) (e.g., as shown in FIG. Figure 9A shown).

[0182] While the first user interface object is displayed in front of the second user interface object, the computer system receives (1008) a request (e.g., 902a) (e.g., via one or more input devices in communication with the computer system) (in some embodiments, the request is: a gesture on a touch-sensitive surface; an air gesture performed with a hand of a user of the computer system; an actuation of a hardware button or key; and / or a voice command) to move the second user interface object in front of the first user interface object.

[0183] In response to receiving (1010) the request to move the second user interface object in front of the first user interface object, the computer system: initiates (1012) a process to move the second user interface object in front of the first user interface object (e.g., Figures 9B to 9E The process includes modifying (e.g., blending) (e.g., before the second user interface object is moved in front of the first user interface object) the visual appearance of the first portion of the first user interface object to include third content based on a first combination of the first content and the second content (e.g., Figures 9B to 9C ). In some embodiments, the combination is based on combining and / or synthesizing the first content and the second content to form the third content. In some embodiments, the third content includes a portion of the first content (e.g., not including all of the first content) and a portion of the second content. In some embodiments, the process includes: moving the second user interface object in front of the first user interface object so that the first portion of the second user interface object covers the first portion of the first user interface object. In some embodiments, when the first user interface object and the second user interface object are displayed at the same depth (e.g., along the z-axis) so that neither object is in front of the other, the first portion of the first user interface object is displayed together with the third content. Modifying the visual appearance of the first portion of the first user interface object to include third content based on the first combination of the first content and the second content provides improved visual feedback about: displaying the relative depths of the two objects, the overlapping area between the two objects, the content of the two objects at the overlapping area, and the current state of the process of moving the second user interface object in front of the first user interface object.

[0184] In some embodiments, the first portion of the second user interface object is at least partially translucent (e.g., as shown in FIG. Figure 9AMaking the first portion of the second user interface object at least partially translucent provides visual feedback about the visual characteristics of any objects displayed behind the portion.

[0185] In some embodiments, the first content includes a first background color (e.g., blue of 708b) (e.g., blue, green, or yellow). In some embodiments, the first content includes foreground content that is different from the background color. The second content includes a second background color (e.g., red of 708c) that is different from the first background color (e.g., red, orange, or brown); and the third content includes an intermediate background color based on a combination of the first background color and the second background color (e.g., when the first background color is red and the second background color is blue, the intermediate background color is purple) (e.g., as shown in reference Figures 9B to 9D Providing the third content with an intermediate background color that is based on a combination of the first background color and the second background color provides improved visual feedback regarding the relative depths of the two objects, the area of overlap between the two objects, and the current state of the process of moving the second user interface object in front of the first user interface object.

[0186] In some embodiments, the second portion of the first user interface object (e.g., the portion of 708b that overlaps with 708c) (e.g., the same or different portion as the first portion of the first user interface object) includes fourth content; the second portion of the second user interface object (e.g., the portion of 708c that overlaps with 708b) (e.g., the same or different portion as the first portion of the second user interface object) includes fifth content; and the process of moving the second user interface object in front of the first user interface object includes: modifying (e.g., while or after moving the second user interface object in front of the first user interface object) the visual appearance of the second portion of the second user interface object to include a combination based on the fourth content and the fifth content (in some embodiments, the combination includes combining colors, patterns and / or details (e.g., details of shapes) of the fourth content and the fifth content) (e.g., as Figure 9DIn some embodiments, a majority of the combination of the fourth content and the fifth content is based on the fifth content. Modifying the visual appearance of the second portion of the second user interface object to include sixth content that is based on the combination of the fourth content and the fifth content provides improved visual feedback regarding: displaying the relative depths of the two objects, the overlapping area between the two objects, the content of the two objects at the overlapping area, and the current state of the process of moving the second user interface object in front of the first user interface object.

[0187] In some embodiments, when the second portion of the second user interface object has the visual appearance including the sixth content, the second portion of the second user interface covers (e.g., overlaps; is displayed in front of) the second portion of the first user interface object (e.g., Figure 9C and Figure 9D In some embodiments, other portions of the second user interface object that do not cover the first user interface object do not include the sixth content.

[0188] In some embodiments, when the first user interface object is displayed in front of the second user interface object: a third portion of the first user interface object including the seventh content covers a first portion of a third user interface object (e.g., 708a) including the eighth content (e.g., different from the first user interface object and the second user interface object), and the visual appearance of the third portion of the first user interface object is not based on the eighth content (e.g., is not affected by or is not the result of a combination of the seventh content and the eighth content) (e.g., as in Figure 9A In some embodiments, the visual appearance of the third portion of the first user interface object is based solely on the seventh content.

[0189] In some embodiments, the first combination of the first content and the second content includes a first percentage (e.g., a first extent or a first amount) of the first content in the first combination. In some embodiments, after modifying the visual appearance of the first portion of the first user interface object to include the third content based on the first combination of the first content and the second content (e.g., after Figure 9B), the computer system modifies (e.g., further modifies and / or gradually modifies) the visual appearance of the first portion of the first user interface object to include ninth content based on a second combination of the first content and the second content, wherein the second combination of the first content and the second content includes a second percentage (e.g., a higher percentage and / or a lower percentage) of the first content in the second combination that is different from the first percentage (e.g., as Figure 9C ). In some embodiments, modifying the visual appearance of the first user interface object includes gradually offsetting (e.g., decreasing or increasing) the extent of the first content in the combination. Modifying the visual appearance of the first portion of the first user interface object to include the ninth content based on the second combination of the first content and the second content (wherein the second combination of the first content and the second content includes a second percentage of the first content in the second combination that is different from the first percentage) provides improved visual feedback regarding: displaying the relative depths of the two objects, the overlapping area between the two objects, the content of the two objects at the overlapping area, and the current state of the process of moving the second user interface object in front of the first user interface object.

[0190] In some embodiments, the process of moving the second user interface object in front of the first user interface object includes: moving the first user interface object (e.g., gradually moving at a predetermined rate) in a first non-depth direction (e.g., as shown in FIG. 1 ) while changing the depth at which the first user interface object is displayed relative to the depth at which the second user interface object is displayed. Figure 9B , shifting the second user interface object in a second non-depth direction (e.g., the same as or different from the first direction) while changing the depth at which the first user interface object is displayed relative to the depth at which the second user interface object is displayed. Figure 9B shown offset to the left).

[0191] In some embodiments, the process of moving the second user interface object in front of the first user interface object includes: modifying (e.g., gradually modifying at a predetermined rate) the size of the first user interface object (e.g., reducing or increasing the size) while changing the depth at which the first user interface object is displayed relative to the depth at which the second user interface object is displayed (e.g., Figure 9A and Figure 9B and modifying the size of the second user interface object (e.g., as shown in the size change of 708b between ); and upon changing the depth at which the first user interface object is displayed relative to the depth at which the second user interface object is displayed, modifying the size of the second user interface object (e.g., as shown in Figure 9A and Figure 9B, and 708c between the first and second user interface objects. In some embodiments, the modification of the size of the first user interface object is opposite (e.g., inversely proportional) to the modification of the size of the second user interface object (e.g., decreasing the size of the first user interface object while increasing the size of the second user interface object). Modifying the size of the second user interface object provides improved visual feedback regarding the current state of the process of moving the second user interface object to the front of the first user interface object.

[0192] In some embodiments, the first user interface object corresponds to a first extended reality experience (e.g., as referenced in FIG. Figure 9A discussed) (e.g., an extended reality user interface generated by a first application (e.g., an extended reality media viewer application, an extended reality media capture application, or an extended reality conferencing application)) (in some embodiments, corresponding to Figures 11A to 11I 、 Figures 13A to 13K , method 1200 and / or method 1400); and the second user interface object corresponds to a second extended reality experience different from the first extended reality experience (e.g., as referenced Figure 9A discussed).

[0193] In some embodiments, displaying the corresponding user interface includes displaying: a representation of the physical environment (e.g., 701) (e.g., an optical or video pass-through representation); the first user interface object (e.g., 708b) (in some embodiments, the first user interface object is a viewpoint-locked virtual object) overlaying the representation of the physical environment (e.g., being displayed in front of and / or above it); and the second user interface object (e.g., 708c) (in some embodiments, the second user interface object is a viewpoint-locked virtual object) overlaying the representation of the physical environment.

[0194] In some embodiments, before initiating the process of moving the second user interface object in front of the first user interface object, the first user interface object is displayed at a depth closer to the user than the depth at which the second user interface object is displayed, from the perspective of the user of the computer system (e.g., as shown in FIG. Figure 9A discussed above) (e.g., the first user interface object appears closer to the user than the second user interface object); and after completing the process of moving the second user interface object in front of the first user interface object, the second user interface object is displayed at a depth closer to the user than the depth at which the first user interface object is displayed, from the perspective of the user of the computer system (e.g., as discussed above); Figure 9EChanging the relative depths at which the first and second user interface objects are displayed before and after the moving process is completed provides improved visual feedback about the process of moving the second user interface object in front of the first user interface object.

[0195] In some embodiments, before initiating the process of moving the second user interface object in front of the first user interface object, the second user interface object is displayed with a first blur amount (e.g., a blur amount of 0 or non-zero); and in some embodiments, before initiating the process of moving the second user interface object in front of the first user interface object, the first user interface object is displayed without blur. After completing the process of moving the second user interface object in front of the first user interface object, the first user interface object is displayed with a second blur amount that is greater than the first blur amount (e.g., as shown in FIG. 2 ). Figure 9A and Figure 9E In some embodiments, after the process of moving the second user interface object in front of the first user interface object is completed, the second user interface object is displayed without any blur, with the first blur amount, and / or with a blur amount that is less than the second blur amount. Displaying the first user interface object with a greater amount of blur provides improved visual feedback about the relative positions of the first and second user interface objects in z-space as the first user interface object moves backward in z-order.

[0196] In some embodiments, displaying the respective user interface includes: displaying a plurality of user interface objects in a sequentially ordered arrangement (e.g., the order of 708a through 708c); the first user interface object and the second user interface object being part of the plurality of user interface objects; and, prior to initiating the process of moving the second user interface object in front of the first user interface object, displaying the first user interface object (e.g., as shown in FIG. 1 ) at a depth closer to the user from the perspective of a user of the computer system than user interface objects that precede or follow the first user interface object in the sequentially ordered arrangement (e.g., including the second user interface object). Figure 9A ). In some embodiments, after the process of moving the second user interface object in front of the first user interface object is completed, the second user interface object is displayed at a depth closer to the user from the perspective of the user of the computer system than user interface objects that are located before or after the first user interface object in the sequentially ordered arrangement. Displaying the first user interface object closer to the foreground than elements that are located before and after the first user interface object in the sequentially ordered arrangement provides improved visual feedback of the position of the first user interface object in the sequentially ordered arrangement.

[0197] In some embodiments, receiving the request to move the second user interface object in front of the first user interface object includes detecting activation (e.g., actuation (e.g., pressing, sliding, or rotating)) of a hardware input mechanism (e.g., 700b) that communicates with the computer system (e.g., a button (e.g., an actuated mechanical button or solid-state button that detects input pressure (which, in some embodiments, provides tactile feedback to indicate detected pressure / input)), a dial, a slider, or a knob). Detecting the request including activation of a hardware mechanism provides tactile feedback to the user (e.g., via actuation of the mechanism and / or tactile feedback) that the input has been properly received. Providing improved feedback that input has been received enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0198] In some embodiments, receiving the request to move the second user interface object in front of the first user interface object includes detecting an air gesture (e.g., as referenced in FIG. Figure 7A and input 902a) (e.g., detected via one or more input mechanisms (e.g., cameras, hand motion sensors) in communication with the computer system). Detecting requests including detecting air gestures provides the user with an input modality that does not require manual contact with the device's sensors and / or input mechanisms, which reduces the risk that the user will fail to provide input or fail to provide input appropriately (e.g., via errors in locating the sensors / input mechanisms, such as when operating an HMD with input mechanisms that are not visible to the user). Reducing the risk of input failures enhances device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors in operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0199] In some embodiments, receiving the request to move the second user interface object in front of the first user interface object includes: detecting, via one or more gaze tracking sensors in communication with the computer system, that the gaze of the user of the computer system is directed toward a corresponding portion of the corresponding user interface (e.g., a portion corresponding to the second user interface object) (e.g., directed toward a direction corresponding to the user's gaze that intersects with the corresponding portion) (in some embodiments, directed toward the corresponding portion for at least a predetermined period of time (e.g., 0.25 seconds, 0.5 seconds, or 1 second)) (e.g., as described in reference Figure 7ADetecting a request including detecting the user's gaze provides the user with an input modality that does not require manual contact with the device's sensors and / or input mechanisms, which reduces the risk that the user will fail to provide input or fail to provide input appropriately (e.g., via errors in positioning the sensor / input mechanism, e.g., when operating an HMD with an input mechanism that is not visible to the user). Reducing the risk of input failure enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors in operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0200] Figures 11A to 11I An example of a technique for transitioning the mode of a camera capture user interface is shown. Figure 12 is a flow chart of an exemplary method 1200 for transitioning a mode of a camera capture user interface. Figures 11A to 11H The user interface in the diagram is used to illustrate the process described below, including Figure 12 in the process.

[0201] exist Figure 11A 1 , a user 1101 is using the device 700 while the user is in a physical environment 1103. A subject 1105 and a light 1107 are also located in the physical environment 1103. Figure 11A 1100 depicts the relative positions of a user 1101, a subject 1105, and a light 1107 within an environment 1103. The device 700 includes a display 700a, buttons 700b to 700d, sensors 306 (e.g., accelerometers and / or gyroscopes for detecting the position and / or movement of the device 700), and multiple cameras (e.g., as part of sensors 314), including Figure 11A Two or more rear cameras pointing in the direction of the main body 1105 and the light 1107.

[0202] exist Figure 11A At , the device 700 displays a user interface 1102 overlaid on a representation 1104a of the physical environment 1103. In some embodiments, in response to Figure 7A1103 is displayed upon selection of the central virtual object of the device 700. Representation 1104a is based on image data captured by one or more of the multiple cameras of the device 700. In some embodiments, representation 1104a is an optical pass-through representation of the physical environment (e.g., one or more portions of the device 700 are transparent, and the physical environment 1103 is visible through these portions). User interface 1102 is an interface for a camera application that can be used to capture media (e.g., photos, stereoscopic photos, videos, and / or stereoscopic videos). The user interface 1102 includes several virtual objects (e.g., objects for assisting a user in capturing media, modifying one or more settings of the application and / or camera, and / or reviewing captured media), including a guideline 1106 (e.g., a framing virtual object for assisting a user in framing content for capture), a control virtual object 1108 (e.g., for modifying a camera preview mode), a control virtual object 1110 (e.g., for modifying a camera capture mode (e.g., photo or video)), a shutter button 1112 (e.g., for initiating a media capture process), and a face detection indicator 1114 (e.g., indicating one or more faces detected within the representation 1104a). Figure 11A 1104a, the graticule 1106 indicates an area of the representation 1104a that will be included in the captured media, but the appearance of the area within the graticule 1106 is not otherwise modified compared to the rest of the representation 1104a. In some embodiments, the user interface 1102 is an XR interface and presents an XR experience to the user 1101. In such embodiments, one or more virtual objects of the user interface 1102 can be environment-locked virtual objects (e.g., face detection indication 1114) or viewpoint-locked virtual objects (e.g., shutter button 1112); in addition, one or more virtual objects of the user interface 1102 can be displayed at different depths relative to the user's viewing angle and elements of the physical environment 1103, as depicted in representation 1104a.

[0203] exist Figure 11A At , user 1101 is interested in capturing media that includes subject 1105, but user 1101 wishes to have subject 1105 more toward the center of the resulting captured media. Figure 11A At , device 700 detects that device 700 is moving to the left, as indicated by indication 1109a.

[0204] exist Figure 11B, in response to detecting that the device 700 has moved to the left, the representation 1104a has been updated with the subject 1105 positioned more toward the center of the guideline 1106. The diagram 1100 reflects the updated relative positions of the user 1101 and the device 700 relative to the subject 1105 and the light 1107. The face detection indicator 1114 has also shifted as it continuously tracks the position of the detected face. Figure 11B At , the device 700 detects input 1111a (e.g., actuation of button 700b) while the device 700 detects that the gaze of the user of the device 700 is directed toward the shutter button 1112, as indicated by gaze indication 1113a. Alternatively, the device 700 detects input 1111a while the device 700 detects that the gaze of the user of the device 700 is directed toward controlling the virtual object 1108, as indicated by gaze indication 1113b. In some embodiments, the input 1111a is an air gesture, a touch on a touch-sensitive surface, or a speech input. In some embodiments, detecting that the user's gaze is directed toward the virtual object includes determining that the user's gaze has remained on the virtual object for a predetermined period of time (e.g., 0.1 seconds, 0.25 seconds, 0.5 seconds, or 1 second).

[0205] exist Figure 11C At 1101, in response to detecting input 1111a when device 700 detects that the gaze of the user of device 700 is directed toward shutter button 1112, device 700 initiates a process for capturing media using multiple cameras of device 700. The process includes modifying the appearance of reticle 1106 (e.g., by thickening reticle 1106) to provide feedback to user 1101 about the media being captured.

[0206] Located in time Figure 11C After Figure 11D At this point, the process for capturing media initiated by detecting input 1111a is now complete. The display of the reticle 1106 is similar to Figure 11B , and user interface 1102 now includes a preview virtual object 1116a, which is a representation of the media item (e.g., a photo) captured during the process initiated by input 1111a. The media item corresponding to preview virtual object 1116 has a field of view of the physical environment 1103 that is different from the field of view included within graticule 1106 and different from the field of view of the physical environment 1103 as a whole presented in representation 1104a. In some embodiments, the media item corresponding to preview virtual object 1116 has a field of view that encompasses more of the physical environment 1103 than the physical environment 1103 demarcated by graticule 1106 (e.g., graticule 1106 indicates the general area that will be captured); for example, in Figure 11D, preview virtual object 1116 includes all of light 1107, while only a portion of light 1107 is within guideline 1116. In some embodiments, the media item corresponding to preview virtual object 1116 has a field of view that encompasses less of the physical environment 1103 than the physical environment 1103 delineated by guideline 1106. In either embodiment, preview virtual object 1116 provides an indication to the user of device 700 of exactly what was captured, which the user can review to determine whether the captured content is acceptable and, if not, delete the media.

[0207] exist Figure 11E In response to detecting input 1111a when device 700 detects that the gaze of the user of device 700 is directed toward controlling virtual object 1108, as indicated by Figure 11B , the device 700 modifies the appearance of the control virtual object 1108 to indicate that the corresponding camera preview display mode is now active. While in the corresponding camera preview display mode, the device 700 displays representation 1104b, which is a picture-in-picture ("PiP") representation of the physical environment 1103 overlaid on a portion of representation 1104a. In some embodiments, selecting the control virtual object 1108 while the device 700 is in the corresponding camera preview display mode will revert to 11A to 11D In some embodiments, after a predetermined period of time has elapsed, the device 700 automatically transitions back to the camera preview display mode shown in FIG. 11A to 11D . In some embodiments, representation 1104a presents image data from a first set of cameras in the plurality of cameras of device 700, and representation 1104b presents image data from a second, different set of cameras in the plurality of cameras of device 700. In some embodiments, the second set of cameras is a set of cameras used when capturing media (e.g., in response to input corresponding to shutter button 1112). In some embodiments, representation 1104b includes a field of view of physical environment 1103 that is different from the field of view included within reticle 1106 and different from the field of view of physical environment 1103 as a whole presented in representation 1104a. In some embodiments, in which representation 1104a is an optical pass-through representation of physical environment 1103, representation 1104b is a virtual object (e.g., a viewpoint-locked virtual object) overlaid on a portion of optical pass-through representation 1104a. In some embodiments, representation 1104b is partially translucent or semi-transparent, and the color and / or details of the underlying optical pass-through representation can be perceived through representation 1104b. Representation 1104b includes extended control 1104a1. Figure 11E At 1109 b, device 700 detects that device 700 is moving backward (eg, away from body 1105), as indicated by indication 1109b.

[0208] exist Figure 11F In response to detecting that the device 700 has moved backward, representation 1104a has been updated, wherein the subject 1105 now appears to be larger than Figure 11E 1107. The face detection indicator 1114 is now smaller because the size of the subject 1105's face is now smaller and because the face detection indicator 1114 is a stereoscopic virtual object displayed at the same depth at which the detected face is displayed. Diagram 1100 reflects the updated relative positions of the user 1101 and device 700 with respect to the subject 1105 and light 1107. Figure 11F At , the device 700 detects that the gaze of the user of the device 700 is directed toward the expansion control 1104a1 of the representation 1104b, as indicated by the gaze indication 1113c.

[0209] exist Figure 11G In response to detecting that the gaze of the user of device 700 is directed toward expansion control 1104a1 of representation 1104b, device 700 expands representation 1104b to occupy the area delimited by guide lines 1106. In some embodiments, when representation 1104b is expanded, device 700 stops displaying guide lines 1106. In some embodiments, representation 1104b, when expanded, is smaller than or larger than the area delimited by guide lines 1106 and / or has a geometry that is different from the geometry of the area delimited by guide lines 1106. Figure 11G As shown in FIG, the expanded representation 1104b is partially translucent or semi-transparent, and the colors and / or details of the underlying optical pass-through representation can be perceived through the representation 1104b; for example, before the representation 1104b is expanded, the body 1105 of the representation 1104a is still visible under the expanded representation 1104b, although with a larger size. Figure 11F Less color and / or shape detail in ; the same is true for lamp 1107. Figure 11G , the appearance of the subject 1105 in the enlarged representation 1104b (labeled 1105a) is clearer and more distinct than the appearance of the subject 1105 in the representation 1104a, and the same is true for the light 1107 (labeled 1107a in the enlarged representation 1104b). The enlarged representation 1104b includes a folding control virtual object 1104b2 that, when selected (e.g., via gazing), returns the representation 1104b to its non-enlarged size. In some embodiments, the enlarged representation 1104b provides the user with a representation of the field of view from the second set of cameras that will be used when capturing media (e.g., in response to input corresponding to a shutter button 1112). In some embodiments, after the enlarged representation has been displayed for a period of time, the device 700 automatically redisplays the representation 1104b in the non-enlarged state. Figure 11GAt , device 700 detects input 1111b (e.g., actuation of button 700b) while device 700 detects that the gaze of the user of device 700 is directed toward shutter button 1112, as indicated by gaze indication 1113d. In some embodiments, input 1111a is an air gesture, a touch on a touch-sensitive surface, or speech input.

[0210] exist Figure 11H At , in response to detecting input 1111b when device 700 detects that the gaze of the user of device 700 is directed toward shutter button 1112, device 700 initiates a second process for capturing media using multiple cameras of device 700. The process includes modifying the area delineated by guideline 1106 by displaying virtual object 1118 that whitens the area and temporarily ceasing to display the enlarged representation 1104b. Figure 11H As shown in FIG, virtual object 1118 is partially translucent or semi-transparent, and the color and / or details of the underlying optical pass-through representation can be perceived through representation 1104b (e.g., the faded representation of subject 1105 is not visible in FIG). Figure 11H is visible in the ).

[0211] Located in time Figure 11H After Figure 11I At this point, the process for capturing media initiated by detecting input 1111b is now complete. Representation 1104a, enlarged representation 1104b, and guideline 1106 are displayed in the same manner as Figure 11G Device 700 has ceased displaying virtual object 1118 and now displays preview virtual object 1116b, which is a representation of the media item (e.g., a photo) captured during the second process (e.g., initiated by input 1111b). In some embodiments, when the media is captured using the same second set of cameras used to present the magnified representation 1104b, preview virtual object 1116b includes a representation of the same field of view of the physical environment 1103 that was represented by the magnified representation 1104b when the media was captured. Thus, Figures 11A to 11I The techniques described in provide multiple camera preview modes to the user to assist the user in composing and capturing media items. In some embodiments where the device 700 is an HMD, particularly an HMD having optical pass-through (e.g., via one or more transparent portions of the device 700), Figures 11A to 11I The technology described in provides a user with a single representation for utilizing a physical environment (e.g., according to Figures 11B to 11D 1104a) or two representations of the physical environment (e.g., according to Figures 11E to 11I1104b) to compose and capture media, where one of the representations (e.g., 1104b) is from the set of cameras to be used for media capture. In some embodiments, the selected camera preview mode is persistent between sessions. For example, if the application generating the user interface 1102 is closed or suspended while displaying representation 1104b, representation 1104b will be redisplayed on subsequent redisplays of the user interface 1102.

[0212] about Figures 11A to 11I See below for additional description of Figures 11A to 11I Method 1200 is described.

[0213] Figure 12 is a flow chart of an exemplary method 1200 for transitioning a mode of a camera capture user interface according to some embodiments. In some embodiments, the method 1200 is performed on a computer system (e.g., Figure 1 In some embodiments, the method 1200 is executed at a computer system 101 in a device; a head-mounted display; an optical head-mounted display; a personal computer; a smart phone; and / or a tablet computer) that communicates with a display generation component (e.g., a display generation component 120; a display controller; a touch-sensitive display system; a pass-through display (e.g., integrated and / or connected), a 3D display, a transparent display, a projector, a head-up display, and / or a head-mounted display) and one or more cameras. In some embodiments, the one or more cameras are multiple cameras with different perspectives (e.g., with partially overlapping fields of view) capable of capturing stereoscopic media (e.g., photos and / or videos). In some embodiments, the method 1200 is performed by storing in a non-transitory (or transient) computer-readable storage medium and executed by one or more processors of a computer system (such as one or more processors 202 of the computer system 101) (e.g., Figure 1 Some operations in method 800 may be optionally combined, and / or the order of some operations may be optionally changed.

[0214] The computer system (e.g., 700) displays (1202) via the display generation component (e.g., 700a) and in a mixed reality environment (e.g., 1102 and 1104a, combined) a camera capture user interface (e.g., 1102) overlaid on a portion of a physical environment (e.g., 1103) visible to a user (e.g., 1101) of the computer system (e.g., when operating the computer system) (e.g., visible through a transparent portion of the computer system and / or visible as a pass-through representation generated by the computer system), wherein: the camera capture user interface is in a first mode (e.g., as in 11A to 11Das shown) (display mode; camera capture user interface configuration and / or layout mode); when in the first mode, the camera capture user interface includes a set of one or more framed virtual objects (e.g., 1106) (e.g., a hollow geometric shape (e.g., a hollow rectangle); multiple unconnected corners defining a rectangular area), the set of one or more framed virtual objects being viewpoint-locked and indicating a first sub-portion of the physical environment to be captured by the one or more cameras upon receiving a first media capture request (e.g., input 1111a) (e.g., a request to capture still images and / or video when in the first mode); in some embodiments, the captured media includes at least a sub-portion of the physical environment (e.g., capturing the first sub-portion and the second sub-portion). In some embodiments, the set of one or more framed virtual objects does not block a substantial amount of the first sub-portion (e.g., a substantial portion (e.g., 80%, 85%, 90%, or 95%) of the first sub-portion is not covered and / or obscured by the set of one or more framed virtual objects).

[0215] While displaying the camera capture user interface in the first mode, the computer system receives (1204) (e.g., via one or more input devices in communication with the computer system) a request (e.g., 1113b) (in some embodiments, the request is a gesture on a touch-sensitive surface; an air gesture performed with a hand of a user of the computer system; an actuation of a hardware button or key; a gaze-based input; and / or a voice command) (in some embodiments, the request is multiple inputs (e.g., a first input followed by a second input)) to transition the camera capture user interface to a second mode different from the first mode (e.g., such as Figures 11G to 11I ).

[0216] In response to receiving the request to transition the camera capture user interface to the second mode, the computer system displays (1206) the camera capture user interface in the second mode (e.g., Figure 11G). In some embodiments, the camera capture user interface in the second mode does not include the set of one or more framing elements. When in the second mode, the camera capture user interface includes a first representation (e.g., 1104b) of the field of view of at least a first camera of the one or more cameras (e.g., a portion or all of the field of view); in some embodiments, the first representation is a representation of the area where the fields of view of multiple cameras overlap. In some embodiments, the first representation is a live camera feed. The first representation is overlaid on (e.g., displayed on top of) a second sub-portion (e.g., the portion within 1106) of the physical environment (in some embodiments, the first sub-portion and the second sub-portion are the same), which is to be captured by the one or more cameras upon receiving a second media capture request (e.g., 1111b) (e.g., a request to capture still images and / or video when in the second mode). In some embodiments, the first representation is semi-transparent and / or translucent such that aspects of the sub-portion of the physical environment are visible through the first representation. In some embodiments, the captured media includes at least the second subportion of the physical environment (e.g., capturing the second subportion and another subportion). In some embodiments, the first representation is viewpoint-locked. Displaying a corresponding user interface operable in two different modes of providing different camera-captured synthetic visual aids indicating different subportions of the physical environment provides the user with multiple ways of synthesizing media capture events. Doing so enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently. In addition, doing so reduces the risk that transient media capture opportunities (e.g., opportunities to capture transient events / syntheses) will be captured by mistake.

[0217] In some embodiments, while displaying the camera capture user interface in the first mode, the computer system receives the first media capture request (e.g., 1111a). In response to receiving the first media capture request, the computer system captures first media (e.g., still media and / or video media; spatial media / stereoscopic media; two-dimensional media or three-dimensional media) via the one or more cameras, the first media including content corresponding to at least the first sub-portion of the physical environment (e.g., media corresponding to 1116a).

[0218] In some embodiments, the first media includes content corresponding to at least a third subportion of the physical environment, at least the third subportion of the physical environment corresponding to an area of the physical environment outside of the one or more framed virtual objects (e.g., as discussed in reference 1116a) (e.g., the set of one or more framed virtual objects does not indicate (e.g., the set of one or more framed virtual objects depicts an area of the physical environment that does not include the third subportion (e.g., the third subportion is outside of the set of one or more framed virtual objects)) that the third subportion of the physical environment will be captured by the one or more cameras upon receiving the first media request, or the set of one or more framed virtual objects indicates that the third subportion of the physical environment will be captured based on the proximity of the third subportion of the physical environment to the one or more framing elements). Capturing the additional third subportion of the physical environment outside of the set of one or more framing elements reduces the risk of not capturing and / or miscapturing the intended content due to errors in operating the computer system and / or due to unexpected movement of the computer system and / or composition or subject. Doing so enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently. Additionally, doing so reduces the risk that transient media capture opportunities (e.g., opportunities to capture transient events / compositions) will be captured by mistake.

[0219] In some embodiments, while displaying the camera capture user interface in the second mode, the computer system receives the second media capture request (e.g., 1111b). In response to receiving the second media capture request, the computer system captures second media (e.g., still media and / or video media; spatial media / stereoscopic media; two-dimensional media or three-dimensional media) via the one or more cameras, the second media including content corresponding to at least the second sub-portion of the physical environment (e.g., media corresponding to 1116b).

[0220] In some embodiments, the second media does not include content corresponding to any sub-portion of the physical environment that is not represented in the first representation (e.g., as discussed with reference to 1116b) (e.g., the first representation is a true indication / preview of content that will be included in the media captured while in the second mode). Providing the first representation as a true indication / preview of content that will be included in the media captured while in the second mode provides the user with a visual composition assistance that accurately reflects the content that will be captured, which helps the user frame the desired capture event. Doing so enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently. In addition, doing so reduces the risk that transient media capture opportunities (e.g., opportunities to capture transient events / compositions) will be captured in error.

[0221] In some embodiments, while displaying the first representation and based on determining that the first set of one or more criteria are met, the computer system stops displaying the first representation (e.g., as described in reference to Figure 11E as discussed above) (in some embodiments, and transitioning the camera capture user interface to the first mode), wherein the first set of one or more criteria includes criteria that are satisfied when the first representation has been displayed for a first predetermined time period (in some embodiments, the first set of one or more criteria includes criteria that are satisfied when no capture request is received during the first predetermined time period). Ceasing to display the first representation when the first set of one or more criteria is satisfied performs an operation when a set of conditions has been met without requiring further user input. Doing so also reduces power consumption by reducing the number of elements displayed.

[0222] In some embodiments, while displaying the first representation and based on determining that a second set of one or more criteria are met, the computer system reduces the size of the first representation from the first size to a second size that is smaller than the first size (e.g., as described in reference to FIG. Figure 11GIn some embodiments, the first representation is displayed for a second predetermined period of time (as discussed above), wherein the second set of one or more criteria includes criteria that are satisfied when the first representation has been displayed for a second predetermined period of time (in some embodiments, the second set of one or more criteria includes criteria that are satisfied when no capture request is received during the second predetermined period of time). In some embodiments, reducing the size includes displaying the first representation at a reduced size at a predetermined location (e.g., within the viewpoint of a user of the computer system and / or relative to one or more other elements of the camera capture user interface). Reducing the size of the first representation when the second set of one or more criteria is met performs an operation when a set of conditions have been met without requiring further user input. Doing so also reduces power consumption by reducing the displayed size of the element.

[0223] In some embodiments, while displaying the first representation at the second size, the computer system receives a request to expand the size of the first representation (e.g., a gaze-based request, actuation of a hardware input mechanism while the first representation is in focus, touch input, and / or a voice command); and in response to the request to expand the size of the first representation, expands the size of the first representation from the second size to a third size that is larger than the second size (e.g., Figures 11F to 11G ) (in some embodiments, the third size is the first size; in some embodiments, the third size is a predetermined size).

[0224] In some embodiments, when the first representation is displayed at the second size, the camera capture user interface includes a first selectable virtual object (e.g., 11104b1) (e.g., an expanded enable representation displayed over or adjacent to the first representation); and the request to expand the size of the first representation includes input (e.g., 1113c) corresponding to the first selectable virtual object (e.g., a gesture on a touch-sensitive surface; an in-air gesture performed with a hand of a user of the computer system; an actuation of a hardware button or key; gaze-based input; and / or a voice command).

[0225] In some embodiments, when displaying the first representation at the third size, the computer system modifies the visual appearance of at least the second sub-portion of the physical environment (e.g., Figure 11G) (e.g., obscuring, darkening, applying a mask over the portion, and / or applying a tint layer) (in some embodiments, at least some visual details of the second subportion remain visible). In some embodiments, the second subportion corresponds to the portion of the camera capture user interface displaying the first representation; in some embodiments, the second subportion includes or overlaps with the first subportion. Modifying the visual appearance of at least the second subportion of the physical environment while the first representation is magnified is performed when a set of conditions have been met, without requiring further user input. Doing so also improves the visibility of the first representation by reducing potential interference from visual elements of the physical environment (e.g., bright lights).

[0226] In some embodiments, while displaying the set of one or more framed virtual objects, the computer system receives a third media capture request (e.g., a gesture on a touch-sensitive surface; an in-air gesture performed with a hand of a user of the computer system; an actuation of a hardware button or key; a gaze-based input; and / or a voice command). In response to receiving the third media capture request, the computer system: captures third media (e.g., still media and / or video media; spatial media / stereoscopic media; or two-dimensional media or three-dimensional media) via the one or more cameras; and displays an animation that includes modifying the visual appearance of at least a portion of the set of one or more framed virtual objects (e.g., as Figure 11C and Figure 11H ) (e.g., an animation that includes blurring an area including and / or adjacent to the set of one or more framed virtual objects; or an animation that changes size of the set of one or more framed virtual objects). Displaying an animation that includes modifying the visual appearance of at least a portion of the set of one or more framed virtual objects at the time of capture provides improved visual feedback about the capture event.

[0227] In some embodiments, the first representation includes representations of the first subportion of the physical environment and a fourth subportion of the physical environment; and the set of one or more framed virtual objects does not indicate (e.g., the set of one or more framed virtual objects depicts an area of the physical environment that does not include the fourth subportion (e.g., the third subportion is outside the set of one or more framed virtual objects)) the fourth subportion of the physical environment to be captured by the one or more cameras upon receiving the first media request (e.g., the first representation includes one or more portions of the physical environment not indicated by the set of one or more framed virtual objects). In some embodiments, the field of view represented in the first representation is wider than the field of view indicated by the set of one or more framed virtual objects (e.g., of the one or more cameras). Including an additional fourth sub-portion of the physical environment beyond the set of one or more framed elements in the first representation provides the user with different composite assistance encompassing different portions of the physical environment, enhancing the operability of the device and making the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently. Furthermore, doing so reduces the risk that transient media capture opportunities (e.g., opportunities to capture transient events / compositions) will be captured in error.

[0228] In some embodiments, displaying the camera capture user interface (e.g., when in the first mode or the second mode) includes: displaying a set of one or more tracking elements (e.g., 1114) based on determining that a set of one or more tracking criteria are satisfied, wherein the set of one or more tracking criteria includes criteria satisfied when determining that the portion of the physical environment includes a first type of object (e.g., a face, a hand, and / or a person), wherein: displaying the set of one or more tracking elements at a location in the camera capture user interface that is based on a location of the first type of object in the physical environment (e.g., a face, a hand, and / or a person). Figure 11A and Figure 11F ); and the position (e.g., position in the z-axis, x-axis, and / or y-axis) of the one or more tracking elements displayed in the camera-captured user interface shifts as the position of the first type of object in the physical environment shifts (e.g., as Figure 11A 、 Figure 11B and Figure 11F). In some embodiments, a displayed size of the set of one or more tracking elements is based on the detected size of the first type of object (e.g., such that the size of the set of one or more tracking elements changes as the size of the first type of object changes (e.g., due to the object moving closer to or farther away from the computer system). Displaying the set of one or more tracking elements when the set of one or more tracking criteria is met performs an operation when a set of conditions has been met without requiring further user input. Doing so also provides improved visual feedback about objects detected in the physical environment and helps the user compose media capture events that include the first type of object.

[0229] In some embodiments, the set of one or more tracking elements are locked to (e.g., the positions of the set of one or more tracking elements are displayed as the object to which they are locked shifts (in some embodiments, the portion of the physical environment is represented with stereo depth and the set of one or more tracking elements are displayed as offset relative to their distance from the user within the user's viewpoint and / or within the physical environment (e.g., offset in z-space) because the object to which they are locked is offset in distance relative to the user) a set of one or more environment-locked virtual objects of the first type of object (e.g., the face of a subject within the user's viewpoint). Displaying the set of one or more tracking elements as environment-locked objects provides improved visual feedback about the position of the first type of object and performs an operation (e.g., moving the displayed position of the tracking element as the object shifts position) without requiring further user input when a set of conditions have been met.

[0230] In some embodiments, the first representation is a live feed of the field of view of at least the first camera of the one or more cameras (e.g., a continuously updated representation based on the field of view of the first camera; an instantaneous / simultaneous live feed or a delayed live feed) (e.g., Figure 11A 、 Figure 11B and Figure 11F ).

[0231] In some embodiments, while displaying the first representation at a fourth size, the computer system receives a request to modify the size of the first representation. In response to the request to modify the size of the first representation, the computer system: based on determining that the request to modify the size of the first representation is a request to expand the size of the first representation (in some embodiments, and based on determining that the fourth size is a sixth size (e.g., the first representation is currently at a reduced size)), expands the size of the first representation from the fourth size to a fifth size that is larger than the fourth size (e.g., as shown in FIG. 1 ). Figures 11F to 11G); and based on determining that the request to modify the size of the first representation is a request to reduce the size of the first representation (in some embodiments, and based on determining that the fourth size is the fifth size (e.g., the first representation is currently at an enlarged size)), reducing the size of the first representation from the fourth size to a sixth size that is smaller than the fourth size (e.g., as discussed in reference 1104b2).

[0232] In some embodiments, while displaying the first representation (e.g., while in the second mode), the computer system receives a fourth media capture request (e.g., a gesture on a touch-sensitive surface; an air gesture performed with a hand of a user of the computer system; an actuation of a hardware button or key; a gaze-based input; and / or a voice command); and in response to receiving the fourth media capture request: capturing third media (e.g., still media and / or video media; spatial media / stereoscopic media; or two-dimensional media or three-dimensional media) via the one or more cameras; and displaying an animation that includes modifying the visual appearance of at least a portion of the first representation (e.g., as Figure 11H , and as discussed with reference to 1118) (e.g., including an animation that brightens (e.g., whitens) an area including the first representation and / or adjacent to the first representation). Displaying an animation that includes modifying the visual appearance of at least a portion of the first representation upon capture provides improved visual feedback regarding the capture event.

[0233] In some embodiments, while displaying the first representation (e.g., while in the second mode), the computer system modifies the visual appearance of at least the second sub-portion of the physical environment (e.g., as Figure 11G ) (e.g., obscuring, darkening, applying a mask over the portion, and / or applying a tint layer over the portion) (in some embodiments, at least some visual details of the second subportion remain visible). In some embodiments, the second subportion corresponds to the portion of the camera capture user interface displaying the first representation; in some embodiments, the second subportion includes or overlaps with the first subportion. Modifying the visual appearance of at least the second subportion of the physical environment performs an operation when a set of conditions have been met, without requiring further user input. Doing so also improves the visibility of the first representation by reducing potential interference from visual elements of the physical environment (e.g., bright lights).

[0234] In some embodiments, the first media capture request is a request to capture stereoscopic media; and the second media capture request is a request to capture stereoscopic media. In some embodiments, spatial media includes spatial visual media (also known as stereoscopic media) and / or spatial audio. In some embodiments, spatial capture is the capture of spatial media. In some embodiments, spatial visual media (e.g., spatial images and / or spatial video) is media that includes two different images or two sets of images representing two perspectives with the same or overlapping fields of view for simultaneous display. A first image representing a first perspective is presented to a first eye of a viewer, and a second image representing a second perspective different from the first perspective is simultaneously presented to a second eye of the viewer. The first image and the second image have the same or overlapping fields of view. In some embodiments, the computer system displays the first image via a first display positioned for viewing by the viewer's first eye, and simultaneously displays the second image via a second display different from the first display positioned for viewing by the viewer's second eye. In some embodiments, when viewed together, the first image and the second image create a depth effect and provide the viewer with a depth perception of the content of the images. In some embodiments, a first video representing a first perspective is presented to a first eye of a viewer, and a second video representing a second perspective different from the first perspective is simultaneously presented to a second eye of the viewer. The first video and the second video have the same or overlapping fields of view. In some embodiments, when viewed together, the first video and the second video create a depth effect and provide the viewer with a depth perception of the content of the videos.

[0235] In some embodiments, displaying the camera capture user interface includes displaying a first instance of the camera capture user interface. In some embodiments, after stopping displaying the first instance of the camera capture user interface (e.g., after closing the capture application that generates the camera capture user interface), the computer system receives a request to display a second instance of the camera capture user interface (e.g., a request to start the capture application). In response to receiving the request to display the second instance of the camera capture user interface, the computer system displays the second instance of the camera capture user interface via the display generation component, wherein displaying the second instance of the camera capture user interface includes: displaying the second instance of the camera capture user interface in the first mode based on determining that the previous (e.g., the immediately previous) instance of the camera capture user interface was in the first mode when the display of the camera capture user interface was stopped; and displaying the second instance of the camera capture user interface in the second mode based on determining that the previous instance of the camera capture user interface was in the second mode when the display of the previous instance of the camera capture user interface was stopped (e.g., as described in reference to FIG. 1 ). Figure 11IIn some embodiments, the mode state of the camera capture user interface is persistent between sessions in which the interface is displayed. Making the mode of the respective user interface persistent reduces the number of inputs required to configure the respective user interface to a user's likely preferred mode, and also reduces operations when a set of conditions have been met (e.g., configuring the respective user interface to a previously used mode) without requiring further user input.

[0236] Figures 13A to 13K An example of a technique for interacting with a graphical user interface using gaze is shown. Figure 14 is a flow diagram of an exemplary method 1400 for interacting with a graphical user interface using gaze. Figures 13A to 13K The user interface in the diagram is used to illustrate the process described below, including Figure 14 in the process.

[0237] exist Figure 13A At , the device 700 displays a user interface 1302a including a background 1301 on the display 700a. The user interface 1302a includes container virtual objects 1304a to 1304c. In some embodiments, the container virtual object is a folder that organizes one or more files and / or other digital items. In some embodiments, the container virtual object is an album that organizes one or more media items (e.g., photos, videos, and / or audio media). For example, the container virtual object 1304a may be an album of photos and videos taken in 2020, the container virtual object 1304b may be an album of favorite photos and videos, and so on. Figure 13A , each container virtual object includes an extension control, such as the extension control 1304a1 of the container virtual object 1304a. In some embodiments, the background 1301 is a representation of a physical environment, and the user interface 1302a is an XR user interface that presents virtual objects (such as the container virtual object) overlaid on portions of the representation of the physical environment. In such embodiments, one or more virtual objects of the user interface 1302a are viewpoint-locked virtual objects and / or environment-locked virtual objects. Figure 13A At , the device 700 detects (e.g., via the eye tracking device 130 and / or the sensor 306) that the gaze of the user of the device 700 is directed toward the expanded space 1304a1 of the container virtual object 1304a, as indicated by the gaze indication 1310a. In some embodiments, detecting that the user's gaze is directed toward the virtual object includes determining that the user's gaze has remained on the virtual object for a predetermined period of time (e.g., 0.1 seconds, 0.25 seconds, 0.5 seconds, or 1 second).

[0238] exist Figure 13BIn response to detecting that the gaze of the user of device 700 is directed toward the expansion control 1304a1 of the container virtual object 1304a, device 700 displays user interface 1302b, which is also displayed on background 1301. In some embodiments, user interface 1302a and user interface 1302b are user interfaces of the same application (e.g., a media viewer and / or manager application). User interface 1302b includes a media viewer virtual object 1306 that presents an ordered set of media item representations, the ordered set of media item representations being displayed in the virtual object. Figure 13B 1306A1 through 1306A6. In some embodiments, representations 1306A1 through 1306A6 are representations of XR experiences and / or applications that can be executed on device 700. User interface 1302b also includes a scroll bar 1308, which includes a movable slider 1308a within a slot 1308b. The media collection of container virtual object 1304a includes additional media items in addition to those represented by 1306A1 through 1306A6, and scroll bar 1308 can be used to quickly navigate within the collection. Figure 13B , slider 1308a is at the top of slot 1308b, indicating that representations 1306A1 through 1306A6 correspond to the first media item in the media collection (e.g., the first media item in the ordered group) of container virtual object 1304a. User interface 1302b also includes a collapse control 1312 that, when selected, causes user interface 1302a to be redisplayed. Figure 13B In some embodiments, representations 1306A1 through 1306A3 and 1306A5 through 1306A6 are presented in a first visual manner, while representation 1306A4 is presented in a second visual manner that is different from the first visual manner. For example, representations 1306A1 through 1306A3 and 1306A5 through 1306A6 are presented at a first size, while representation 1306A4 is presented at a second, larger size. In some embodiments, representation 1306A4 is presented with a stereoscopic effect, animation effect, and / or filter effect that is not applied to representations 1306A1 through 1306A3. In some embodiments, representations 1306A1 through 1306A3 are presented with a stereoscopic effect, animation effect, and / or filter effect that is not applied to representation 1306A4. In some embodiments, representation 1306A4 is presented in a different second visual manner based on characteristics of the media item represented by representation 1306A4 (e.g., the item is the most recent item, the item is a stereo media item, and / or the item is a favorite item) or based on the position of the media item within the ordered group (every fourth media item is presented in the second visual manner). In some embodiments, only a predetermined number (e.g., 1, 2, or 3) of representations are presented in the second visual manner in user interface 1302b at any given time. Figure 13BAt , device 700 detects that the gaze of the user of device 700 is directed toward representation 1306A6, as indicated by gaze indication 1310b, which is interpreted as a request to move representation 1306A6 to a predetermined position (e.g., a horizontally centered position) within the media viewer virtual object 1306.

[0239] exist Figure 13C At , in response to detecting that the gaze of the user of device 700 is directed toward representation 1306A6, device 700 begins to shift the representation displayed within media viewer virtual object 1306 upward (e.g., the relative positioning of the media items remains consistent as the representation shifts) so that representation 1306A6 is closer to horizontal center, and the device updates the position of slider 1308a of scroll bar 1308 to reflect the shift. In some embodiments, device 700 shifts the representations in media viewer virtual object 1306 at a predetermined speed while device 700 continues to detect the user's gaze directed toward a given representation until the representation reaches the predetermined position. Figure 13C At , the set of representations has shifted so that representation 1306A6 is closer to being horizontally centered, but representation 1306A6 is not yet horizontally centered. Figure 13C At , device 700 detects that the gaze of the user of device 700 continues to be directed toward representation 1306A6, as indicated by gaze indication 1310c.

[0240] exist Figure 13D In response to detecting that the gaze of the user of device 700 continues to be directed toward representation 1306A6, device 700 shifts the representation displayed within media viewer virtual object 1306 upward, thereby bringing representation 1306A6 to a horizontally centered position, and the device updates the position of slider 1308a of scroll bar 1308 to reflect this further shift. In some embodiments, if device 700 detects that the user's gaze is no longer directed toward representation 1306A6, then at Figure 13C At , the device 700 will stop offsetting the representation displayed within the media viewer virtual object 1306. In some embodiments, once representation 1306A6 is displayed at the horizontally centered position, the device 700 will stop offsetting the representation displayed within the viewer virtual object 1306 even if the device 700 continues to detect that the user of the device 700 is looking toward representation 1306A6. Figure 13D At this point, device 700 detects that the gaze of the user of device 700 is now directed toward representation 1306A6, as indicated by gaze indication 1310d, which is interpreted as a request to move representation 1306A5 to the predetermined position (e.g., the horizontally centered position) within the media viewer virtual object 1306.

[0241] exist Figure 13EAt , in response to detecting that the gaze of the user of device 700 is directed toward representation 1306A5, device 700 begins to shift the representation displayed within media viewer virtual object 1306 downward so that representation 1306A5 is more nearly horizontally centered, and the device updates the position of slider 1308a of scroll bar 1308 to reflect the shift. Figure 13E At , the device 700 detects that the user's gaze is no longer directed toward representation 1306A5 and therefore does not continue to shift the representation displayed within the media viewer virtual object 1306 downward. 13B to 13E As shown in , a user can cause the device 700 to shift a target representation to a horizontally centered position by directing their gaze toward the target representation. Using these techniques, a user can scroll through a collection of media items without having to access manual controls, and with varying degrees of precision and speed. In embodiments where the device 700 is an HMD, the user is able to avoid having to access manual controls that may not be visible when operating the device 700, and also free up his or her hands to interact with the environment or other devices. Figure 13E At this point, the device 700 detects that the gaze of the user of the device 700 is now directed toward a portion of the slot 1308b that is substantially below the current position of the slider 1308a, as indicated by the gaze indication 1310e, which is interpreted as a request to navigate to a location within the media collection of the container virtual object 1304a that corresponds to the location of the slot 1308b.

[0242] exist Figure 13F In response to detecting that the gaze of the user of the device 700 is directed to the portion of the slot 1308b indicated by the gaze indication 1310e, the device 700 displays representations 1306A45 to 1306A50 that are further away in the ordered group of the media collection of the container virtual object 1304a. In some embodiments, the device 700 immediately transitions to the display of representations 1306A45 to 1306A50. In some embodiments, the device 700 displays a fast (e.g., faster than the media viewer virtual object 1306) displayed within the media viewer virtual object 1306. 13B to 13E The device 700 also updates the position of the slider 1308a to reflect the current position within the media collection of the container virtual object 1304a. Thus, the user can gaze at the scroll bar 1308 to make faster and / or more obvious shifts within the media collection of the container virtual object 1304a, while directly gazing at a specific representation to make more subtle and / or slower shifts. Figure 13F At , device 700 detects that the gaze of the user of device 700 is directed toward the top portion of slot 1308b, as indicated by gaze indication 1310f, which is interpreted as a request to navigate to the beginning of the media collection of container virtual object 1304a corresponding to the top of slot 1308b.

[0243] exist Figure 13G In response to detecting that the gaze of the user of device 700 is directed to the portion of slot 1308b indicated by gaze indication 1310f, device 700 redisplays representations 1306Aa through 1306A6 at the beginning of the ordered group representation of the media collection of container virtual object 1304a. Device 700 also updates the position of slider 1308a to indicate the offset. Figure 13G At this point, the device 700 detects that the gaze of the user of the device 700 is directed toward the representation 1306A1, as indicated by the gaze indication 1310g, which is interpreted as a request to move the representation 1306A1 to a predetermined position within the media viewer virtual object 1306 (e.g., a horizontally centered position or a position within a threshold distance of the horizontal center).

[0244] exist Figure 13H , in response to detecting that the gaze of the user of device 700 is directed toward representation 1306A1, device 700 has offset the representation displayed within media viewer virtual object 1306 downward, thereby bringing representation 1306A1 to a horizontally centered position. Device 700 also displays a blank area 1314, which is an area beyond the edge of the ordered arrangement of representations of the media collection of container virtual object 1304a. In some embodiments, doing so provides a visual indication to the user that the end of the collection has been reached, and also provides feedback that device 700 has detected the user's gaze. In some embodiments, device 700 shortens the length of slider 1308a to indicate that a smaller portion of the representation of the media collection of container virtual object 1304a is currently being displayed (e.g., with Figure 13G ). Figure 13G At , device 700 detects that the gaze of the user of device 700 continues to be directed toward representation 1306A1 , as indicated by gaze indication 1310h .

[0245] exist Figure 13I In response to detecting that the gaze of the user of device 700 continues to be directed toward representation 1306A1, device 700 maintains display of representation 1306A1 at a horizontally centered position and maintains display of blank area 1314. Figure 13I At , the device 700 detects that the gaze of the user of the device 700 is no longer directed toward representation 1306A1 (eg, no longer directed toward any representation in the media viewer virtual object 1306).

[0246] exist Figure 13JIn response to detecting that the gaze of the user of device 700 is no longer directed toward representation 1306A1, device 700 shifts the representations displayed in media viewer virtual object 1306 upward so that the top edges of representations 1306A1 through 1306A3 are at the top edge of media viewer virtual object 1306, thereby causing blank area 1314 to no longer be displayed. In some embodiments, when device 700 determines that the user no longer requests that representation 1306A1 be displayed at the predetermined horizontally centered position, the effect is referred to as a bounce effect or a rubber band effect, and doing so optimizes the use of the display area. Figure 13J At , the device 700 detects input 1316a (e.g., actuation of button 700b) while the device 700 detects a gaze-directed representation 1306A5 of the user of the device 700, as indicated by gaze indication 1310i. In some embodiments, the input 1316a is an air gesture, a touch on a touch-sensitive surface, or speech input.

[0247] exist Figure 13K At, in response to detecting input 1316a when device 700 detects that the gaze of the user of device 700 is directed toward representation 1306A5, device 700 displays an enlarged view 1316 of representation 1306A5. In some embodiments, in addition to being larger than representation 1306A5, enlarged view 1316 is displayed with stereo effects, animation effects, and / or filter effects that are not applied to representation 1306A5 (e.g., representation 1306A5 is displayed as a two-dimensional object and enlarged view 1316 is displayed as a three-dimensional object). In some embodiments, additional input (e.g., 1316a) other than gaze is required to enlarge a representation because, when browsing a representation in the media viewer virtual object 1306, the user's gaze typically rests on the representation without the user having the intention to display the enlarged view. In some embodiments, reference Figures 13A to 13K The discussed techniques provide a user interface and control scheme for navigating through a collection of items (eg, media items) without having to use manual controls to navigate to a representation / item of interest.

[0248] about Figures 13A to 13K See below for additional description of Figures 13A to 13K Method 1400 is described.

[0249] Figure 14 is a flow chart of an exemplary method 1400 for interacting with a graphical user interface using gaze, according to some embodiments. In some embodiments, the method 1400 is performed on a computer system (e.g., Figure 1In some embodiments, method 1400 is performed at a computer system 101 in a computer system; a head-mounted display; an optical head-mounted display; a personal computer; a smart phone; and / or a tablet computer) that communicates with one or more gaze tracking sensors (e.g., an optical and / or IR camera configured to track the direction of gaze of a user of the computer system; an eye tracking device 130; and / or a sensor 306) and a display generation component (e.g., a display generation component 120; a display controller; a touch-sensitive display system; a pass-through display (e.g., integrated and / or connected), a 3D display, a transparent display, a projector, a head-up display, and / or a head-mounted display). In some embodiments, method 1400 is performed by storing in a non-transitory (or transient) computer-readable storage medium and by one or more processors of a computer system (such as one or more processors 202 of computer system 101) (e.g., Figure 1 Some operations in method 800 may be optionally combined, and / or the order of some operations may be optionally changed.

[0250] The computer system (e.g., 700) displays (1402) a corresponding user interface (e.g., 1302b) via the display generation component (e.g., 700a), wherein displaying the corresponding user interface includes: displaying a group of one or more virtual objects (e.g., 1306A1 to 1306A6), the group of one or more virtual objects including a first virtual object (e.g., 1306A6) (e.g., a media item (e.g., a representation of a photo or video); an icon; and / or a text box), the first virtual object being displayed at a first position (in some embodiments, a first position in the corresponding user interface; in some embodiments, a position along an edge of the displayable area) within a displayable area (e.g., an area of 700a and / or 1306) where the display generation component can display content (e.g., a displayable area of a display screen, an area where a projector can project content).

[0251] When the first virtual object is displayed at the first position within the displayable area, the computer system detects (1404) via the one or more gaze tracking sensors that the gaze of the user of the computer system is directed toward the first virtual object (e.g., as indicated by 1310b) (e.g., directed toward a direction corresponding to the gaze of the user that intersects with the first virtual object) (in some embodiments, directed toward the first virtual object for at least a predetermined time period (e.g., 0.01 seconds, 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.25 seconds, 0.5 seconds, or 1 second)).

[0252] In response to detecting that the gaze of the user of the computer system is directed toward the first virtual object, the computer system moves the first virtual object from the first position within the displayable area toward a second position within the displayable area that is different from the first position (for example, Figure 13C The horizontally centered position discussed above) (e.g., a predetermined position) (in some embodiments, the second position is the center of the displayable area) is moved (1406) (e.g., as Figure 13C ).

[0253] While moving the first virtual object toward the second position within the displayable area and before the first virtual object reaches the second position, the computer system detects (1408) the movement of the gaze via the one or more gaze tracking sensors.

[0254] In response (1410) to detecting the movement of the gaze, the computer system: continues to move the first virtual object toward the second position (e.g., as shown in FIG. 14A ) based on determining (1412) that the gaze of the user of the computer system continues to be directed toward the first virtual object (e.g., the user's gaze is tracking the first virtual object as the first virtual object moves). Figure 13D and stopping moving the first virtual object toward the second position within the displayable area (e.g., as discussed in reference to FIG); ... Figure 13C (in some embodiments, stopping the movement of the first virtual object within the displayable area). Moving the first virtual object from the first position within the displayable area toward a second position based on the gaze brings the object currently focused on by the user (e.g., as indicated by the gaze) to a predetermined second position in the displayable area (e.g., a more central area in the displayable area). Doing so also provides improved visual feedback regarding the currently detected location of the user's gaze.

[0255] In some embodiments, before detecting the movement of the gaze, while moving the first virtual object toward the second location and before the first virtual object reaches the second location, the computer system: based on determining that the gaze of the user of the computer system is substantially stationary and continues to be directed toward the first virtual object (e.g., the user's gaze tracks the first virtual object as it moves), continues to move the first virtual object toward the second location (e.g., as described in reference to FIG). Figure 13DIn some embodiments, based on determining that the gaze of the user of the computer system has stopped being directed toward (e.g., is no longer directed toward) the first virtual object, the first virtual object is stopped from moving toward the second position within the displayable area (in some embodiments, movement of the first virtual object within the displayable area is stopped). Continuing to move the first virtual object from the first position within the displayable area toward the second position based on continued gaze directed toward the first virtual object brings the object currently focused by the user (e.g., as indicated by the gaze) to a predetermined second position in the displayable area (e.g., a more central area in the displayable area). Doing so also provides improved visual feedback regarding the currently detected location of the user's gaze.

[0256] In some embodiments, the group of one or more virtual objects includes a second virtual object (e.g., 1306A5); before detecting that the gaze of the user of the computer system is directed toward the first virtual object and before moving the first virtual object from the first position to the second position within the displayable area, arranging the first virtual object and the second virtual object in a first predetermined spatial relationship (e.g., a grid in which the first virtual object and the second virtual object are in predetermined positions relative to each other); and moving the first virtual object from the first position toward the second position within the displayable area includes: moving the second virtual object so that the first virtual object and the second virtual object maintain the first predetermined spatial relationship (e.g., as Figure 13C Moving the first virtual object in a manner that maintains the first predetermined spatial relationship between the first virtual object and the second virtual object helps the user track objects by preserving the existing spatial relationship, which enhances device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0257] In some embodiments, the second position within the displayable area is a substantially central position along at least one axis (e.g., x-axis, y-axis, or z-axis) of the displayable area (e.g., as shown in FIG. FIG. 13C to FIG. 13DIn some embodiments, the second position is not a substantially central position on at least one other axis of the displayable area (e.g., the second position is centered about the x-axis rather than the y-axis, or vice versa). In some embodiments, moving the first virtual object occurs at a first predetermined speed (e.g., a speed that is not based on the direction of the user's gaze). Moving the first virtual object from the first position within the displayable area to a central position based on gaze brings the object currently focused by the user into a more central portion of the displayable area, making it easier for the user to interact with the object. Doing so also provides improved visual feedback about the currently detected location of the user's gaze.

[0258] In some embodiments, after detecting that the gaze of the user of the computer system is directed toward the first virtual object and after moving the first virtual object from the first position toward the second position within the displayable area, the computer system detects via the one or more gaze tracking sensors that the gaze of the user of the computer system has moved from pointing toward the first virtual object to pointing toward (e.g., pointing in a direction corresponding to the user's gaze that intersects with the third virtual object) (in some embodiments, pointing toward the third virtual object for at least a predetermined period of time (e.g., 0.01 seconds, 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.25 seconds, 0.5 seconds, or 1 second)) a third virtual object (e.g., Figure 13D In response to detecting that the gaze of the user of the computer system has moved from being directed at the first virtual object to being directed at the third virtual object, the computer system moves the third virtual object from the third position within the displayable area toward a fourth position (e.g., a predetermined position) within the displayable area that is different from the third position (in some embodiments, the fourth position is the center of the displayable area) (e.g., as shown in FIG. 1306A5). Figure 13E ).

[0259] In some embodiments, displaying the corresponding user interface includes: displaying an interactive navigation virtual object (e.g., 1308) (e.g., a scroll bar, a drag bar, an index bar, or a document or interface map) that, when selected via a first input (e.g., via a gesture on a touch-sensitive surface; an air gesture performed with a hand of a user of the computer system; an actuation of a hardware button or key; a gaze-based input; and / or a voice command), causes the computer system to navigate to a location within the corresponding user interface based on characteristics of the first input (e.g., a speed of movement of the first input (e.g., the first input includes a movement component (e.g., a slide, hold, and drag)); a location of the first input within the interactive scroll virtual object (e.g., a click at a location along the groove / track of the scroll bar); a direction of the first input and / or a duration of the input (e.g., maintaining gaze at a location for a period of time)). Figures 13E to 13F Displaying interactive navigational virtual objects that allow navigation through the corresponding user interface based on input characteristics provides the user with an input mechanism / scheme for (potentially) quickly navigating through the corresponding user interface, which enhances the operability of the device and makes the user-device interface more efficient, thereby reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0260] In some embodiments, detecting that the gaze of the user of the computer system is directed toward the first virtual object includes: detecting that the gaze of the user of the computer system has been directed toward the first virtual object for more than a first predetermined period of time (e.g., 0.01 seconds, 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.25 seconds, 0.5 seconds, or 1 second) (e.g., as described in reference Figure 13A discussed).

[0261] In some embodiments, detecting that the gaze of the user of the computer system is directed toward the first virtual object includes: detecting a first type of air gesture (e.g., an air pinch, an air tap, or an air double tap) when the gaze of the user of the computer system is directed toward the first virtual object.

[0262] In some embodiments, the computer system receives a request to enlarge the first virtual object (e.g., 1310i and 1316a). In response to receiving the request to enlarge the first virtual object, the computer system enlarges the first virtual object from a first size (and, in some embodiments, stops displaying other objects in the group of one or more virtual objects) to a second size that is larger than the first size, wherein: when displayed at the second size, the first virtual object is displayed as a three-dimensional object (in some embodiments, when displayed at the second size, the first virtual object is a stereoscopic object (e.g., the display generation component presents the object to the user's right eye differently than the object presented to the user's left eye)); and when displayed at the first size, the first virtual object is displayed as a two-dimensional object (in some embodiments, when displayed at the first size, the first visual object is a non-stereoscopic object) (e.g., as described in reference 1316 and Figure 13K Displaying the first virtual object as a three-dimensional object when the first virtual object is expanded (eg, at the second size) provides improved visual feedback regarding the received request to enlarge the first virtual object.

[0263] In some embodiments, the group of one or more virtual objects includes: a fourth virtual object corresponding to a first two-dimensional media item (e.g., 1306A1) (e.g., a photo or video); a fifth virtual object (e.g., 1306A2) corresponding to a second two-dimensional media item; a sixth virtual object corresponding to a first three-dimensional media item (e.g., 1306A4) (e.g., a stereoscopic photo or stereoscopic video); and a seventh virtual object (e.g., 1306A7) corresponding to the second three-dimensional media item. The fourth virtual object and the fifth virtual object are both displayed with a first type of visual appearance (e.g., visual treatment; display style or theme (e.g., color, pattern, brightness level)). In some embodiments, all virtual objects corresponding to 2D media items are displayed using the same visual appearance / treatment. The sixth virtual object and the seventh virtual object are both displayed with a second type of visual appearance that is different from the first type of visual appearance (e.g., different size, different brightness or contrast, different lighting effects (e.g., glow effects), presence or absence of a border or difference in border appearance, and / or different shape). In some embodiments, all virtual objects corresponding to 3D media items are displayed using the same visual appearance / processing. Displaying virtual objects corresponding to three-dimensional media items with a different visual appearance than virtual objects corresponding to two-dimensional media items provides improved visual feedback about the properties of the media items.

[0264] In some embodiments, after moving the first virtual object from the first position toward the second position within the displayable area and while the first virtual object is at the second position, the computer system detects via the one or more gaze tracking sensors that the gaze of the user of the computer system is directed toward the first virtual object (e.g., Figure 13I In response to detecting that the gaze of the user of the computer system is directed toward the first virtual object and when the first virtual object is at the second position, the computer system maintains the first virtual object at the second position (e.g., as shown in FIG. Figure 13I Once the first virtual object completes its movement, maintaining the first virtual object at the second position allows the user to continue interacting with (e.g., viewing) the first virtual object without having to track the moving object, which enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently. Doing so also provides improved visual feedback about the object's position relative to the second position.

[0265] In some embodiments, in response to detecting that the gaze of the user of the computer system is directed toward the first virtual object, the computer system displays an eighth virtual object (e.g., 1306A7) (in some embodiments, an animation of the eighth virtual object transitioning into the displayable area is displayed), wherein the eighth virtual object is not displayed when the first virtual object is displayed at the first position. In some embodiments, the first virtual object and the eighth virtual object are part of a plurality of virtual objects arranged in a predetermined spatial arrangement, wherein only a subset of the plurality of virtual objects is displayed at a given time.

[0266] In some embodiments, the computer system detects a first input (e.g., 1316a) while the gaze of the user of the computer system is directed toward the first virtual object, wherein the first input is selected from the group consisting of: an air gesture (in some embodiments, an air gesture detected via one or more sensors of an external electronic device in communication with the computer system (e.g., a smartwatch or smart phone)), actuation of a hardware input mechanism in communication with the computer system (e.g., an external or integrated button, dial, or switch), and continued detection of the gaze of the user of the computer system being directed toward the first virtual object for more than a second predetermined time period (e.g., 0.01 seconds, 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.25 seconds, 0.5 seconds, or 1 second). In response to detecting the first input while the gaze of the user of the computer system is directed toward the first virtual object, the computer system performs a first operation on the first virtual object (e.g., Figure 13K ) (e.g., magnifying the object, focusing / selecting the object, and / or modifying the appearance of the object).

[0267] In some embodiments, the group of one or more virtual ob...

Claims

1. A method comprising: At a computer system in communication with one or more gaze tracking sensors and a display generation component: Displaying, via the display generation component, a corresponding user interface, wherein displaying the corresponding user interface comprises displaying: a plurality of edges of the corresponding user interface including a first edge and a second edge different from the first edge; a first user interface object positioned along the first edge corresponding to a first operation; and a second user interface object positioned along the second edge corresponding to a second operation different from the first operation; detecting, via the one or more gaze tracking sensors, that a gaze of a user of the computer system is directed toward respective portions of the respective user interfaces while displaying the first user interface object and the second user interface object; as well as In response to detecting that the gaze of the user of the computer system is directed toward the corresponding portion of the corresponding user interface: According to determining that the corresponding portion of the corresponding user interface corresponds to the first user interface object: performing the first operation; and continuing to display the first user interface object while ceasing to display the second user interface object; According to determining that the corresponding portion of the corresponding user interface corresponds to the second user interface object: performing the second operation; and Continue to display the second user interface object while stopping displaying the first user interface object.

2. The method according to claim 1, wherein: The corresponding user interface is an extended reality user interface; and Displaying the corresponding user interface includes displaying: a representation of the physical environment; and A third user interface object, wherein the third user interface object is an environment-locked virtual object.

3. The method according to any one of claims 1 to 2, wherein: Continuing to display the first user interface object while ceasing to display the second user interface object includes: visually emphasizing the first user interface object; and Continuing to display the second user interface object while ceasing to display the first user interface object includes visually emphasizing the second user interface object.

4. The method according to any one of claims 1 to 2, wherein performing the first operation comprises: The first user interface object is expanded from an unexpanded state to a first expanded state by expanding at least a first portion of the first user interface object, wherein the first expanded state of the first user interface object includes a first control object and first information that are not included in the unexpanded state of the first user interface object.

5. The method according to claim 4, wherein: The first control object, when selected, causes the computer system to perform an operation related to display brightness; and The first information corresponds to brightness information.

6. The method according to claim 4, wherein: The first control object, when selected, causes the computer system to perform an operation related to audio output volume; and The first information corresponds to volume information.

7. The method according to claim 4, wherein: the first control object, when selected, causes the computer system to perform an operation related to an energy storage component of the computer system; and The first information corresponds to energy storage information.

8. The method according to claim 4, further comprising: detecting, via the one or more gaze tracking sensors, that the gaze of the user of the computer system is directed toward the first control object; In response to detecting that the gaze of the user of the computer system is directed toward the first control object, the first control object is expanded from a first control object unexpanded state to a first control object expanded state.

9. The method according to claim 8, wherein: When the first user interface object is displayed with the first control object in the first control object unexpanded state, the first user interface object includes a second control object in a second control object expanded state, and Expanding the first control object from the first control object unexpanded state to the first control object expanded state includes contracting the second control object from the second control object expanded state to the second control object unexpanded state.

10. The method according to any one of claims 8 to 9, wherein the first control object includes a third control object when in the first control object expanded state, and the third control object is not included in the first control object when the first control object is in the first control object unexpanded state.

11. The method according to any one of claims 8 to 9, wherein the first control object includes second information when in the first control object extended state, and the second information is not included in the first control object when the first control object is in the first control object unextended state.

12. The method of any one of claims 1 to 2, wherein the first control object, when selected, causes the display generating component to transition from a first mode to a second mode.

13. The method of any one of claims 1 to 2, wherein the first control object, when selected, causes the computer system to disable a set of one or more functions activated by detecting the gaze of the user of the computer system.

14. The method of claim 13, wherein causing the computer system to disable the set of one or more functions activated by detecting the gaze of the user of the computer system comprises: disabling a first function that is activated when the computer system detects that the gaze of the user of the computer system is directed toward a first location of the corresponding user interface; as well as A second function is maintained available for activation, the second function being activated when the computer system detects that the gaze of the user of the computer system is directed toward a second location of the corresponding user interface.

15. The method according to claim 13, further comprising: displaying an indication that the set of one or more functions activated by detecting the gaze of the user of the computer system are available for activation based on determining that the set of one or more functions activated by detecting the gaze of the user of the computer system are available for activation; as well as Based on determining that the set of one or more functions activated by detecting the gaze of the user of the computer system are disabled, displaying an indication that the set of one or more functions activated by detecting the gaze of the user of the computer system are disabled.

16. The method of any one of claims 1 to 2, wherein the corresponding user interface further comprises a current time indicator.

17. The method according to any one of claims 1 to 2, wherein: The corresponding user interface includes a plurality of application user interface objects displayed in a first spatial arrangement; and The first control object, when selected, causes the plurality of application user interface objects to transition from being displayed in the first spatial arrangement to being displayed in a second spatial arrangement different from the first spatial arrangement.

18. The method of any one of claims 1 to 2, wherein the corresponding user interface further comprises a first representation of an application user interface of a first application.

19. The method according to claim 18, further comprising: detecting a first input corresponding to the first representation while the first representation is displayed at a first location in the corresponding user interface; In response to the first input, the first representation is moved to a second position in the respective user interface that is different from the first position.

20. The method of claim 19, wherein the first position is predefined and the second position is predefined.

21. The method according to claim 18, further comprising: detecting a second input corresponding to the first representation when the first representation is in a first representation unexpanded state; In response to detecting the first input, the first representation is expanded to a first representation expanded state, wherein the first representation is larger in the first representation expanded state than in the first representation unexpanded state.

22. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more gaze tracking sensors and a display generation component, the one or more programs comprising instructions for performing the method of any one of claims 1 to 21.

23. A computer system configured to communicate with one or more gaze tracking sensors and a display generation component, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for executing the method according to any one of claims 1 to 21.

24. A computer system configured to communicate with one or more gaze tracking sensors and a display generation component, the computer system comprising: Components for carrying out the method according to any one of claims 1 to 21.

25. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more gaze tracking sensors and a display generation component, the one or more programs comprising instructions for performing the method according to any one of claims 1 to 21.