Devices, methods, and graphical user interfaces for interacting with window controls in three-dimensional environment

By detecting user gaze input and interaction behavior in a computer system, dynamically managing user interface controls is solved, and the problems of low interaction efficiency and high power consumption in the prior art are achieved, and more intuitive and efficient user interface interaction is achieved.

CN119998764APending Publication Date: 2025-05-13APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380068172.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-21
Filing Date
2023-09-22
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The methods of interacting with virtual/augmented reality environments in the prior art are inefficient and insufficient feedback, resulting in a large cognitive burden on users and high power consumption.

Method used

By using display generation components and input devices in a computer system, detect user gaze input and user interface interactions, dynamically display and hide control elements to reduce the number of user inputs and improve interaction efficiency.

Benefits of technology

Achieve more intuitive and efficient user interface interaction, reduce battery consumption and improve the experience of virtual/augmented reality environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998764A_ABST
    Figure CN119998764A_ABST
Patent Text Reader

Abstract

A computer system displays a first object, the first object comprising at least a first portion of the first object and a second portion of the first object, and the computer system detects a first gaze input that meets a first criterion, where the first criterion requires the first gaze input to point to the first portion of the first object so as to meet the first criterion. In response, the computer system: displays a first control element corresponding to a first operation associated with the first object, where the first control element is not displayed prior to detecting that the first gaze input satisfies a first criterion; and detecting a first user input directed to the first control element. In response to detecting a first user input directed to the first control element, the computer system performs a first operation with respect to the first object.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related patent applications

[0002] This application claims priority to U.S. patent application No. 18 / 371,368 filed on September 21, 2023, U.S. patent application No. 18 / 371,372 filed on September 21, 2023, U.S. patent application No. 18 / 371,374 filed on September 21, 2023, U.S. patent application No. 18 / 371,378 filed on September 21, 2023, U.S. Provisional Patent Application No. 63 / 539,566 filed on September 20, 2023, U.S. Provisional Patent Application No. 63 / 535,012 filed on August 28, 2023, U.S. Provisional Patent Application No. 63 / 470,909 filed on June 4, 2023, and U.S. Provisional Patent Application No. 63 / 409,600 filed on September 23, 2022. Technical Field

[0003] The present disclosure generally relates to computer systems providing computer-generated experiences in communication with display generation components and optionally with one or more input devices, including but not limited to electronic devices providing virtual reality and mixed reality experiences via displays. Background Art

[0004] In recent years, the development of computer systems for augmented reality has increased significantly. Example augmented reality environments include at least some virtual elements that replace or augment the physical world. Input devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touch screen displays) for computer systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Some examples of virtual elements include virtual objects such as digital images, videos, text, icons, and control elements (such as buttons and other graphics). Summary of the invention

[0005] Some methods and interfaces for interacting with environments that include at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve desired results in augmented reality environments, and systems that make virtual object manipulation complex, cumbersome, and error-prone can impose a huge cognitive burden on users and detract from the experience of virtual / augmented reality environments. In addition, these methods take longer than necessary, wasting energy on the computer system. This latter consideration is particularly important in battery-powered devices.

[0006] Therefore, there is a need for computing systems with improved methods and interfaces to provide users with controls and user interface elements that indicate conditional display of information about content, thereby making interaction with a computer system more efficient and intuitive for the user. Such methods and interfaces optionally supplement or replace conventional methods for providing extended reality experiences to users. Such methods and interfaces reduce the amount, degree, and / or nature of inputs from a user by helping the user understand the connection between the inputs provided and the device's response to those inputs, thereby forming a more effective human-computer interface.

[0007] The above-mentioned defects and other problems associated with the user interface of the computer system are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a laptop, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device, such as a watch or a head-mounted device). In some embodiments, the computer system has a touch pad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touch screen" or "touch screen display"). In some embodiments, the computer system has one or more eye tracking components. In some embodiments, the computer system has one or more hand tracking components. In some embodiments, in addition to displaying a generating component, the computer system also has one or more output devices, which include one or more tactile output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, a memory, and one or more modules, a program or instruction set stored in the memory for performing multiple functions. In some embodiments, the user interacts with the GUI through contact and gestures of a stylus and / or finger on a touch-sensitive surface, movement of the user's eyes and hands in space relative to the GUI (and / or computer system) or the user's body (as captured by a camera and other mobile sensors), and / or voice input (as captured by one or more audio input devices). In some embodiments, the functions performed by interaction optionally include image editing, drawing, presentations, word processing, spreadsheet creation, playing games, making and receiving calls, video conferencing, sending and receiving emails, instant messaging, testing support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. Executable instructions for performing these functions are optionally included in a transient and / or non-transient computer-readable storage medium or other computer program product configured for execution by one or more processors.

[0008] There is a need for electronic devices with improved methods and interfaces for interacting with a three-dimensional environment. Such methods and interfaces can supplement or replace conventional methods for interacting with a three-dimensional environment. Such methods and interfaces reduce the amount, degree, and / or nature of input from a user and produce a more efficient human-computer interface. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charges.

[0009] A method is performed at a computer system in communication with a first display generation component and one or more input devices. The method includes: displaying a first object in a first view of a three-dimensional environment via the first display generation component, wherein the first object includes at least a first portion of the first object and a second portion of the first object. The method includes: while displaying the first object, detecting via the one or more input devices a first gaze input that satisfies a first criterion, wherein the first criterion requires that the first gaze input be directed to the first portion of the first object in order to satisfy the first criterion. The method includes: in response to detecting that the first gaze input satisfies the first criterion, displaying a first control element corresponding to a first operation associated with the first object, wherein the first control element is not displayed before detecting that the first gaze input satisfies the first criterion. The method includes: while displaying the first control element, detecting via the one or more input devices a first user input directed to the first control element, and in response to detecting the first user input directed to the first control element, performing the first operation relative to the first object.

[0010] A method is performed at a computer system in communication with a first display generation component and one or more input devices. The method includes: displaying, via the first display generation component, a first user interface object and a first control element associated with performing a first operation relative to the first user interface object in a first view of a three-dimensional environment, wherein the first control element is spaced apart from the first user interface object in the first view of the three-dimensional environment, and wherein the first control element is displayed with a first appearance. The method includes: while displaying the first control element with the first appearance, detecting, via the one or more input devices, a first gaze input directed to the first control element. The method includes: in response to detecting the first gaze input directed to the first control element, updating the appearance of the first control element from the first appearance to a second appearance different from the first appearance. The method includes: while displaying the first control element with the second appearance, detecting, via the one or more input devices, a first user input directed to the first control element. The method includes: in response to detecting the first user input pointing to the first control element, and based on determining that the first user input satisfies a first criterion, updating the appearance of the first control element from the second appearance to a third appearance, the third appearance being different from the first appearance and the second appearance and indicating that additional movement associated with the first user input will cause the first operation associated with the first control element to be performed.

[0011] A method is performed at a computer system in communication with a first display generation component and one or more input devices. The method includes: displaying a first application window and a first title bar of the first application window simultaneously via the first display generation component, wherein the first title bar of the first application window is separated from the first application window on a first side of the first application window and displays a corresponding identifier of the first application window. The method includes: when the first application window and the first title bar separated from the first application window are displayed together, detecting, via the one or more input devices, that a user's attention is directed to the first title bar. The method includes: in response to detecting that the user's attention is directed to the first title bar, based on determining that the user's attention meets a first criterion relative to the first title bar, expanding the first title bar to display one or more first selectable controls for interacting with a first application corresponding to the first application window, wherein the one or more first selectable controls are not displayed before the first title bar is expanded.

[0012] A method is performed at a computer system in communication with a first display generation component having a first display area and one or more input devices. The method includes: displaying a first application window of a first application at a first window location in the first display area via the first display generation component. The method includes: displaying a first indicator at a first indicator location in the first display area that has a first spatial relationship with the first application window as an indication that the first application is accessing the one or more sensors of the computer system based on determining that the first application is accessing the one or more sensors of the computer system. The method includes: while displaying the first indicator at the first indicator location in the first display area that has the first spatial relationship with the first application window, detecting a first user input corresponding to a request to move the first application window of the first application to a second window location in the first display area, the second window location being different from the first window location. The method includes: in response to detecting the first user input corresponding to the request to move the first application window of the first application from the first window position in the first display area to the second window position, displaying the first application window of the first application at the second window position in the first display area; and based on determining that the first application is accessing the one or more sensors of the computer system, displaying the first indicator at a second indicator position in the first display area that is different from the first indicator position in the first display area, wherein the second indicator position in the first display area has the first spatial relationship with the first application window displayed at the second window position.

[0013] A method is performed at a computer system in communication with a first display generation component and one or more input devices. The method includes: displaying a user interface, wherein displaying the user interface includes simultaneously displaying a content area having first content, a first user interface object, and a second user interface object in the user interface, wherein: the corresponding content in the content area is constrained to have an appearance in which a corresponding parameter is within a first value range, the first user interface object is displayed with an appearance in which the corresponding parameter has a value outside the first value range, and the second user interface object is displayed with an appearance in which the corresponding parameter has a value outside the first value range. The method includes: while simultaneously displaying the first content, the first user interface object, and the second user interface object, updating the user interface, including: changing the first content to the second content, while the corresponding content in the content area continues to be constrained to have an appearance in which the corresponding parameter is within the first value range; updating the appearance of the first user interface object and continuing to display the first user interface object with an appearance in which the corresponding parameter has a value outside the first value range; and updating the appearance of the second user interface object and continuing to display the second user interface object with an appearance in which the corresponding parameter has a value outside the first value range.

[0014] A method is performed at a computer system in communication with a first display generation component and one or more input devices. The method includes: displaying a first view of a three-dimensional environment via the first display generation component, the first view corresponding to a first viewpoint of a user. The method also includes: while displaying the first view of the three-dimensional environment corresponding to the first viewpoint of the user via the first display generation component, detecting a first event corresponding to a request to display a first virtual object in the first view of the three-dimensional environment. The method also includes: in response to detecting the first event corresponding to the request to display the first virtual object in the first view of the three-dimensional environment, displaying the first virtual object at a first location in the three-dimensional environment in the first view of the three-dimensional environment, wherein the first virtual object is displayed together with a first object management user interface corresponding to the first virtual object, and wherein the first object management user interface has a first appearance relative to the first virtual object at the first location in the three-dimensional environment. The method includes: detecting a first user input corresponding to a request to move the first virtual object in the three-dimensional environment via the one or more input devices. The method also includes: in response to detecting the first user input corresponding to the request to move the first virtual object in the three-dimensional environment: displaying the first virtual object at a second position in the three-dimensional environment that is different from the first position in the first view of the three-dimensional environment, wherein the first virtual object is displayed simultaneously with the first object management user interface at the second position in the three-dimensional environment, and wherein the first object management user interface has a second appearance relative to the first virtual object that is different from the first appearance.

[0015] A method is performed at a computer system that communicates with one or more display generation components and one or more input devices. The method includes: when a user interface of a first application and a close enable indication associated with the user interface of the first application are simultaneously displayed via the one or more display generation components, detecting a first input directed to the close enable indication. The method includes: in response to detecting the first input, based on determining that the first input is an input of a first type, displaying a first option for closing an application other than the first application.

[0016] A method is performed at a computer system in communication with a first display generation component and one or more input devices. The method includes: displaying a first object at a first location in a first view of a three-dimensional environment via the first display generation component. The method also includes: when displaying the first object at the first location in the first view of the three-dimensional environment via the first display generation component, displaying a first set of one or more control objects, wherein a corresponding control object in the first set of one or more control objects corresponds to a corresponding operation that can be applied to the first object. The method includes: detecting a first user input corresponding to a request to move the first object in the three-dimensional environment via the one or more input devices. The method also includes: in response to detecting the first user input corresponding to the request to move the first object in the three-dimensional environment: moving the first object from the first location to a second location, and when moving the first object from the first location to the second location, visually weakening at least one control object in the first set of one or more control objects corresponding to the corresponding operation that can be applied to the first object relative to the first object.

[0017] A method is performed at a computer system in communication with a display generation component and one or more input devices. The method includes: while displaying a first application user interface at a first location in a three-dimensional environment via the display generation component, detecting a first input corresponding to a request to close the first application user interface via the one or more input devices at a first time. The method includes: in response to detecting the first input corresponding to the request to close the first application user interface: closing the first application user interface, including stopping displaying the first application user interface in the three-dimensional environment; and displaying a main menu user interface at a corresponding main menu location determined based on the first location of the first application user interface in the three-dimensional environment in accordance with determining that a corresponding criterion is met.

[0018] A method is performed at a computer system in communication with a first display generation component and one or more input devices. The method includes: displaying a first user interface object via the display generation component, wherein the first user interface object includes first content. The method includes: while displaying the first user interface object including the first content via the display generation component, detecting a first user input directed to the first user interface object via the one or more input devices. The method also includes: in response to detecting the first user input pointing to the first user interface object, based on determining that the first user input corresponds to a request to adjust the size of the first user interface object, adjusting the size of the first user interface object based on the first user input, wherein adjusting the size of the first user interface object based on the first user input includes: one or more temporary resizing operations, including: based on determining that the characteristic refresh rate of the first content within the first user interface object is a first refresh rate, before updating the first content within the first user interface object according to a first updated size of the first user interface object specified by the first user input, scaling the first user interface object with the first content by a first scaling amount different from the first scaling amount before updating the first content within the first user interface object according to the first updated size of the first user interface object specified by the first user input; and based on determining that the characteristic refresh rate of the first content within the first user interface object is a second refresh rate different from the first refresh rate, scaling the first user interface object with the first content by a second scaling amount different from the first scaling amount. The method includes: after the one or more temporary resizing operations, displaying the first user interface object at the first updated size specified by the first user input, and updating the first content within the first user interface object according to the first updated size specified by the first user input.

[0019] A method is performed at a first computer system in communication with one or more display generation components and one or more input devices. The method includes: displaying a first application window at a first scale via the one or more display generation components. The method includes: while displaying the first application window at the first scale, detecting, via the one or more input devices, a first gesture directed to the first application window. The method includes: in response to detecting the first gesture, changing a corresponding scale of the first application window from the first scale to a second scale different from the first scale based on determining that the first gesture is directed to a corresponding portion of the first application window that is not associated with an application-specific response to the first gesture.

[0020] It should be noted that the various embodiments described above can be combined with any other embodiments described herein. The features and advantages described in this specification are not comprehensive, and in particular, many additional features and advantages will be apparent to those of ordinary skill in the art based on the drawings, the specification, and the claims. In addition, it should be noted that the language used in this specification is selected in principle for readability and instructional purposes, and may not be selected to describe or define the subject matter of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] For a better understanding of the various described embodiments, reference should be made to the following detailed description taken in conjunction with the following drawings, wherein like reference numerals designate corresponding parts throughout the several views.

[0022] Figure 1A is a block diagram illustrating an operating environment of a computer system for providing an extended reality (XR) experience according to some embodiments.

[0023] Figure 1B to Figure 1P is used in Figure 1A An example of a computer system that provides an XR experience in an operating environment.

[0024] Figure 2 is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's XR experience according to some embodiments.

[0025] Figure 3 is a block diagram illustrating display generation components of a computer system configured to provide a visual component of an XR experience to a user according to some embodiments.

[0026] Figure 4 is a block diagram illustrating a hand tracking unit of a computer system configured to capture gesture input from a user according to some embodiments.

[0027] Figure 5 is a block diagram illustrating an eye tracking unit of a computer system configured to capture gaze input of a user according to some embodiments.

[0028] Figure 6 is a flow chart illustrating a flash-assisted gaze tracking pipeline according to some embodiments.

[0029] FIG. 7A to FIG. 7R Example techniques for conditionally displaying controls for an application window and updating visual properties of the controls in response to user interactions according to some embodiments are illustrated.

[0030] Figure 7S to Figure 7AD3 Example techniques for displaying a title bar with information about an application according to some embodiments are illustrated.

[0031] Figure 7AE to Figure 7AH Example techniques for displaying privacy indicators according to some embodiments are illustrated.

[0032] Figure 7AI to Figure 7AP Example techniques for displaying some user interface elements having values ​​for a parameter constrained to a first range and displaying other user interface elements having values ​​for the parameter outside of the first range are illustrated according to some embodiments.

[0033] Figure 7AQ to Figure 7BB Example techniques are illustrated for displaying controls for a virtual object and automatically updating the displayed controls as the virtual object moves within an AR / VR environment in accordance with some embodiments.

[0034] Figure 7BC to Figure 7BL Example techniques for closing different applications or groups of applications in a three-dimensional environment in response to different inputs are illustrated in accordance with some embodiments.

[0035] Figure 7BM1 to Figure 7CH Example techniques for displaying a main menu user interface after closing one or more application user interfaces according to some embodiments are illustrated.

[0036] Figure 8 is a flowchart of a method of conditionally displaying controls for an application according to various embodiments.

[0037] Fig. 9 is a flow diagram of a method of updating visual properties of a control in response to user interaction according to various embodiments.

[0038] Fig.10 is a flowchart of a method of displaying a title bar with information about an application according to various embodiments.

[0039] Fig.11 is a flow chart of a method of displaying a privacy indicator according to various embodiments.

[0040] Fig.12 is a flow diagram of a method of displaying some user interface elements having values ​​of a parameter constrained to a first range and displaying other user interface elements having values ​​of the parameter outside of the first range, according to various embodiments.

[0041] Fig.13 is a flow chart of a method of displaying controls for a virtual object and automatically updating the displayed controls as the virtual object moves within an AR / VR environment according to various embodiments.

[0042] Fig.14is a flow chart of a method of closing different applications or groups of applications in a three-dimensional environment in response to different inputs according to various embodiments.

[0043] Fig.15 is a flow chart of a method of displaying controls for an object and automatically reducing the prominence of the controls as the object moves within an AR / VR environment, according to various embodiments.

[0044] Fig.16 is a flowchart of a method for displaying a main menu user interface after closing one or more application user interfaces according to various embodiments.

[0045] 17A1 to 17T Example techniques are illustrated for scaling a user interface object by different scaling factors based on a refresh rate of content in the user interface object during a resize operation of the user interface object in accordance with various embodiments.

[0046] Fig.18 is a flowchart of a method for scaling a user interface object by different scaling factors based on a refresh rate of content in the user interface object during a resize operation of the user interface object, according to various embodiments.

[0047] FIG. 19A to FIG. 19S Example techniques for scaling user interface objects based on corresponding gestures according to various embodiments are illustrated.

[0048] Fig. 20 is a flowchart of a method for scaling a user interface object based on a corresponding gesture according to various embodiments. DETAILED DESCRIPTION

[0049] According to some embodiments, the present disclosure relates to a user interface for providing an extended reality (XR) experience to a user.

[0050] The systems, methods, and GUIs described herein improve user interface interactions with virtual / augmented reality environments in a number of ways.

[0051] In some embodiments, a computer system displays an application window. In response to detecting that a user gaze is directed to a corresponding portion of the application window, the computer system displays a corresponding control associated with the corresponding portion of the application window. Conditionally displaying controls in response to detecting a user gaze directed to an area of ​​a control without requiring additional user input enables a user to access a specific control to perform an operation by shifting the user gaze without cluttering the user interface with the display of all available controls.

[0052] In some embodiments, the computer system displays a control for an application in a first appearance in response to a user looking at the control. The computer system updates the display of the control to be displayed in a second appearance after detecting that the user is interacting with the control (such as performing a gesture) to perform an operation. Automatically updating the appearance of the control when the user is looking at the control and further updating the appearance of the mobile control when the user is interacting with the control provides the user with improved visual feedback of the user interaction.

[0053] In some embodiments, the computer system displays an application window for a first application. The computer system displays a title bar for the application window simultaneously with the application window. In response to detecting that the user's attention is directed to the title bar, the size of the title bar is dynamically increased and additional controls are displayed in the title bar. The size of the title bar is dynamically increased to provide the user with access rights to the additional controls, which reduces the number of inputs required to access the additional controls for the application window and provides visual feedback on the device status.

[0054] In some embodiments, the computer system displays an application window in the display area. The computer system determines whether the application is accessing sensitive user data and displays a privacy indicator near the application window of the application that is accessing sensitive user data. In some embodiments, the computer system detects user input that moves the application window in the display area. The privacy indicator for the application window continues to be provided even when the application window is repositioned to be displayed at a different location in the display area, improving the security and privacy of the system by providing real-time information about a specific application window that is accessing sensitive user data and maintaining the information as the application window moves within the display area.

[0055] In some embodiments, a computer system displays a user interface, wherein displaying the user interface includes simultaneously displaying a content area having first content, a first user interface object, and a second user interface object in the user interface. The corresponding content in the content area is constrained to have an appearance in which the corresponding parameter is within a first value range, the first user interface object is displayed with an appearance in which the corresponding parameter has a value outside the first value range, and the second user interface object is displayed with an appearance in which the corresponding parameter has a value outside the first value range. When the first content, the first user interface object, and the second user interface object are displayed simultaneously, the user interface is updated. Updating the user interface includes: changing the first content to the second content, while the corresponding content in the content area continues to be constrained to have an appearance in which the corresponding parameter is within the first value range; updating the appearance of the first user interface object and continuing to display the first user interface object with an appearance in which the corresponding parameter has a value outside the first value range; and updating the appearance of the second user interface object and continuing to display the second user interface object with an appearance in which the corresponding parameter has a value outside the first value range.

[0056] In some embodiments, in response to detecting a first event corresponding to a request to display a first virtual object in a three-dimensional environment, a computer system displays the first virtual object at a first location in the three-dimensional environment together with a first object management user interface corresponding to the first virtual object, and wherein the first object management user interface has a first appearance relative to the first virtual object at the first location. In response to detecting a first user input corresponding to a request to move the first virtual object in the three-dimensional environment, the computer system displays the first virtual object at a second location in a first view of the three-dimensional environment, while displaying the first object management user interface at the second location in the three-dimensional environment, the first object management user interface having a second appearance relative to the first virtual object. In response to detecting that the virtual object is moving within the AR / VR environment, a control for the virtual object is automatically updated without requiring additional user input, so that the user can continue to access the control to perform an operation, and improved visual feedback is provided by dynamically adjusting the control to make it easy for the user to view even when the location of the virtual object changes.

[0057] In some embodiments, the computer system simultaneously displays a user interface of a first application and a close enable indication associated with the user interface of the first application. In response to detecting a first input directed to the close enable indication, the computer system displays a first option for closing applications other than the first application. Providing different options for closing one or more user interfaces reduces the number of inputs required to display one or more user interfaces of interest.

[0058] In some embodiments, the computer system displays a first set of one or more control objects when displaying a first object at a first position in a first view of a three-dimensional environment, wherein a corresponding control object in the first set of one or more control objects corresponds to a corresponding operation that can be applied to the first object. In response to detecting a first user input corresponding to the request to move the first object in the three-dimensional environment, the computer system: moves the first object from the first position to the second position, and when moving the first object from the first position to the second position, visually weakens at least one control object in the first set of one or more control objects corresponding to the corresponding operation that can be applied to the first object relative to the first object. In response to detecting that a virtual object is moving within an AR / VR environment, controls for the object are automatically updated without additional user input, so that the user can continue to access the controls to perform operations, and improved visual feedback is provided by indicating that the controls are not available for interaction when the object is being moved.

[0059] In some embodiments, the computer system displays a first application user interface at a first location in the three-dimensional environment. In response to detecting a first input corresponding to a request to close the first application user interface, the computer system closes the first application user interface and displays the main menu user interface at a corresponding main menu location determined based on the first location of the first application user interface in the three-dimensional environment. When no application user interfaces remain in the viewport of the three-dimensional environment, automatically displaying the main menu user interface allows the user to continue navigating through one or more sets of selectable representations of the main menu user interface without displaying additional controls.

[0060] In some embodiments, in response to detecting a user input corresponding to a request to resize a first user interface object, a computer system resizes the first user interface object in accordance with the user input, including performing one or more temporary resizing operations, including: scaling the first user interface object having the first content by a first amount in accordance with a determination that a characteristic refresh rate of the first content is a first refresh rate; and scaling the first user interface object having the first content by a second amount in accordance with a determination that a characteristic refresh rate of the first content is a second refresh rate. After the one or more temporary resizing operations, the computer system displays the first user interface object at a first updated size specified by the first user input, and updates the first content within the first user interface object.

[0061] In some embodiments, the computer system displays the first application window at a first scale. While displaying the first application window at the first scale, the computer system detects a first gesture directed to the first application window. In response to detecting the first gesture, based on determining that the first gesture is directed to a corresponding portion of the first application window that is not associated with an application-specific response to the first gesture, the computer system changes a corresponding scale of the first application window from the first scale to a second scale that is different from the first scale.

[0062] Figures 1A to 6 A description of an example computer system for providing an XR experience to a user is provided. FIG. 7A to FIG. 7R Example techniques for conditionally displaying controls for an application window and updating visual properties of the controls in response to user interactions according to some embodiments are illustrated. Figure 8 is a flowchart of a method of conditionally displaying controls for an application window according to various embodiments. Fig. 9 is a flow diagram of a method of updating visual properties of a control in response to user interaction according to various embodiments. FIG. 7A to FIG. 7R The user interface in Figure 8 and Fig. 9 process. Figure 7S to Figure 7AD3 Example techniques for displaying a title bar with information about an application according to some embodiments are illustrated. Fig.10 is a flowchart of a method of displaying a title bar with information about an application according to various embodiments. Figure 7S to Figure 7AD3 The user interface in Fig.10 FIG7T (for example, Figure 7T1 , Figure 7T2 and Figure 7T3 )to Figure 7AH Example techniques for displaying privacy indicators according to some embodiments are illustrated. Fig.11 is a flow chart of a method of displaying a privacy indicator according to various embodiments. Figure 7AH The user interface in Fig.11 process. Figure 7AI to Figure 7AP Example techniques for displaying some user interface elements having values ​​for a parameter constrained to a first range and displaying other user interface elements having values ​​for the parameter outside of the first range are illustrated according to some embodiments. Fig.12 is a flowchart of a method for displaying some user interface elements having values ​​of a parameter constrained to a first range and displaying other user interface elements having values ​​of the parameter outside of the first range, according to various embodiments. Figure 7AI to Figure 7AP The user interface in Fig.12 process. Figure 7AQ to Figure 7BBExample techniques are illustrated for displaying controls for a virtual object and automatically updating the displayed controls as the virtual object moves within an AR / VR environment, and for displaying controls for an object and automatically reducing the prominence of the controls as the object moves within an AR / VR environment, according to some embodiments. Fig.13 is a flowchart of a method of displaying controls for a virtual object and automatically updating the displayed controls as the virtual object moves within an AR / VR environment according to various embodiments. Fig.15 is a flow chart of a method of displaying controls for an object and automatically reducing the prominence of the controls as the object moves within an AR / VR environment, according to various embodiments. Figure 7AQ to Figure 7BB The user interface in Fig.13 and Fig.15 process. Figure 7BC to Figure 7BL Example techniques are illustrated for closing different applications or groups of applications in a three-dimensional environment in response to different inputs. Fig.14 is a flow chart of a method of closing different applications or groups of applications in a three-dimensional environment in response to different inputs according to various embodiments. Figure 7BC to Figure 7BL The user interface in Fig.14 process. Figure 7B M to Figure 7CH Example techniques for displaying a main menu user interface after closing one or more application user interfaces according to some embodiments are illustrated. Fig.16 is a flowchart of a method for displaying a main menu user interface after closing one or more application user interfaces according to various embodiments. Figure 7B M to Figure 7CH The user interface in Fig.16 process. 17A1 to 17T Example techniques are illustrated for scaling a user interface object by different scaling factors based on a refresh rate of content in the user interface object during a resize operation of the user interface object in accordance with various embodiments. Fig.18 is a flowchart of a method for scaling a user interface object by different scaling factors based on a refresh rate of content in the user interface object during a resize operation of the user interface object, according to various embodiments. 17A1 to 17T The user interface in Fig.18 process. FIG. 19A to FIG. 19S Example techniques for scaling user interface objects based on corresponding gestures according to various embodiments are illustrated. Fig. 20 is a flow chart for scaling a user interface object based on a corresponding gesture according to various embodiments. FIG. 19A to FIG. 19S The user interface in Fig. 20 process.

[0063] The process described below enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device) through various techniques, including by providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional display controls, performing an operation without further user input when a set of conditions have been met, improving privacy and / or security, providing a richer, more detailed and / or more realistic user experience while saving storage space, and / or additional techniques. These techniques also reduce power usage and extend the battery life of the device by enabling the user to use the device faster and more efficiently. Saving battery power, and therefore saving weight, improves the ergonomics of the device. These techniques also enable real-time communication, allow the use of fewer and / or less accurate sensors, thereby producing a more compact, lighter and cheaper device, and enable the device to be used under various lighting conditions. These techniques reduce energy usage and thereby reduce the heat emitted by the device, which is particularly important for wearable devices where if the device generates too much heat well within the operating parameters of the device components, it may become uncomfortable for the user to wear the device.

[0064] In addition, in the method described herein where one or more steps depend on one or more conditions being met, it should be understood that the method can be repeated in multiple repetitions so that in the process of repetition, all conditions for determining the steps in the method have been met in different repetitions of the method. For example, if the method needs to perform the first step (if the condition is met), and perform the second step (if the condition is not met), then the ordinary technician will know that the steps stated are repeated until both the condition is met and the condition is not met (in no particular order). Therefore, the method described as having one or more steps depending on one or more conditions being met can be rewritten as a method of repeating until each condition described in the method is met. However, this does not require the system or computer-readable medium to declare that the system or computer-readable medium contains instructions for performing a contingent operation based on the satisfaction of the corresponding one or more conditions, and is therefore able to determine whether the possible situation has been met without explicitly repeating the steps of the method until all conditions for determining the steps in the method have been met. It will also be understood by ordinary technicians in the art that, similar to the method with the contingent step, the system or computer-readable storage medium can repeat the steps of the method as many times as needed to ensure that all the contingent steps have been performed.

[0065] In some embodiments, such as Figure 1AAs shown, an XR experience is provided to a user via an operating environment 100 including a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a position sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., a household appliance, a wearable device, etc.). In some embodiments, one or more of the input device 125, the output device 155, the sensor 190, and the peripheral device 195 are integrated with the display generation component 120 (e.g., in a head mounted device or a handheld device).

[0066] In describing an XR experience, various terms are used to distinguishably refer to several related but distinct environments that a user can sense and / or with which the user can interact (e.g., using inputs detected by the computer system 101 generating the XR experience, which inputs cause the computer system generating the XR experience to generate audio, visual, and / or tactile feedback corresponding to the various inputs provided to the computer system 101). The following is a subset of these terms:

[0067] Physical environment: The physical environment refers to the physical world that people can sense and / or interact with without the help of electronic systems. A physical environment, such as a physical park, includes physical objects, such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment, such as through vision, touch, hearing, taste, and smell.

[0068] Extended Reality: In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via an electronic system. In XR, a subset of a person's physical movements or a representation thereof is tracked, and in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner that complies with at least one physical law. For example, an XR system may detect a person's head turn, and in response, adjust the graphical content and sound field presented to the person in a manner similar to the way such views and sounds change in a physical environment. In some cases (e.g., for accessibility reasons), adjustments to the characteristics of virtual objects in the XR environment may be made in response to a representation of physical movement (e.g., a voice command). People can sense and / or interact with XR objects using any of their senses, including vision, hearing, touch, taste, and smell. For example, people can sense and / or interact with audio objects, which create a 3D or spatial audio environment that provides the perception of a point audio source in 3D space. As another example, an audio object can enable audio transparency that selectively introduces ambient sounds from the physical environment with or without computer-generated audio. In some XR environments, a person can sense and / or interact only with an audio object.

[0069] Examples of XR include virtual reality and mixed reality.

[0070] Virtual Reality: A virtual reality (VR) environment refers to a simulated environment designed to be based entirely on computer-generated sensory input to one or more senses. A VR environment includes a plurality of virtual objects that a person can sense and / or interact with. For example, trees, buildings, and computer-generated images representing avatars of people are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment through a simulation of the person's presence within the computer-generated environment and / or through a simulation of a subset of the person's physical movement within the computer-generated environment.

[0071] Mixed Reality: In contrast to VR environments, which are designed to be based entirely on computer-generated sensory input, a mixed reality (MR) environment refers to a simulated environment that is designed to include sensory input from the physical environment, or representations thereof, in addition to computer-generated sensory input (e.g., virtual objects). On the virtuality continuum, a mixed reality environment is anything between, but not including, a fully physical environment at one end and a virtual reality environment at the other end. In some MR environments, computer-generated sensory input can respond to changes in sensory input from the physical environment. In addition, some electronic systems used to render MR environments can track position and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical items from the physical environment, or representations thereof). For example, the system can cause motion so that virtual trees appear stationary relative to the physical ground.

[0072] Examples of mixed reality include augmented reality and augmented virtuality.

[0073] Augmented reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on a physical environment or a representation of a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on a transparent or translucent display so that a person uses the system to perceive virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system combines the image or video with the virtual object and presents the composition on an opaque display. A person uses the system to indirectly view the physical environment via an image or video of the physical environment and perceives virtual objects superimposed on the physical environment. As used herein, a video of a physical environment displayed on an opaque display is referred to as a "transparent video," meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on an opaque display. Further alternatively, the system may have a projection system that projects virtual objects into a physical environment, such as as a hologram or on a physical surface, so that a person using the system perceives virtual objects superimposed on the physical environment. An augmented reality environment also refers to a simulated environment in which a representation of a physical environment is transformed by computer-generated sensory information. For example, in providing a pass-through video, the system may transform one or more sensor images to apply a selected perspective (e.g., a viewpoint) that is different from the perspective captured by the imaging sensor. For another example, the representation of the physical environment may be transformed by graphically modifying (e.g., enlarging) a portion thereof so that the modified portion may be a representative but not real version of the original captured image. For another example, the representation of the physical environment may be transformed by graphically eliminating a portion thereof or blurring a portion thereof.

[0074] Augmented Virtual: An augmented virtual (AV) environment refers to a simulated environment in which a virtual environment or a computer-generated environment incorporates one or more sensory inputs from the physical environment. The sensory input may be a representation of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, but the faces of people are realistically reproduced from images taken of physical people. For another example, a virtual object may take the shape or color of a physical object imaged by one or more imaging sensors. For another example, a virtual object may take a shadow that conforms to the positioning of the sun in the physical environment.

[0075] In an augmented reality, mixed reality or virtual reality environment, a view of a three-dimensional environment is visible to a user. The view of a three-dimensional environment is usually visible to a user through a virtual viewport via one or more display generation components (e.g., a display or a pair of display modules that provide stereoscopic content to different eyes of the same user), and the virtual viewport has a viewport boundary, which defines the scope of the three-dimensional environment visible to the user via one or more display generation components. In some embodiments, the area defined by the viewport boundary is smaller than the user's visual range in one or more dimensions (e.g., based on the user's visual range, the size of one or more display generation components, optical properties or other physical properties, and / or the position and / or orientation of one or more display generation components relative to the user's eyes). In some embodiments, the area defined by the viewport boundary is larger than the user's visual range in one or more dimensions (e.g., based on the user's visual range, the size of one or more display generation components, optical properties or other physical properties, and / or the position and / or orientation of one or more display generation components relative to the user's eyes). The viewport and viewport boundary usually move with the movement of one or more display generation components (e.g., for a head-mounted device, it moves with the user's head, or for a handheld device such as a tablet or smart phone, it moves with the user's hand). The user's viewpoint determines what is visible in the viewport, the viewpoint typically specifies a position and orientation relative to the three-dimensional environment, and as the viewpoint moves, the view of the three-dimensional environment will also move in the viewport. For head-mounted devices, the viewpoint is typically based on the position and orientation of the user's head, face, and / or eyes to provide a view of the three-dimensional environment that is perceptually accurate and provides an immersive experience when the user is using the head-mounted device. For handheld or fixed devices, the viewpoint moves as the handheld or fixed device moves and / or as the user's positioning relative to the handheld or fixed device changes (e.g., the user moves toward, away from, up, down, right, and / or left). For a device including display generation components with virtual pass-through, the portion of the physical environment visible (e.g., displayed and / or projected) via one or more display generation components is based on the field of view of one or more cameras in communication with the display generation components, which one or more cameras typically move with movement of the display generation components (e.g., with movement of the user's head for a head-mounted device, or with movement of the user's hands for a handheld device such as a tablet or smartphone) because the user's viewpoint moves with movement of the field of view of the one or more cameras (and the appearance of one or more virtual objects displayed via the one or more display generation components is updated based on the user's viewpoint (e.g., the display position and pose of the virtual objects are updated based on movement of the user's viewpoint)).For display generating components with optical transmittance, portions of the physical environment that are visible via one or more display generating components (e.g., optically visible through one or more partially or fully transparent portions of the display generating components) are based on the user's field of view through the partially or fully transparent portions of the display generating components (e.g., moves as the user's head moves for a head-mounted device, or moves as the user's hands move for a handheld device such as a tablet or smart phone) because the user's viewpoint moves as the user moves through the field of view of the partially or fully transparent portions of the display generating components (and the appearance of one or more virtual objects is updated based on the user's viewpoint).

[0076] In some embodiments, the representation of the physical environment (e.g., displayed via virtual transmission or optical transmission) may be partially or completely obscured by the virtual environment. In some embodiments, the amount of the virtual environment displayed (e.g., the amount of the physical environment that is not displayed) is based on the immersion level of the virtual environment (e.g., relative to the representation of the physical environment). For example, increasing the immersion level optionally causes more of the virtual environment to be displayed, replacing and / or obscuring more of the physical environment, and reducing the immersion level optionally causes less of the virtual environment to be displayed, thereby revealing portions of the physical environment that were previously not displayed and / or obscured. In some embodiments, at a particular immersion level, one or more first background objects (e.g., in the representation of the physical environment) are more visually de-emphasized (e.g., dimmed, blurred, displayed with increased transparency) than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, the immersion level includes an associated degree to which virtual content (e.g., a virtual environment and / or virtual content) displayed by a computer system obscures background content (e.g., content other than the virtual environment and / or virtual content) surrounding / behind the virtual environment, optionally including the number of items of background content displayed and / or displayed visual characteristics of the background content (e.g., color, contrast, and / or opacity), the angular range of the virtual content displayed via the display generation component (e.g., 60 degrees for content displayed at low immersion, 120 degrees for content displayed at medium immersion, or 180 degrees for content displayed at high immersion), and / or the proportion of the field of view displayed via the display generation component that is occupied by the virtual content (e.g., 33% of the field of view occupied by the virtual content at low immersion, 66% of the field of view occupied by the virtual content at medium immersion, or 100% of the field of view occupied by the virtual content at high immersion). In some embodiments, the background content is included in the background on which the virtual content is displayed (e.g., background content in a representation of the physical environment). In some embodiments, the background content includes a user interface (e.g., a user interface generated by a computer system corresponding to an application), virtual objects that are not associated with or included in the virtual environment and / or virtual content (e.g., files generated by a computer system or representations of other users, etc.), and / or real objects (e.g., transparent objects representing real objects in the physical environment surrounding the user, which are visible so that they are displayed via the display generation component and / or visible via transparent or translucent components of the display generation component because the computer system does not block / impede their visibility through the display generation component). In some embodiments, at a low immersion level (e.g., a first immersion level), the background, virtual and / or real objects are displayed in an unobstructed manner. For example, a virtual environment with a low immersion level is optionally displayed simultaneously with the background content, which is optionally displayed at full brightness, color and / or translucency.In some embodiments, at a higher immersion level (e.g., a second immersion level higher than the first immersion level), background, virtual and / or real objects are displayed in an obstructed manner (e.g., dimmed, blurred, or removed from the display). For example, a corresponding virtual environment with a high immersion level is displayed without simultaneously displaying background content (e.g., in full screen or full immersion mode). As another example, a virtual environment displayed at a medium immersion level is displayed simultaneously with background content that is dimmed, blurred, or otherwise de-emphasized. In some embodiments, the visual characteristics of background objects vary between background objects. For example, at a particular immersion level, one or more first background objects are more visually de-emphasized (e.g., dimmed, blurred, and / or displayed with increased transparency) than one or more second background objects, and one or more third background objects stop displaying. In some embodiments, zero immersion or zero immersion level corresponds to a virtual environment that stops displaying, and instead displays a representation of the physical environment (optionally with one or more virtual objects, such as applications, windows, or virtual three-dimensional objects), while the representation of the physical environment is not obscured by the virtual environment. Adjusting the immersion level using physical input elements provides a fast and efficient method of adjusting the degree of immersion, which enhances the operability of the computer system and makes the user-device interface more efficient.

[0077] Viewpoint-locked virtual objects: When a computer system displays a virtual object at the same position and / or location in a user's viewpoint, the virtual object is viewpoint-locked even if the user's viewpoint shifts (e.g., changes). In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the forward direction of the user's head (e.g., when the user is looking straight ahead, the user's viewpoint is at least a portion of the user's field of view); thus, without moving the user's head, the user's viewpoint remains fixed even when the user's gaze shifts. In embodiments where the computer system has a display generation component (e.g., a display screen) that is repositionable relative to the user's head, the user's viewpoint is an augmented reality view presented to the user on the display generation component of the computer system. For example, a viewpoint-locked virtual object that is displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's viewpoint even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the position and / or location of the viewpoint-locked virtual object displayed in the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, such that the virtual object is also referred to as a "head-locked virtual object."

[0078] Environment-locked visual objects: A virtual object is environment-locked (alternatively, "world-locked") when a computer system displays a virtual object at a position and / or location in a user's viewpoint that is based on (e.g., selected with reference to and / or anchored to) a location and / or object in a three-dimensional environment (e.g., a physical environment or a virtual environment). As the user's viewpoint moves, the location and / or objects in the environment change relative to the user's viewpoint, which causes the environment-locked virtual object to be displayed at a different location and / or location in the user's viewpoint. For example, an environment-locked virtual object locked to a tree immediately in front of the user is displayed at the center of the user's viewpoint. When the user's viewpoint shifts to the right (e.g., the user's head turns to the right) such that the tree is now to the left of center in the user's viewpoint (e.g., the tree location in the user's viewpoint shifts), the environment-locked virtual object locked to the tree is displayed to the left of center in the user's viewpoint. In other words, the position and / or orientation of the virtual object that is locked to the environment in the user's viewpoint depends on the position and / or orientation of the object in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference system (e.g., a coordinate system anchored to a fixed position and / or object in the physical environment) to determine the position of the virtual object that is locked to the user's viewpoint. The virtual object that is locked to the environment can be locked to a stationary part of the environment (e.g., a floor, wall, table or other stationary object), or can be locked to a movable part of the environment (e.g., a vehicle, animal, person or even a representation of a part of the user's body that moves independently of the user's viewpoint, such as a hand, wrist, arm or foot of the user) so that the virtual object moves as the viewpoint or the part of the environment moves to maintain a fixed relationship between the virtual object and the part of the environment.

[0079] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits an inertial following behavior that reduces or delays the movement of the environment-locked or viewpoint-locked virtual object relative to the movement of the reference point followed by the virtual object. In some embodiments, when exhibiting inertial following behavior, when detecting the movement of a reference point (e.g., a portion of the environment, a viewpoint, or a point fixed relative to the viewpoint, such as a point between 5cm and 300cm from the viewpoint) that the virtual object is following, the computer system intentionally delays the movement of the virtual object. For example, when the reference point (e.g., the portion of the environment or the viewpoint) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point, but moves at a second speed that is slower than the first speed (e.g., until the reference point stops moving or slows down, at which point the virtual object begins to catch up with the reference point). In some embodiments, when the virtual object exhibits inertial following behavior, the device ignores small amounts of movement of the reference point (e.g., ignoring movements of the reference point below a threshold movement amount, such as moving 0 degrees to 5 degrees or moving 0cm to 50cm). For example, when a reference point (e.g., a portion or viewpoint of an environment to which a virtual object is locked) moves a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed so as to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and when the reference point (e.g., the portion or viewpoint of the environment to which the virtual object is locked) moves a second amount greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is being displayed so as to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and then decreases when the amount of movement of the reference point increases above a threshold (e.g., a “lazy follow” threshold) because the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point. In some embodiments, maintaining a substantially fixed position of the virtual object relative to the reference point includes displaying the virtual object within a threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 15 cm, 20 cm, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward relative to the position of the reference point).

[0080] Hardware: There are many different types of electronic systems that enable a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without tactile feedback), smartphones, tablet devices, and desktop / laptop computers. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. Instead of an opaque display, a head-mounted system may have a transparent or translucent display. A transparent or translucent display may have a medium through which light representing an image is directed to a person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium may be an optical waveguide, a hologram medium, an optical combiner, an optical reflector, or any combination thereof. In one embodiment, a transparent or translucent display may be configured to selectively become opaque. Projection-based systems may employ retinal projection technology that projects graphic images onto a person's retina. The projection system may also be configured to project virtual objects into a physical environment, such as as a hologram or on a physical surface. In some embodiments, the controller 110 is configured to manage and coordinate the user's XR experience. In some embodiments, the controller 110 includes a suitable combination of software, firmware, and / or hardware. The following description with respect to Figure 2Controller 110 is described in more detail. In some embodiments, controller 110 is a computing device that is located locally or remotely relative to scene 105 (e.g., physical environment). For example, controller 110 is a local server located within scene 105. As another example, controller 110 is a remote server (e.g., cloud server, central server, etc.) located outside scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., HMD, display, projector, touch screen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE802.16x, IEEE 802.3x, etc.). For example, the controller 110 is included in a housing (e.g., a physical housing) of the display generating component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors, etc.), one or more input devices among the input devices 125, one or more output devices among the output devices 155, one or more sensors among the sensors 190, and / or one or more peripheral devices among the peripheral devices 195, or shares the same physical housing or support structure with one or more of the above devices.

[0081] In some embodiments, the display generation component 120 is configured to provide an XR experience (e.g., at least the visual component of the XR experience) to the user. In some embodiments, the display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Figure 3 Display generation component 120 is described in further detail. In some embodiments, the functionality of controller 110 is provided by and / or combined with display generation component 120.

[0082] According to some embodiments, display generation component 120 provides an XR experience to the user when the user is virtually and / or physically present within scene 105.

[0083] In some embodiments, the display generation component is worn on a part of the user's body (e.g., on his / her head, on his / her hand, etc.). In this way, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smart phone or tablet device) configured to present XR content, and the user holds a device with a display facing the user's field of view and a camera facing the scene 105. In some embodiments, the handheld device is optionally placed in a shell worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is an XR room, shell, or room configured to present XR content, wherein the user does not wear or hold the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interactions with XR content triggered based on interactions occurring in the space in front of a handheld device or a tripod-mounted device may similarly be implemented with an HMD, where the interactions occur in the space in front of the HMD and responses to the XR content are displayed via the HMD. Similarly, a user interface showing interactions with XR content triggered based on movement of a handheld device or a tripod-mounted device relative to a physical environment (e.g., scene 105 or a part of a user's body (e.g., the user's eyes, head, or hands)) may similarly be implemented with an HMD, where the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of a user's body (e.g., the user's eyes, head, or hands)).

[0084] Despite Figure 1A Relevant features of operating environment 100 are shown, but those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the example embodiments disclosed herein.

[0085] Figure 1A to Figure 1PVarious examples of computer systems for performing methods and providing audio, visual, and / or tactile feedback as part of the user interface described herein are illustrated. In some embodiments, the computer system includes one or more display generation components (e.g., a first display component 1-120a and a second display component 1-120b and / or a first optical module 11.1.1-104a and a second optical module 11.1.1-104b) for displaying to a user of the computer system a representation of a virtual element and / or a physical environment, optionally generated based on detected events and / or user input detected by the computer system. The user interface generated by the computer system is optionally corrected by one or more corrective lenses 11.3.2-216, which are optionally removably attached to one or more of the optical modules to make the user interface easier to view by users who would otherwise use glasses or contact lenses to correct their vision. Although many of the user interfaces shown herein show a single view of the user interface, the user interface in the HMD is optionally displayed using two optical modules (e.g., a first display component 1-120a and a second display component 1-120b and / or a first optical module 11.1.1-104a and a second optical module 11.1.1-104b), one optical module for the user's right eye and a different optical module for the user's left eye, and slightly different images are presented to the two different eyes to generate the illusion of stereoscopic depth, the single view of the user interface is typically a right eye view or a left eye view, and the depth effect is explained in the text or using other schematics or views. In some embodiments, the computer system includes one or more external displays (e.g., display components 1-108) for displaying status information of the computer system to a user of the computer system (when the computer system is not worn) and / or to other people near the computer system, the status information being optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more audio output components (e.g., electronic components 1-112) for generating audio feedback, which is optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting information about the physical environment of the device (e.g., one or more sensors in sensor components 1-356, and / or Fig. 1I ), which information can be used (optionally in conjunction with one or more luminaires, such as Fig. 1IThe computer system may generate a digital pass-through image using an illuminator as described in the above, capture visual media (e.g., photos and / or videos) corresponding to the physical environment, or determine the position (e.g., location and / or orientation) of physical objects and / or surfaces in the physical environment, so that virtual objects can be placed based on the detected position of the physical objects and / or surfaces. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting hand positioning and / or movement (e.g., sensor assembly 1-356 and / or Fig. 1I One or more sensors in the apparatus) which may be used (optionally in combination with one or more illuminators, such as Fig. 1I In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting eye movement (e.g., Fig. 1I eye tracking and gaze tracking sensors in the Fig.1OLights 11.3.2-110 in the device) determine attention or gaze location and / or gaze movement, which can be optionally used to detect gaze-only input based on gaze movement and / or dwell. Combinations of the various sensors described above can be used to determine user facial expressions and / or hand movements for generating an avatar or representation of the user, such as an anthropomorphic avatar or representation for a real-time communication session, where the avatar has facial expressions, hand movements, and / or body movements based on or similar to the detected facial expressions, hand movements, and / or body movements of the user of the device. Gaze and / or attention information is optionally combined with hand tracking information to determine interaction between a user and one or more user interfaces based on direct and / or indirect input, such as air gestures or input using one or more hardware input devices, such as one or more buttons (e.g., first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328), knobs (e.g., first button 1-128, button 11.1.1-114, and / or dial or button 1-328), digital crown (e.g., pressable and twistable or rotatable first button 1-128, button 11.1.1-114, and / or dial or button 1-328), touchpad, touch screen, keyboard, mouse, and / or other input devices. One or more buttons (e.g., first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328) are optionally used to perform system operations, such as re-centering content in a three-dimensional environment visible to a user of the device, displaying a primary user interface for launching an application, starting a real-time communication session, or initiating display of a virtual three-dimensional background. A knob or digital crown (e.g., a pressable and twistable or rotatable first button 1-128, button 11.1.1-114, and / or dial or button 1-328) is optionally rotatable to adjust parameters of visual content, such as an immersion level of a virtual three-dimensional environment (e.g., the extent to which virtual content occupies a user's viewport in the three-dimensional environment) or other parameters associated with the three-dimensional environment and virtual content displayed via optical modules (e.g., first display component 1-120a and second display component 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b).

[0086] Figure 1BFront, top, and perspective views of an example of a head mounted display (HMD) device 1-100 configured to be worn by a user and to provide a virtual and altered / mixed reality (VR / AR) experience are illustrated. The HMD 1-100 may include a display unit 1-102 or assembly, an electronic strap assembly 1-104 connected to and extending from the display unit 1-102, and a strap assembly 1-106 secured to the electronic strap assembly 1-104 at either end. The electronic strap assembly 1-104 and the strap 1-106 may be part of a retaining assembly configured to wrap around a user's head to hold the display unit 1-102 against the user's face.

[0087] In at least one example, the strap assembly 1-106 may include a first strap 1-116 configured to wrap around the back of a user's head and a second strap 1-117 configured to extend over the top of the user's head. As shown, the second strap may extend between the first electronic strip 1-105a and the second electronic strip 1-105b of the electronic strip assembly 1-104. The strap assembly 1-104 and the strap assembly 1-106 may be part of a securing mechanism that extends rearward from the display unit 1-102 and is configured to hold the display unit 1-102 against the user's face.

[0088] In at least one example, the securing mechanism includes a first electronic strip 1-105a including a first proximal end 1-134 coupled to the display unit 1-102 (e.g., a housing 1-150 of the display unit 1-102) and a first distal end 1-136 opposite the first proximal end 1-134. The securing mechanism may also include a second electronic strip 1-105b including a second proximal end 1-138 coupled to the housing 1-150 of the display unit 1-102 and a second distal end 1-140 opposite the second proximal end 1-138. The securing mechanism may also include a first band 1-116 and a second band 1-117, the first band including a first end 1-142 coupled to the first distal end 1-136 and a second end 1-144 coupled to the second distal end 1-140, and the second band extending between the first electronic strip 1-105a and the second electronic strip 1-105b. The strips 1-105a-b and the strip 1-116 may be coupled via a connection mechanism or assembly 1-114. In at least one example, the second strip 1-117 includes a first end 1-146 coupled to the first electronic strip 1-105a between the first proximal end 1-134 and the first distal end 1-136 and a second end 1-148 coupled to the second electronic strip 1-105b between the second proximal end 1-138 and the second distal end 1-140.

[0089] In at least one example, the first and second electronic strips 1-105a-b include plastic, metal, or other structural materials formed into substantially rigid strip 1-105a-b shapes. In at least one example, the first band 1-116 and the second band 1-117 are formed of a resilient flexible material including a woven textile, rubber, etc. The first band 1-116 and the second band 1-117 can be flexible to conform to the shape of the user's head when the HMD 1-100 is worn.

[0090] In at least one example, one or more of the first and second electronic strips 1-105a-b can define an inner strip volume and include one or more electronic components disposed in the inner strip volume. Figure 1B As shown, the first electronic strip 1-105a may include an electronic component 1-112. In one example, the electronic component 1-112 may include a speaker. In one example, the electronic component 1-112 may include a computing component, such as a processor.

[0091] In at least one example, the housing 1-150 defines a first front opening 1-152. Figure 1B 1-152 in dashed lines because the display assembly 1-108 is configured to block the first opening 1-152 from the field of view when the HMD 1-100 is assembled. The housing 1-150 may also define a rear-mounted second opening 1-154. The housing 1-150 also defines an internal volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display assembly 1-108, which may include a front cover and a display screen (shown in other figures) disposed in or across the front opening to block the front opening 1-152. In at least one example, the display screen of the display assembly 1-108, and the display assembly 1-108 in general, has a curvature configured to follow the curvature of the user's face. The display screen of the display assembly 1-108 may be curved as shown to complement the user's facial features and the overall curvature from one side of the face to the other, such as from left to right and / or from top to bottom, where the display unit 1-102 is pressed.

[0092] In at least one example, the housing 1-150 may define a first hole 1-126 between the first opening 1-152 and the second opening 1-154 and a second hole 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 may also include a first button 1-126 disposed in the first hole 1-128, and a second button 1-132 disposed in the second hole 1-130. The first button 1-128 and the second button 1-132 can be pressed through the corresponding holes 1-126, 1-130. In at least one example, the first button 1-126 and / or the second button 1-132 can be a twistable dial and a pressable button. In at least one example, the first button 1-128 is a pressable and twistable dial button, and the second button 1-132 is a pressable button.

[0093] Figure 1C A rear perspective view of an HMD 1-100 is illustrated. The HMD 1-100 may include a light seal 1-110 extending rearwardly from a housing 1-150 of a display assembly 1-108 around the periphery of the housing 1-150, as shown. The light seal 1-110 may be configured to extend from the housing 1-150 to the face of the user, around the eyes of the user, to block external light from being visible. In one example, the HMD 1-100 may include a first display assembly 1-120a and a second display assembly 1-120b, which are disposed at or in a rearward-facing second opening 1-154 defined by the housing 1-150 and / or disposed in an internal volume of the housing 1-150 and are configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a-b may include a corresponding display screen 1-122a, 1-122b, which is configured to project light toward the user's eyes through the second opening 1-154 in a rearward direction.

[0094] In at least one example, reference Figure 1B and Figure 1C In both cases, the display assembly 1-108 may be a front-facing forward display assembly including a display screen configured to project light in a first forward direction, and the rear display screens 1-122a-b may be configured to project light in a second rearward direction opposite the first direction. As noted above, the light seal 1-110 may be configured to block light external to the HMD 1-100 from reaching the user's eyes, including by Figure 1B 1-108 is shown in the front perspective view of the HMD 1-100. In at least one example, the HMD 1-100 may also include a curtain 1-124 that blocks the second opening 1-154 between the housing 1-150 and the rear display assembly 1-120a-b. In at least one example, the curtain 1-124 may be elastic or at least partially elastic.

[0095] Figure 1B and Figure 1C Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figures 1D to 1F Any other examples of devices, features, components and parts shown and described herein. Figures 1D to 1F Any of the features, components and / or parts shown or described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1B and Figure 1C Examples of devices, features, components, and parts are shown.

[0096] Figure 1D An exploded view of an example of an HMD 1-200 is illustrated that includes various portions or parts that are separated according to modularization and selective coupling of these parts. For example, the HMD 1-200 may include a band 1-216 that is selectively coupleable to a first electronic strip 1-205a and a second electronic strip 1-205b. The first fixed strip 1-205a may include a first electronic component 1-212a, and the second fixed strip 1-205b may include a second electronic component 1-212b. In at least one example, the first and second strips 1-205a-b can be removably coupled to the display unit 1-202.

[0097] Additionally, the HMD 1-200 may include an optical seal 1-210 configured to be removably coupled to the display unit 1-202. The HMD 1-200 may also include a lens 1-218 that may be removably coupled to the display unit 1-202, for example, over a first component including a display screen and a second display component. The lens 1-218 may include a custom prescription lens configured to correct vision. As noted, in Figure 1D Each of the parts shown in the exploded view of and described above can be removably coupled, attached, reattached, and replaced to update parts or swap out parts for different users. For example, bands such as band 1-216, optical seals such as optical seal 1-210, lenses such as lens 1-218, and electronic strips such as electronic strips 1-205a-b can be swapped out depending on the user so that these parts are customized to fit and correspond to a single user of the HMD 1-200.

[0098] Figure 1D Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figure 1B , Figure 1C and Figure 1E to Figure 1FAny other examples of devices, features, components and parts shown and described herein. Figure 1B , Figure 1C and Figure 1E to Figure 1F Any of the features, components and / or parts shown or described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1D Examples of devices, features, components, and parts are shown.

[0099] Figure 1E An exploded view of an example of a display unit 1-306 of an HMD is illustrated. The display unit 1-306 may include a front display assembly 1-308, a frame / housing assembly 1-350, and a curtain assembly 1-324. The display unit 1-306 may also include a sensor assembly 1-356, a logic board assembly 1-358, and a cooling assembly 1-360 disposed between the frame assembly 1-350 and the front display assembly 1-308. In at least one example, the display unit 1-306 may also include a rear display assembly 1-320, which includes a first rear display screen 1-322a and a second rear display screen 1-322b disposed between the frame 1-350 and the curtain assembly 1-324.

[0100] In at least one example, the display unit 1-306 may also include a motor assembly 1-362 configured as an adjustment mechanism for adjusting the position of the display screens 1-322a-b of the display assembly 1-320 relative to the frame 1-350. In at least one example, the display assembly 1-320 is mechanically coupled to the motor assembly 1-362, with each display screen 1-322a-b having at least one motor such that the motor can translate the display screens 1-322a-b to match the interpupillary distance of the user's eyes.

[0101] In at least one example, the display unit 1-306 may include a dial or button 1-328 that is depressible relative to the frame 1-350 and accessible by a user external to the frame 1-350. The button 1-328 may be electrically connected to the motor assembly 1-362 via a controller such that the button 1-328 may be manipulated by a user to cause a motor of the motor assembly 1-362 to adjust the position of the display screens 1-322a-b.

[0102] Figure 1E Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figures 1B to 1D and Figure 1F Any other examples of devices, features, components and parts shown and described herein. Figures 1B to 1D and Figure 1FAny of the features, components and / or parts shown and described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1E Examples of devices, features, components, and parts are shown.

[0103] Figure 1F An exploded view of another example of a display unit 1-406 of an HMD device similar to other HMD devices described herein is illustrated. The display unit 1-406 may include a front display assembly 1-402, a sensor assembly 1-456, a logic board assembly 1-458, a cooling assembly 1-460, a frame assembly 1-450, a rear display assembly 1-421, and a curtain assembly 1-424. The display unit 1-406 may also include a motor assembly 1-462 for adjusting the position of the first display subassembly 1-420a and the second display subassembly 1-420b of the rear display assembly 1-421, including the first corresponding display screen and the second corresponding display screen for interpupillary adjustment, as described above.

[0104] Figure 1F The various parts, systems and assemblies shown in exploded views herein are referenced Figures 1B to 1E and subsequent figures referenced in this disclosure are described in more detail. Figure 1F The display unit 1-406 shown can be used with Figures 1B to 1E The fixing mechanism shown is assembled and integrated, and the fixing mechanism includes electronic strips, belts and other components including optical seals, connecting components, etc.

[0105] Figure 1F Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figures 1B to 1E Any other examples of devices, features, components and parts shown and described herein. Figures 1B to 1E Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1F Examples of devices, features, components, and parts are shown.

[0106] Figure 1G A perspective exploded view of a front cover assembly 3-100 of an HMD device described herein is illustrated, for example Figure 1G The front cover assembly 3-1 of the HMD 3-100 shown, or any other HMD device shown and described herein. Figure 1GThe front cover assembly 3-100 shown may include a transparent or translucent cover 3-102, a shield 3-104 (or "canopy"), an adhesive layer 3-106, a display assembly 3-108 including a lenticular lens panel or array 3-110, and a structural trim 3-112. The adhesive layer 3-106 may secure the shield 3-104 and / or the transparent cover 3-102 to the display assembly 3-108 and / or the trim 3-112. The trim 3-112 may secure the various components of the front cover assembly 3-100 to a frame or base of the HMD device.

[0107] In at least one example, Figure 1G As shown, the transparent cover 3-102, the shield 3-104 and the display assembly 3-108 including the lenticular lens array 3-110 can be bent to adapt to the curvature of the user's face. The transparent cover 3-102 and the shield 3-104 can be bent in two or three dimensions, for example, vertically bent in the Z direction inside and outside the ZX plane, and horizontally bent in the X direction inside and outside the ZX plane. In at least one example, the display assembly 3-108 may include a lenticular lens array 3-110 and a display panel having pixels that are configured to project light through the shield 3-104 and the transparent cover 3-102. The display assembly 3-108 can be bent in at least one direction (e.g., horizontally) to adapt to the curvature of the user's face from one side of the face (e.g., the left side) to the other side (e.g., the right side). In at least one example, each layer or component of the display assembly 3-108 (which will be shown in subsequent figures and described in more detail, but which may include the lenticular lens array 3-110 and the display layer) may be curved similarly or concentrically in the horizontal direction to accommodate the curvature of the user's face.

[0108] In at least one example, the shield 3-104 may include a transparent or translucent material through which the display assembly 3-108 projects light. In one example, the shield 3-104 may include one or more opaque portions, such as an opaque ink printed portion or other opaque film portion on the back of the shield 3-104. When the HMD device is worn, the rear surface may be the surface of the shield 3-104 facing the user's eyes. In at least one example, the opaque portion may be on the front surface of the shield 3-104 opposite to the rear surface. In at least one example, one or more opaque portions of the shield 3-104 may include a peripheral portion that visually hides any components surrounding the outer periphery of the display screen of the display assembly 3-108. In this way, the opaque portion of the shield hides any other components of the HMD device that would otherwise be visible through the transparent or translucent cover 3-102 and / or the shield 3-104, including electronic components, structural components, etc.

[0109] In at least one example, the shield 3-104 may define one or more aperture transparent portions 3-120 through which the sensor may send and receive signals. In one example, the portion 3-120 is a hole through which the sensor may extend or send and receive signals. In one example, the portion 3-120 is a transparent portion, or a portion that is more transparent than the surrounding translucent or opaque portion of the shield through which the sensor may send and receive signals through the shield and through the transparent cover 3-102. In one example, the sensor may include a camera, an IR sensor, a LUX sensor, or any other visual or non-visual environmental sensor of the HMD device.

[0110] Figure 1G Any of the features, components, and / or parts shown (including arrangements and configurations thereof) may be included alone or in any combination in any other examples of devices, features, components, and parts described herein. Likewise, any of the features, components, and / or parts shown and described herein (including arrangements and configurations thereof) may be included alone or in any combination in any other examples of devices, features, components, and parts described herein. Figure 1G Examples of devices, features, components, and parts are shown.

[0111] Figure 1H An exploded view of an example of an HMD device 6-100 is illustrated. The HMD device 6-100 may include a sensor array or system 6-102 including one or more sensors, cameras, projectors, etc. mounted to one or more components of the HMD 6-100. In at least one example, the sensor system 6-102 may include a bracket 1-338 to which one or more sensors of the sensor system 6-102 may be fixed / secured.

[0112] Fig. 1I A portion of an HMD device 6-100 is illustrated that includes a front transparent cover 6-104 and a sensor system 6-102. The sensor system 6-102 may include a plurality of different sensors, emitters, receivers, including cameras, IR sensors, projectors, etc. The transparent cover 6-104 is illustrated in front of the sensor system 6-102 to illustrate the relative positions of the various sensors and emitters and the orientation of each sensor / emitter of the system 6-102. As referred to herein, "lateral," "sideways," "lateral," "horizontal," and other similar terms refer to the orientation of the sensor system 6-102. Figure 1J The x-axis is used to indicate the orientation or direction. Terms such as "vertical", "upward", "downward" and the like refer to the orientation or direction of the Figure 1J The orientation or direction is indicated by the Z-axis shown. Terms such as "forward," "rearward," "forward," "rearward" and the like refer to Figure 1J The Y-axis shown indicates the orientation or direction.

[0113] In at least one example, a transparent cover 6-104 may define a front exterior surface of the HMD device 6-100, and a sensor system 6-102 including various sensors and components thereof may be disposed in the Y axis / direction behind the cover 6-104. The cover 6-104 may be transparent or translucent to allow light to pass through the cover 6-104, including both light detected by the sensor system 6-102 and light emitted thereby.

[0114] As described elsewhere herein, the HMD device 6-100 may include one or more controllers including processors for electrically coupling the various sensors and transmitters of the sensor system 6-102 to one or more motherboards, processing units, and other electronic devices such as display screens. In addition, as will be shown in more detail below with reference to other figures, the various sensors, transmitters, and other components of the sensor system 6-102 may be coupled to Fig. 1I Various structural frame members, brackets, etc. of the HMD device 6-100 are not shown. For clarity, Fig. 1I Components of the sensor system 6-102 are shown unattached and unelectrically coupled to other components.

[0115] In at least one example, the device may include one or more controllers having a processor configured to execute instructions stored on a memory component electrically coupled to the processor. The instructions may include or cause the processor to execute one or more algorithms for self-correcting the angles and positions of the various cameras described herein over time as the initial position, angle, or orientation of the camera is bumped or deformed due to an accidental drop event or other event.

[0116] In at least one example, the sensor system 6-102 may include one or more scene cameras 6-106. The system 6-102 may include two scene cameras 6-102, respectively disposed on either side of the nose bridge or arch structure of the HMD device 6-100, such that each of the two cameras 6-106 roughly corresponds to the position of the left eye and the right eye of the user behind the cover 6-103. In at least one example, the scene camera 6-106 is generally oriented forward in the Y direction to capture images in front of the user during use of the HMD 6-100. In at least one example, the scene camera is a color camera and provides images and content for MR video pass-through to a display screen facing the user's eyes when the HMD device 6-100 is used. The scene camera 6-106 can also be used for environment and object reconstruction.

[0117] In at least one example, the sensor system 6-102 may include a first depth sensor 6-108 pointing generally forward in the Y direction. In at least one example, the first depth sensor 6-108 may be used for environment and object reconstruction and hand and body tracking of a user. In at least one example, the sensor system 6-102 may include a second depth sensor 6-110 centrally disposed along the width of the HMD device 6-100 (e.g., along the X-axis). For example, the second depth sensor 6-110 may be disposed on an adapting structure above a central nose bridge or above a nose when the user wears the HMD 6-100. In at least one example, the second depth sensor 6-110 may be used for environment and object reconstruction and hand and body tracking. In at least one example, the second depth sensor may include a LIDAR sensor.

[0118] In at least one example, the sensor system 6-102 may include a depth projector 6-112 that is generally forward facing to project electromagnetic waves (e.g., in a predetermined pattern of light dots) into or within the field of view of a user and / or scene camera 6-106, or into or within a field of view that includes and exceeds the field of view of the user and / or scene camera 6-106. In at least one example, the depth projector is capable of projecting electromagnetic waves of light in the form of a pattern of light dots that are reflected from an object and returned to the depth sensors described above, including the depth sensors 6-108, 6-110. In at least one example, the depth projector 6-112 can be used for environment and object reconstruction and hand and body tracking.

[0119] In at least one example, the sensor system 6-102 may include downward facing cameras 6-114 whose fields of view are generally pointed downward on the Z axis relative to the HMD device 6-100. In at least one example, the downward cameras 6-114 may be disposed on the left and right sides of the HMD device 6-100 as shown and used for hand and body tracking, headset tracking, and facial avatar detection and creation for displaying a user avatar on a forward facing display screen of the HMD device 6-100 as described elsewhere herein. For example, the downward cameras 6-114 may be used to capture facial expressions and movements of a user's face below the HMD device 6-100, including cheeks, mouth, and chin.

[0120] In at least one example, the sensor system 6-102 may include a jaw camera 6-116. In at least one example, the jaw camera 6-116 may be disposed on the left and right sides of the HMD device 6-100 as shown and used for hand and body tracking, headset tracking, and facial avatar detection and creation for displaying a user avatar on a forward-facing display screen of the HMD device 6-100 as described elsewhere herein. For example, the jaw camera 6-116 may be used to capture facial expressions and movements of a user's face beneath the HMD device 6-100, including the user's jaw, cheeks, mouth, and chin. For hand and body tracking, headset tracking, and facial avatar detection and creation, the user avatar may be displayed on the front display screen of the HMD device 6-100 as described elsewhere herein.

[0121] In at least one example, the sensor system 6-102 may include a side camera 6-118. The side camera 6-118 may be oriented to capture left and right side views in an X-axis or direction relative to the HMD device 6-100. In at least one example, the side camera 6-118 may be used for hand and body tracking, headset tracking, and facial avatar detection and re-creation.

[0122] In at least one example, the sensor system 6-102 may include a plurality of eye tracking and gaze tracking sensors for determining the identity, status, and gaze direction of the user's eyes during and / or prior to use. In at least one example, the eye / gaze tracking sensors may include a nose-eye camera 6-120 that is disposed on either side of the user's nose and adjacent to the user's nose when the HMD device 6-100 is worn. The eye / gaze sensors may also include a bottom eye camera 6-122 disposed below the respective user's eyes for capturing images of the eyes for facial avatar detection and creation, gaze tracking, and iris identification functions.

[0123] In at least one example, the sensor system 6-102 may include an infrared illuminator 6-124 that points outward from the HMD device 6-100 to illuminate the external environment and any objects therein with IR light for IR detection using one or more IR sensors of the sensor system 6-102. In at least one example, the sensor system 6-102 may include a flicker sensor 6-126 and an ambient light sensor 6-128. In at least one example, the flicker sensor 6-126 may detect the overhead light refresh rate to avoid display flicker. In one example, the infrared illuminator 6-124 may include a light emitting diode and may be particularly useful in low light environments for illuminating the user's hands and other objects in low light for detection by the infrared sensors of the sensor system 6-102.

[0124] In at least one example, multiple sensors (including the scene camera 6-106, the downward camera 6-114, the jaw camera 6-116, the side camera 6-118, the depth projector 6-112, and the depth sensors 6-108, 6-110) can be used in combination with an electrically coupled controller to combine depth data with camera data for hand tracking and for size determination to better perform hand tracking and object recognition and tracking functions of the HMD device 6-100. In at least one example, the above-described and Fig. 1I The downward camera 6-114, the jaw camera 6-116, and the side camera 6-118 shown in the figure can be wide-angle cameras capable of operating in the visible and infrared spectrum. In at least one example, these cameras 6-114, 6-116, 6-118 can operate only in black and white light detection to simplify image processing and gain sensitivity.

[0125] Fig. 1I Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figure 1J to Figure 1L Any other examples of devices, features, components and parts shown and described herein. Figure 1J to Figure 1L Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included alone or in any combination in the Fig. 1I Examples of devices, features, components, and parts are shown.

[0126] Figure 1J A lower perspective view of an example of an HMD 6-200 including a cover or shroud 6-204 secured to a frame 6-230 is illustrated. In at least one example, sensors 6-203 of a sensor system 6-202 may be disposed around the perimeter of the HMD 6-200 such that the sensors 6-203 are disposed outwardly around the perimeter of a display area or region 6-232 so as to not obstruct viewing of displayed light. In at least one example, the sensors may be disposed behind the shroud 6-204 and aligned with a transparent portion of the shroud, thereby allowing the sensor and projector to allow light to pass back and forth through the shroud 6-204. In at least one example, opaque ink or other opaque material or film / layer may be disposed on the shroud 6-204 around the display area 6-232 to hide components of the HMD 6-200 outside of the display area 6-232 rather than the transparent portion defined by the opaque portion through which the sensor and projector send and receive light and electromagnetic signals during operation. In at least one example, the shield 6-204 allows light to pass from the display (eg, within the display area 6-232), but does not allow light to pass radially outward from the display area around the display and the perimeter of the shield 6-204.

[0127] In some examples, the shield 6-204 includes a transparent portion 6-205 and an opaque portion 6-207, as described above and elsewhere herein. In at least one example, the opaque portion 6-207 of the shield 6-204 may define one or more transparent areas 6-209 through which the sensor 6-203 of the sensor system 6-202 may send and receive signals. In the illustrated example, the sensor 6-203 of the sensor system 6-202 sends and receives signals through the shield 6-204, or more specifically, sends and receives signals through (or defined by) the transparent area 6-209 of the opaque portion 6-207 of the shield 6-204, which sensor may include a Fig. 1I The same or similar sensors as those shown in the example of FIG. 6-108, such as depth sensors 6-108 and 6-110, depth projector 6-112, first and second scene cameras 6-106, first and second downward cameras 6-114, first and second side cameras 6-118, and first and second infrared illuminators 6-124. These sensors are also Figure 1K and Figure 1L Other sensors, sensor types, number of sensors, and their relative positions may be included in one or more other examples of the HMD.

[0128] Figure 1J Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Fig. 1I and Figure 1K to Figure 1L Any other examples of devices, features, components and parts shown and described herein. Fig. 1I and Figure 1K to Figure 1L Any of the features, components and / or parts shown or described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1J Examples of devices, features, components, and parts are shown.

[0129] Figure 1K A front view of a portion of an example of an HMD device 6-300 is illustrated, including a display 6-334, brackets 6-336, 6-338, and a frame or housing 6-330. Figure 1K The example shown does not include a front cover or shield in order to illustrate the brackets 6-336, 6-338. For example, Figure 1J The illustrated shield 6-204 includes an opaque portion 6-207 that would visually cover / block viewing of anything outside (e.g., radially / peripherally outside) the display / display area 6-334, including the sensor 6-303 and the bracket 6-338.

[0130] In at least one example, the various sensors of the sensor system 6-302 are coupled to brackets 6-336, 6-338. In at least one example, the scene cameras 6-306 include tight tolerances on angles relative to each other. For example, the tolerance on the mounting angles between two scene cameras 6-306 may be 0.5 degrees or less, such as 0.3 degrees or less. In order to achieve and maintain such tight tolerances, in one example, the scene camera 6-306 may be mounted to the bracket 6-338 instead of the shield. The bracket may include a cantilever on which the scene camera 6-306 and other sensors of the sensor system 6-302 may be mounted to maintain position and orientation in the event of a drop by a user causing any deformation of the other brackets 6-226, the housing 6-330, and / or the shield.

[0131] Figure 1K Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figures 1I to 1J and Figure 1L Any other examples of devices, features, components and parts shown and described herein. Figures 1I to 1J and Figure 1L Any of the features, components and / or parts shown or described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1K Examples of devices, features, components, and parts are shown.

[0132] Figure 1L A bottom view of an example of an HMD 6-400 including a front display / cover assembly 6-404 and a sensor system 6-402 is illustrated. The sensor system 6-402 may be similar to other sensor systems described above and elsewhere herein, including reference Figures 1I to 1K As described. In at least one example, the jaw camera 6-416 may face downward to capture images of the user's lower facial features. In one example, the jaw camera 6-416 may be directly coupled to the frame or housing 6-430 or one or more internal brackets that are directly coupled to the frame or housing 6-430 as shown. The frame or housing 6-430 may include one or more holes / openings 6-415 through which the jaw camera 6-416 can send and receive signals.

[0133] Figure 1L Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figures 1I to 1K Any other examples of devices, features, components and parts shown and described herein. Figures 1I to 1KAny of the features, components and / or parts shown and described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1L Examples of devices, features, components, and parts are shown.

[0134] Figure 1M A rear perspective view of an interpupillary distance (IPD) adjustment system 11.1.1-102 is shown, the IPD adjustment system comprising first and second optical modules 11.1.1-104a-b slidably engaged / coupled to respective guide rods 11.1.1-108a-b and motors 11.1.1-110a-b of left and right adjustment subsystems 11.1.1-106a-b. The IPD adjustment system 11.1.1-102 may be coupled to a bracket 11.1.1-112 and include a button 11.1.1-114 in electrical communication with the motors 11.1.1-110a-b. In at least one example, the button 11.1.1-114 may be in electrical communication with the first and second motors 11.1.1-110a-b via a processor or other circuit component to cause the first and second motors 11.1.1-110a-b to activate and respectively cause the first and second optical modules 11.1.1-104a-b to change position relative to each other.

[0135] In at least one example, the first and second optical modules 11.1.1-104a-b may include respective display screens configured to project light toward the user's eyes when the HMD 11.1.1-100 is worn. In at least one example, the user may manipulate (e.g., press and / or rotate) the button 11.1.1-114 to activate position adjustment of the optical modules 11.1.1-104a-b to match the interpupillary distance of the user's eyes. The optical modules 11.1.1-104a-b may also include one or more cameras or other sensors / sensor systems for imaging and measuring the user's IPD so that the optical modules 11.1.1-104a-b may be adjusted to match the IPD.

[0136] In one example, a user may manipulate the button 11.1.1-114 to cause automatic position adjustment of the first and second optical modules 11.1.1-104a-b. In one example, a user may manipulate the button 11.1.1-114 to cause manual adjustment such that the optical modules 11.1.1-104a-b move farther or closer (e.g., as the user rotates the button 11.1.1-114 one way or another) until the user visually matches her / his own IPD. In one example, the manual adjustment is communicated electronically via one or more circuits, and power for moving the optical modules 11.1.1-104a-b via the motors 11.1.1-110a-b is provided by a power source. In one example, adjustment and movement of the optical modules 11.1.1-104a-b via the manipulation button 11.1.1-114 is mechanically actuated via the movement button 11.1.1-114.

[0137] Figure 1M Any of the features, components and / or parts shown (including arrangements and configurations thereof) may be included, alone or in any combination, in any other examples of devices, features, components and parts shown in any other figures and described herein. Similarly, any of the features, components and / or parts shown or described with reference to any other figures (including arrangements and configurations thereof) may be included, alone or in any combination, in any other examples of devices, features, components and parts shown or described herein. Figure 1M Examples of devices, features, components, and parts are shown.

[0138] Figure 1N A front perspective view of a portion of an HMD 11.1.2-100 is shown, including an outer structural frame 11.1.2-102 and an inner or intermediate structural frame 11.1.2-104 defining a first aperture 11.1.2-106a and a second aperture 11.1.2-106b. Figure 1N 106a-b may be blocked by one or more other components of the HMD 11.1.2-100 coupled to the internal frame 11.1.2-104 and / or the external frame 11.1.2-102, as shown. In at least one example, the HMD 11.1.2-100 may include a first mounting bracket 11.1.2-108 coupled to the internal frame 11.1.2-104. In at least one example, the mounting bracket 11.1.2-108 is coupled to the internal frame 11.1.2-104 between the first and second apertures 11.1.2-106a-b.

[0139] The mounting bracket 11.1.2-108 may include a middle or center portion 11.1.2-109 coupled to the internal frame 11.1.2-104. In some examples, the middle or center portion 11.1.2-109 may not be the geometric middle or center of the bracket 11.1.2-108. Instead, the middle / center portion 11.1.2-109 may be disposed between a first cantilevered extension arm and a second cantilevered extension arm extending away from the middle portion 11.1.2-109. In at least one example, the mounting bracket 108 includes a first cantilever 11.1.2-112 and a second cantilever 11.1.2-114 extending away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 coupled to the internal frame 11.1.2-104.

[0140] like Figure 1N As shown, the external frame 11.1.2-102 may define a curved geometry on its underside to accommodate the user's nose when the user wears the HMD 11.1.2-100. The curved geometry may be referred to as a nose bridge 11.1.2-111 and is centrally located on the underside of the HMD 11.1.2-100 as shown. In at least one example, the mounting bracket 11.1.2-108 may be connected to the internal frame 11.1.2-104 between the holes 11.1.2-106a-b so that the cantilevers 11.1.2-112, 11.1.2-114 extend downwardly and laterally outward away from the middle portion 11.1.2-109 to complement the nose bridge 11.1.2-111 geometry of the external frame 11.1.2-102. In this way, the mounting bracket 11.1.2-108 is configured to accommodate the user's nose, as described above. The geometry of the nose bridge 11.1.2-111 adapts to the nose as the nose bridge 11.1.2-111 provides a curvature that conforms to the shape of the user's nose, providing a comfortable fit from above, over, and around.

[0141] The first cantilever arm 11.1.2-112 may extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a first direction, and the second cantilever arm 11.1.2-114 may extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a second direction opposite to the first direction. The first cantilever arm 11.1.2-112 and the second cantilever arm 11.1.2-114 are referred to as "cantilever" or "cantilever" arms because each arm 11.1.2-112, 11.1.2-114 includes a free distal end 11.1.2-116, 11.1.2-118, respectively, which are not attached to the inner frame 11.1.2-102 and the outer frame 11.1.2-104. In this way, the arms 11.1.2-112, 11.1.2-114 depend from the middle portion 11.1.2-109, which may be connected to the inner frame 11.1.2-104, while the distal ends 11.1.2-102, 11.1.2-104 are unattached.

[0142] In at least one example, the HMD 11.1.2-100 may include one or more components coupled to a mounting bracket 11.1.2-108. In one example, the components include a plurality of sensors 11.1.2-110a-f. Each of the plurality of sensors 11.1.2-110a-f may include various types of sensors, including cameras, IR sensors, and the like. In some examples, one or more of the sensors 11.1.2-110a-f may be used for object recognition in three-dimensional space, such that maintaining accurate relative positions of two or more of the plurality of sensors 11.1.2-110a-f is important. The cantilever nature of the mounting bracket 11.1.2-108 may protect the sensors 11.1.2-110a-f from damage and change of position in the event of an accidental drop by a user. Because the sensors 11.1.2-110a-f are cantilevered on the arms 11.1.2-112, 11.1.2-114 of the mounting bracket 11.1.2-108, stresses and deformations of the inner frame and / or outer frame 11.1.2-104, 11.1.2-102 are not transmitted to the cantilevered arms 11.1.2-112, 11.1.2-114 and therefore do not affect the relative positions of the sensors 11.1.2-110a-f coupled / mounted to the mounting bracket 11.1.2-108.

[0143] Figure 1NAny of the features, components, and / or parts shown (including arrangements and configurations thereof) may be included alone or in any combination in any other example of a device, feature, component described herein. Likewise, any of the features, components, and / or parts shown and described herein (including arrangements and configurations thereof) may be included alone or in any combination in any other example of a device, feature, component described herein. Figure 1N Examples of devices, features, components, and parts are shown.

[0144] Fig.1O An example of an optical module 11.3.2-100 for use in an electronic device (such as an HMD, including the HDM devices described herein) is illustrated. As shown in one or more other examples described herein, the optical module 11.3.2-100 can be one of two optical modules within the HMD, where each optical module is aligned to project light toward an eye of a user. In this way, a first optical module can project light to a first eye of a user via a display screen, and a second optical module of the same device can project light to a second eye of the user via another display screen.

[0145] In at least one example, the optical module 11.3.2-100 may include an optical frame or housing 11.3.2-102, which may also be referred to as a barrel or optical module barrel. The optical module 11.3.2-100 may also include a display 11.3.2-104 coupled to the housing 11.3.2-102, the display including one or more display screens. The display 11.3.2-104 may be coupled to the housing 11.3.2-102 such that the display 11.3.2-104 is configured to project light toward the eyes of a user when the HMD to which the display module 11.3.2-100 belongs is worn during use. In at least one example, the housing 11.3.2-102 may surround the display 11.3.2-104 and provide connection features for coupling other components of the optical module described herein.

[0146] In one example, the optical module 11.3.2-100 may include one or more cameras 11.3.2-106 coupled to the housing 11.3.2-102. The camera 11.3.2-106 may be positioned relative to the display 11.3.2-104 and the housing 11.3.2-102 such that the camera 11.3.2-106 is configured to capture one or more images of a user's eye during use. In at least one example, the optical module 11.3.2-100 may also include a light strip 11.3.2-108 surrounding the display 11.3.2-104. In one example, the light strip 11.3.2-108 is disposed between the display 11.3.2-104 and the camera 11.3.2-106. The light strip 11.3.2-108 may include a plurality of lights 11.3.2-110. The plurality of lights may include one or more light emitting diodes (LEDs) or other lights configured to project light toward the eyes of the user when the HMD is worn. The individual lights 11.3.2-110 in the light strip 11.3.2-108 may be spaced around the light strip 11.3.2-108 and thus evenly or unevenly spaced around the display 11.3.2-104 at various locations on the light strip 11.3.2-108 and around the display 11.3.2-104.

[0147] In at least one example, the housing 11.3.2-102 defines a viewing opening 11.3.2-101 through which a user can view the display 11.3.2-104 when wearing the HMD device. In at least one example, the LEDs are configured and arranged to emit light through the viewing opening 11.3.2-101 onto the user's eyes. In one example, the camera 11.3.2-106 is configured to capture one or more images of the user's eyes through the viewing opening 11.3.2-101.

[0148] As noted above, Fig.1O Each of the components and features of the illustrated optical module 11.3.2-100 may be replicated in another (eg, second) optical module provided with the HMD to interact with (eg, project light and capture images) the user's other eye.

[0149] Fig.1O Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figure 1P Any other examples of devices, features, components, and parts shown or otherwise described herein. Figure 1P Any of the features, components and / or parts shown or described herein (including arrangements and configurations thereof) may be included alone or in any combination. Fig.1OExamples of devices, features, components, and parts are shown.

[0150] Figure 1P A cross-sectional view of an example of an optical module 11.3.2-200 is shown, including a housing 11.3.2-202, a display assembly 11.3.2-204 coupled to the housing 11.3.2-202, and a lens 11.3.2-216 coupled to the housing 11.3.2-202. In at least one example, the housing 11.3.2-202 defines a first hole or channel 11.3.2-212 and a second hole or channel 11.3.2-214. The channels 11.3.2-212, 11.3.2-214 can be configured to slidably engage corresponding tracks or guides of the HMD device to allow the optical module 11.3.2-200 to adjust its position relative to the user's eyes to match the user's interpupillary distance (IPD). The housing 11.3.2-202 can slidably engage the guides to secure the optical module 11.3.2-200 in place within the HMD.

[0151] In at least one example, the optical module 11.3.2-200 may also include a lens 11.3.2-216 coupled to the housing 11.3.2-202 and disposed between the display assembly 11.3.2-204 and the user's eyes when the HMD is worn. The lens 11.3.2-216 may be configured to direct light from the display assembly 11.3.2-204 to the user's eyes. In at least one example, the lens 11.3.2-216 may be part of a lens assembly including a corrective lens that is removably attached to the optical module 11.3.2-200. In at least one example, the lens 11.3.2-216 is disposed above the light strip 11.3.2-208 and the one or more eye tracking cameras 11.3.2-206 such that the camera 11.3.2-206 is configured to capture images of the user's eyes through the lens 11.3.2-216, and the light strip 11.3.2-208 includes lights configured to project light into the user's eyes through the lens 11.3.2-216 during use.

[0152] Figure 1P Any of the features, components, and / or parts shown (including arrangements and configurations thereof) may be included alone or in any combination in any other examples of devices, features, components, and parts described herein. Likewise, any of the features, components, and / or parts shown and described herein (including arrangements and configurations thereof) may be included alone or in any combination in any other examples of devices, features, components, and parts described herein. Figure 1P Examples of devices, features, components, and parts are shown.

[0153] Figure 2is a block diagram of an example of a controller 110 according to some embodiments. While some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the embodiments disclosed herein. To this end, as a non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a central processing unit (CPU), a processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., a universal serial bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), infrared (IR), Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, a memory 220, and one or more communication buses 204 for interconnecting these components and various other components.

[0154] In some embodiments, the one or more communication buses 204 include circuits that interconnect and control communications between system components. In some embodiments, the one or more I / O devices 206 include at least one of a keyboard, a mouse, a touch pad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, and the like.

[0155] The memory 220 includes a high-speed random access memory, such as a dynamic random access memory (DRAM), a static random access memory (SRAM), a double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some embodiments, the memory 220 includes a non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 220 optionally includes one or more storage devices located away from the one or more processing units 202. The memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, the memory 220 or the non-transitory computer-readable storage medium of the memory 220 stores the following programs, modules, and data structures or subsets thereof, including an optional operating system 230 and an XR experience module 240.

[0156] The operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate single or multiple XR experiences of one or more users (e.g., a single XR experience of one or more users, or multiple XR experiences of corresponding groups of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 242, a tracking unit 244, a coordination unit 246, and a data transmission unit 248.

[0157] In some embodiments, the data acquisition unit 242 is configured to obtain Figure 1A 120, and optionally acquires data (e.g., presentation data, interaction data, sensor data, location data, etc.) from one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, data acquisition unit 242 includes instructions and / or logic for instructions as well as heuristics and metadata for the heuristics.

[0158] In some embodiments, tracking unit 244 is configured to map scene 105 and track at least display generation component 120 relative to Figure 1A The tracking unit 244 may include instructions and / or logic for instructions and heuristics and metadata for the heuristics. In some embodiments, the tracking unit 244 includes a hand tracking unit 245 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 245 is configured to track the position / location of one or more parts of the user's hand and / or the position of one or more parts of the user's hand relative to the scene 105. Figure 1A The movement of the scene 105 relative to the display generation component 120 and / or relative to a coordinate system (the coordinate system is defined relative to the user's hand). Figure 4 The hand tracking unit 245 is described in more detail. In some embodiments, the eye tracking unit 243 is configured to track the position or movement of the user's gaze (or more generally, the user's eyes, face, or head) relative to the scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hands)) or relative to the XR content displayed via the display generation component 120. Figure 5 The eye tracking unit 243 is described in more detail.

[0159] In some embodiments, coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by display generation component 120, and optionally by one or more of output device 155 and / or peripherals 195. To this end, in various embodiments, coordination unit 246 includes instructions and / or logic for instructions and heuristics and metadata for the heuristics.

[0160] In some embodiments, the data sending unit 248 is configured to send data (e.g., presentation data, position data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For this purpose, in various embodiments, the data sending unit 248 includes instructions and / or logic for the instructions as well as heuristics and metadata for the heuristics.

[0161] Although the data acquisition unit 242, the tracking unit 244 (e.g., including the eye tracking unit 243 and the hand tracking unit 245), the coordination unit 246, and the data transmission unit 248 are illustrated as residing on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 242, the tracking unit 244 (e.g., including the eye tracking unit 243 and the hand tracking unit 245), the coordination unit 246, and the data transmission unit 248 may be located in separate computing devices.

[0162] also, Figure 2 It serves more as a functional description of various features that may be present in a particular implementation, as opposed to a block diagram of the embodiments described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 2 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation, and in some embodiments, depends in part on the specific combination of hardware, software and / or firmware selected for a specific implementation.

[0163] Figure 31 is a block diagram of an example of a display generation component 120 according to some embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the embodiments disclosed herein. For this purpose, as a non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., a microprocessor, an ASIC, an FPGA, a GPU, a CPU, a processing core, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional internal-facing and / or external-facing image sensors 314, a memory 320, and one or more communication buses 304 for interconnecting these and various other components.

[0164] In some embodiments, one or more communication buses 304 include circuits for interconnecting and controlling communications between various system components. In some embodiments, one or more I / O devices and sensors 306 include an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, and / or one or more depth sensors (e.g., structured light, time of flight, etc.), etc.

[0165] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to the user. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LcoS), organic light emitting field effect transistor (OLET), organic light emitting diode (OLED), surface conduction electron emission display (SED), field emission display (FED), quantum dot light emitting diode (QD-LED), microelectromechanical system (MEMS) and / or similar display types. In some embodiments, one or more XR displays 312 correspond to diffraction, reflection, polarization, holographic and other waveguide displays. For example, the display generation component 120 (e.g., HMD) includes a single XR display. As another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, one or more XR displays 312 can present MR and VR content. In some embodiments, one or more XR displays 312 can present MR or VR content.

[0166] In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye tracking camera). In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to the user's hands and, optionally, at least a portion of the user's arms (and may be referred to as a hand tracking camera). In some embodiments, the one or more image sensors 314 are configured to face forward so as to acquire image data corresponding to a scene that the user would see in the absence of the display generating component 120 (e.g., an HMD) (and may be referred to as a scene camera). The one or more optional image sensors 314 may include one or more RGB cameras (e.g., with a complementary metal oxide semiconductor (CMOS) image sensor or a charge coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.

[0167] The memory 320 includes a high-speed random access memory, such as a DRAM, SRAM, DDR RAM, or other random access solid-state memory device. In some embodiments, the memory 320 includes a non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 320 optionally includes one or more storage devices located away from the one or more processing units 302. The memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, the memory 320 or the non-transitory computer-readable storage medium of the memory 320 stores the following programs, modules, and data structures or a subset thereof, including an optional operating system 330 and an XR rendering module 340.

[0168] The operating system 330 includes processes for handling various basic system services and for performing hardware-related tasks. In some embodiments, the XR rendering module 340 is configured to present XR content to a user via one or more XR displays 312. For such purposes, in various embodiments, the XR rendering module 340 includes a data acquisition unit 342, an XR rendering unit 344, an XR map generation unit 346, and a data transfer unit 348.

[0169] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 of Figure 1. For such purposes, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.

[0170] In some embodiments, the XR rendering unit 344 is configured to render XR content via one or more XR displays 312. For such purposes, in various embodiments, the XR rendering unit 344 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.

[0171] In some embodiments, the XR map generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which a computer-generated object can be placed to generate an extended reality) based on the media content data. For this purpose, in various embodiments, the XR map generation unit 346 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.

[0172] In some embodiments, the data sending unit 348 is configured to send data (e.g., presentation data, position data, etc.) to at least the controller 110, and optionally one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For such purposes, in various embodiments, the data sending unit 348 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.

[0173] Although the data acquisition unit 342, the XR rendering unit 344, the XR mapping generation unit 346, and the data sending unit 348 are shown as residing on a single device (e.g., the display generation component 120 of Figure 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR rendering unit 344, the XR mapping generation unit 346, and the data sending unit 348 may be located in separate computing devices.

[0174] also, Figure 3 More as a functional description of various features that may be present in a particular embodiment, rather than a schematic diagram of the structures of the embodiments described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 3 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation, and in some embodiments, depends in part on the specific combination of hardware, software and / or firmware selected for a specific implementation.

[0175] Figure 4 is a schematic illustration of an example implementation of the hand tracking device 140. In some implementations, the hand tracking device 140 (FIG. 1) is controlled by the hand tracking unit 245 ( Figure 2 ) to track the position / location of one or more parts of a user's hand and / or one or more parts of a user's hand relative to Figure 1A The hand tracking device 140 can be used to monitor movement of the scene 105 (e.g., relative to a portion of the physical environment surrounding the user, relative to the display generation component 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system (which is defined relative to the user's hands)). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to the head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0176] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures three-dimensional scene information including at least a human user's hand 406. The image sensor 404 captures hand images with sufficient resolution so that fingers and their corresponding positioning can be distinguished. The image sensor 404 typically captures images of other parts of the user's body, or may capture images of all parts of the body, and may have zoom capabilities or a dedicated sensor with increased magnification to capture images of the hand with a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors to capture the physical environment of the scene 105, or is used as an image sensor to capture the physical environment of the scene 105. In some embodiments, the image sensor is positioned relative to the user or the user's environment in a manner that the field of view of the image sensor 404 or a portion thereof is used to define an interaction space, in which the hand movements captured by the image sensor are considered as input to the controller 110.

[0177] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D image data (and, in addition, possible color image data) to the controller 110, which extracts high-level information from the image data. The high-level information is typically provided to an application running on the controller via an application program interface (API), which drives the display generation component 120 accordingly. For example, a user can interact with software running on the controller 110 by moving their hands 406 and / or changing their hand postures.

[0178] In some embodiments, the image sensor 404 projects a speckled pattern onto a scene containing the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral offset of the spots in the pattern. This method is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. The method gives the depth coordinates of points in the scene relative to a predetermined reference plane at a specific distance from the image sensor 404. In the present disclosure, it is assumed that the image sensor 404 defines an orthogonal set of x-axis, y-axis, and z-axis, so that the depth coordinates of points in the scene correspond to the z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) may use other 3D mapping methods based on a single or multiple cameras or other types of sensors, such as stereo imaging or time-of-flight measurement.

[0179] In some embodiments, the hand tracking device 140 captures and processes a time series of depth maps containing the user's hand as the user moves their hand (e.g., the entire hand or one or more fingers). Software running on the image sensor 404 and / or a processor in the controller 110 processes the 3D map data to extract image patch descriptors of the hand in these depth maps. The software can match these descriptors with image patch descriptors stored in the database 408 based on a previous learning process to estimate the pose of the hand in each frame. The pose typically includes the 3D location of the user's hand joints and finger tips.

[0180] The software can also analyze the trajectory of the hand and / or finger over multiple frames in the sequence to identify gestures. The pose estimation function described herein can be alternated with the motion tracking function so that the image block-based pose estimation is performed only once every two (or more) frames, and the tracking is used to find the changes in pose that occur on the remaining frames. The pose, motion, and gesture information is provided to the application running on the controller 110 via the above-mentioned API. The program can, for example, move and modify the image presented on the display generation component 120 in response to the pose and / or gesture information, or perform other functions.

[0181] In some embodiments, gestures include air gestures. An air gesture is a gesture detected without the user touching an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140) (or independent of an input element that is part of a device) and based on detected movement of a part of the user's body (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs) through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of a user's finger relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture in which the hand moves a predetermined amount and / or speed in a predetermined posture, or a shake gesture including a predetermined speed or amount of rotation of a part of the user's body)).

[0182] In some embodiments, according to some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures for interacting with an XR environment (e.g., a virtual or mixed reality environment) performed by movement of a user's fingers relative to other fingers or parts of the user's hand. In some embodiments, an air gesture is a gesture detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and based on detected movement of a part of the user's body through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture in which the hand moves a predetermined amount and / or speed in a predetermined posture, or a shake gesture including a predetermined speed or rotation amount of a part of the user's body)).

[0183] In some embodiments where the input gesture is an in-air gesture (e.g., in the absence of physical contact with an input device that provides information to the computer system about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touch screen, or contact with a mouse or trackpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Thus, in embodiments involving in-air gestures, for example, the input gesture is detection of attention (e.g., gaze) toward a user interface element combined with movement of the user's fingers and / or hand (e.g., simultaneously) to perform a pinch and / or tap input, as described in more detail below.

[0184] In some embodiments, an input gesture pointing to a user interface object is performed directly or indirectly with reference to the user interface object. For example, the user input is performed directly on the user interface object according to performing the input with the user's hand at a location corresponding to the location of the user interface object in the three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, when the user's attention (e.g., gaze) to the user interface object is detected, the input gesture is performed indirectly on the user interface object according to the user's hand being located not at the location corresponding to the location of the user interface object in the three-dimensional environment while the user performs the input gesture. For example, for a direct input gesture, the user can direct the user's input to the user interface object by initiating a gesture at or near a location corresponding to the display location of the user interface object (e.g., within 0.5 cm, 1 cm, 5 cm, or a distance between 0 and 5 cm measured from the outer edge of the option or the center portion of the option). For an indirect input gesture, the user can direct the user's input to the user interface object by focusing on the user interface object (e.g., by gazing at the user interface object), and while focusing on the option, the user initiates an input gesture (e.g., at any location detectable by the computer system) (e.g., at a location that does not correspond to the display location of the user interface object).

[0185] In some embodiments, according to some embodiments, the input gestures (e.g., air gestures) used in the various examples and embodiments described herein include pinch inputs and tap inputs for interacting with a virtual or mixed reality environment. For example, the pinch inputs and tap inputs described below are performed as air gestures.

[0186] In some embodiments, the pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch and drag gesture, or a double pinch gesture. For example, a pinch gesture as an air gesture includes movement of two or more fingers of a hand to contact each other, that is, optionally followed by an immediate (e.g., within 0 seconds to 1 second) interruption of contact with each other. A long pinch gesture as an air gesture includes movement of two or more fingers of a hand in contact with each other for at least a threshold amount of time (e.g., at least 1 second) before an interruption of contact with each other is detected. For example, a long pinch gesture includes the user maintaining a pinch gesture (e.g., in which two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between the two or more fingers is detected. In some embodiments, a double pinch gesture as an air gesture includes two (e.g., or more) pinch inputs (e.g., performed by the same hand) that are detected consecutively immediately (e.g., within a predefined time period) with each other. For example, the user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., interrupts contact between two or more fingers), and performs a second pinch input within a predefined time period after releasing the first pinch input (e.g., within 1 second or within 2 seconds).

[0187] In some embodiments, the pinch and drag gesture as an air gesture (e.g., an air drag gesture or an air swipe gesture) includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in conjunction with (e.g., following) a drag input that changes the position of the user's hand from a first position (e.g., a starting position for dragging) to a second position (e.g., an ending position for dragging). In some embodiments, the user maintains the pinch gesture while performing the drag input, and releases the pinch gesture (e.g., opens their two or more fingers) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., the user pinches two or more fingers to contact each other and moves the same hand to the second position in the air using the drag gesture). In some embodiments, a pinch input is performed by a first hand of a user, and a drag input is performed by a second hand of the user (e.g., the second hand of the user moves from a first position to a second position in the air while the user continues the pinch input with the first hand of the user. In some embodiments, the input gesture as an air gesture includes an input performed using both hands of the user (e.g., a pinch and / or tap input). For example, the input gesture includes two (e.g., or more) pinch inputs performed in conjunction with each other (e.g., concurrently or within a predefined time period). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch and drag input) is performed using the first hand of the user, and a second pinch input is performed using another hand (e.g., a second hand of the user's two hands) in conjunction with the pinch input performed using the first hand. In some embodiments, movement between the two hands of the user is performed (e.g., increasing and / or decreasing the distance or relative orientation between the two hands of the user).

[0188] In some embodiments, a tap input performed as an air gesture (e.g., pointing to a user interface element) includes movement of a user's finger toward the user interface element, movement of a user's hand toward the user interface element (optionally, extension of the user's finger toward the user interface element), downward movement of a user's finger (e.g., mimicking a mouse click motion or a tap on a touch screen), or other predefined movement of a user's hand. In some embodiments, a tap input performed as an air gesture is detected based on movement characteristics of the finger or hand performing the tap gesture movement of the finger or hand, which is a movement of the finger or hand away from the user's viewpoint and / or toward an object that is the target of the tap input, followed by an end of the movement. In some embodiments, the end of the movement is detected based on a change in the movement characteristics of the finger or hand performing the tap gesture (e.g., the end of the movement away from the user's viewpoint and / or toward an object that is the target of the tap input, a reversal of the direction of movement of the finger or hand, and / or a reversal of the acceleration direction of the movement of the finger or hand).

[0189] In some embodiments, it is determined that the user's attention is directed to a portion of the three-dimensional environment based on detection of a gaze directed to that portion of the three-dimensional environment (optionally, no other conditions are required). In some embodiments, it is determined that the user's attention is directed to a portion of the three-dimensional environment based on detection of a gaze directed to that portion of the three-dimensional environment using one or more additional conditions, such as requiring the gaze to be directed to the portion of the three-dimensional environment for at least a threshold duration (e.g., a dwell duration) and / or requiring the gaze to be directed to the portion of the three-dimensional environment when the user's viewpoint is within a distance threshold from the portion of the three-dimensional environment, so that the device determines that the user's attention is directed to the portion of the three-dimensional environment, wherein if one of these additional conditions is not met, the device determines that the attention is not directed to the portion of the three-dimensional environment to which the gaze is directed (e.g., until the one or more additional conditions are met).

[0190] In some embodiments, detection of a ready state configuration of a user or a portion of a user is detected by a computer system. Detection of a ready state configuration of a hand is used by the computer system as an indication that the user may be preparing to interact with the computer system using one or more air gesture inputs (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein) performed by the hand. For example, the ready state of the hand is determined based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape with the thumb and one or more fingers extended and spaced apart in preparation for a pinch or grab gesture, or a pre-tap shape with one or more fingers extended and the palm facing away from the user), based on whether the hand is in a predetermined position relative to the user's viewpoint (e.g., below the user's head and above the user's waist and extending at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body) and / or based on whether the hand has moved in a particular manner (e.g., toward an area in front of the user above the user's waist and below the user's head or away from the user's body or legs). In some embodiments, the ready state is used to determine whether an interactive element of a user interface responds to attention (e.g., gaze) input.

[0191] In scenarios where input is described with reference to in-air gestures, it should be understood that similar gestures may be detected using a hardware input device attached to or held by one or more hands of a user, where the positioning of the hardware input device in space may be tracked using optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units, and the positioning and / or movement of the hardware input device is used instead of the positioning and / or movement of the one or more hands in the corresponding in-air gesture. In scenarios where input is described with reference to air gestures, it should be understood that similar gestures may be detected using a hardware input device attached to or held by one or more hands of a user, user input may be detected using controls contained in the hardware input device, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, one or more hand or finger overlays that can detect the positioning or changes in positioning of parts of a hand and / or finger relative to each other, relative to the user's body and / or relative to the user's physical environment, and / or other hardware input device controls, where user input performed using controls contained in the hardware input device is used in place of hand and / or finger gestures such as an air tap or air pinch in a corresponding air gesture. For example, a selection input described as being performed using an air tap or air pinch input may alternatively be detected using a button press, a tap on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input. As another example, movement input described as being performed using an air pinch and drag (e.g., an air drag gesture or an air swipe gesture) may alternatively be detected based on interaction with a hardware input control (such as a button press and hold, a touch on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input following movement of a hardware input device (e.g., along with a hand associated with the hardware input device) through space. Similarly, two-handed input involving movement of hands relative to each other may be performed using one air gesture and one hardware input device in the hand that is not performing the air gesture, two hardware input devices held in different hands, or two air gestures performed by different hands using various combinations of air gestures and / or inputs detected by one or more of the above hardware input devices.

[0192] In some embodiments, the software may be downloaded to the controller 110 in electronic form, for example, over a network, or may alternatively be provided on a tangible, non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is also stored in a memory associated with the controller 110. Alternatively or in addition, some or all of the described functions of the computer may be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). Although in Figure 4Controller 110 is shown in FIG. 1 , but some or all of the processing functions of the controller may be performed by a suitable microprocessor and software or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand tracking device) or other device associated with the image sensor 404, for example, as a separate unit from the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television receiver, handheld device, or head-mounted device) or with any other suitable computerized device (such as a game console or media player). The sensing functions of the image sensor 404 may also be integrated into a computer or other computerized device to be controlled by the sensor output.

[0193] Figure 4 Also included is a schematic diagram of a depth map 410 captured by the image sensor 404 according to some embodiments. As described above, the depth map includes a matrix of pixels with corresponding depth values. Pixels 412 corresponding to the hand 406 have been segmented from the background and wrist in the figure. The brightness of each pixel in the depth map 410 is inversely proportional to its depth value (i.e., the measured z distance from the image sensor 404), where the gray shades become darker as the depth increases. The controller 110 processes these depth values ​​in order to identify and segment the components of the image that have human hand characteristics (i.e., a group of adjacent pixels). These characteristics may include, for example, overall size, shape, and movement from frame to frame in the depth map sequence.

[0194] Figure 4 Also schematically illustrated is a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406 according to some embodiments. Figure 4 In the hand skeleton 414, a hand background 416 that has been segmented from the original depth map is superimposed. In some embodiments, key feature points of the hand and optionally on the wrist or arm connected to the hand (e.g., points corresponding to knuckles, finger tips, center of the palm, end of the hand connected to the wrist, etc.) are identified and located on the hand skeleton 414. In some embodiments, the controller 110 uses the position and movement of these key feature points over multiple image frames to determine the gesture performed by the hand or the current state of the hand according to some embodiments.

[0195] Figure 5 An example implementation of the eye tracking device 130 (FIG. 1) is illustrated. In some implementations, the eye tracking device 130 is comprised of an eye tracking unit 243 ( Figure 2) controls to track the position and movement of the user's gaze relative to the scene 105 or relative to the XR content displayed via the display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as headphones, helmets, goggles, or glasses) or a handheld device placed in a wearable frame, the head-mounted device includes both components for generating XR content for viewing by the user and components for tracking the user's gaze relative to the XR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a handheld device or an XR room, the eye tracking device 130 is optionally a device separate from the handheld device or the XR room. In some embodiments, the eye tracking device 130 is a head-mounted device or a part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used in conjunction with a display generation component that is also head-mounted or a display generation component that is not head-mounted. In some embodiments, the eye tracking device 130 is not a head mounted device, and is optionally used in conjunction with a head mounted display generation component. In some embodiments, the eye tracking device 130 is not a head mounted device, and is optionally part of a non-head mounted display generation component.

[0196] In some embodiments, the display generation component 120 uses a display mechanism (e.g., a left near-eye display panel and a right near-eye display panel) to display a frame including a left image and a right image in front of the user's eyes, thereby providing a 3D virtual view to the user. For example, the head-mounted display generation component may include a left optical lens and a right optical lens (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external cameras that capture a video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display, and display a virtual object on the transparent or translucent display, and the user can directly view the physical environment through the transparent or translucent display. In some embodiments, the display generation component projects the virtual object into the physical environment. The virtual object may, for example, be projected on a physical surface or projected as a hologram so that an individual using the system observes the virtual object superimposed on the physical environment. In this case, separate display panels and image frames for the left and right eyes may not be required.

[0197] like Figure 5As shown, in some embodiments, the eye tracking device 130 (e.g., gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) or near infrared (NIR) camera), and an illumination source (e.g., an IR or NIR light source, such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera can be pointed at the user's eyes to receive the IR or NIR light reflected directly from the eyes by the light source, or alternatively can be pointed at "hot" mirrors located between the user's eyes and the display panel, which reflect the IR or NIR light from the eyes to the eye tracking camera while allowing visible light to pass through. The eye tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the images to generate gaze tracking information, and transmits the gaze tracking information to the controller 110. In some embodiments, the user's two eyes are tracked separately by corresponding eye tracking cameras and illumination sources. In some embodiments, only one eye of the user is tracked by the corresponding eye tracking camera and illumination source.

[0198] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye tracking device for a specific operating environment 100, such as the 3D geometric relationships and parameters of the LED, camera, thermal mirror (if present), eye lens, and display screen. The device-specific calibration process can be performed at a factory or another facility before the AR / VR equipment is delivered to the end user. The device-specific calibration process can be an automatic calibration process or a manual calibration process. According to some embodiments, the user-specific calibration process may include an estimate of the eye parameters of a specific user, such as pupil position, foveal position, optical axis, visual axis, eye spacing, etc. According to some embodiments, once the device-specific parameters and user-specific parameters are determined for the eye tracking device 130, a flash-assisted method can be used to process the images captured by the eye tracking camera to determine the current visual axis and the user's gaze point relative to the display.

[0199] like Figure 5As shown, the eye tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system, which includes at least one eye tracking camera 540 (e.g., an infrared (IR) or near infrared (NIR) camera) positioned on the side of the user's face on which eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source, such as an array or ring of NIR light emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye 592. The eye tracking camera 540 can be directed toward a mirror 550 (which reflects IR or NIR light from the eye 592 while allowing visible light to pass) located between the user's eye 592 and a display 510 (e.g., a left display panel or a right display panel of a head-mounted display, or a display of a handheld device, a projector, etc.) (e.g., as Figure 5 ), or alternatively may be directed toward the user's eye 592 to receive reflected IR or NIR light from the eye 592 (e.g., as shown in the top portion of Figure 5 ).

[0200] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames for left and right display panels) and provides the frames 562 to the display 510. The controller 110 uses the gaze tracking input 542 from the eye tracking camera 540 for various purposes, such as for processing the frames 562 for display. The controller 110 optionally estimates the user's gaze point on the display 510 based on the gaze tracking input 542 obtained from the eye tracking camera 540 using a flash-assisted method or other suitable method. The gaze point estimated from the gaze tracking input 542 is optionally used to determine the direction in which the user is currently looking.

[0201] Several possible use cases for the user's current gaze direction are described below and are not intended to be limiting. As an example use case, the controller 110 may render virtual content differently based on the determined direction of the user's gaze. For example, the controller 110 may generate virtual content at a higher resolution in the foveal area determined according to the user's current gaze direction than in the peripheral area. As another example, the controller may position or move virtual content in a view based at least in part on the user's current gaze direction. As another example, the controller may display specific virtual content in a view based at least in part on the user's current gaze direction. As another example use case in an AR application, the controller 110 may guide an external camera for capturing the physical environment of the XR experience to focus in the determined direction. The autofocus mechanism of the external camera may then focus on an object or surface in the environment that the user is currently looking at on the display 510. As another example use case, the eye lens 520 may be a focusable lens, and the controller uses gaze tracking information to adjust the focus of the eye lens 520 so that the virtual object that the user is currently looking at has an appropriate degree of convergence to match the convergence of the user's eye 592. The controller 110 may utilize the gaze tracking information to guide the eye lenses 520 to adjust the focus so that nearby objects that the user is looking at appear at the correct distance.

[0202] In some embodiments, the eye tracking device is part of a head mounted device that includes a display (e.g., display 510), two eye lenses (e.g., eye lenses 520), an eye tracking camera (e.g., eye tracking camera 540), and a light source (e.g., light source 530 (e.g., IR or NIR LED)). The light source emits light (e.g., IR or NIR light) toward the user's eye 592. In some embodiments, the light sources may be arranged in a ring or circle around each of the lenses, such as Figure 5 In some embodiments, for example, eight light sources 530 (eg, LEDs) are arranged around each lens 520. However, more or fewer light sources 530 may be used, and other arrangements and locations of the light sources 530 may be used.

[0203] In some embodiments, the display 510 emits light in the visible range and does not emit light in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. Note that the positions and angles of the eye tracking cameras 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.

[0204] like Figure 5 Embodiments of the gaze tracking system illustrated in the drawings may be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a user with a computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experience.

[0205] Figure 6 A flash-assisted gaze tracking pipeline according to some embodiments is illustrated. In some embodiments, the gaze tracking pipeline is implemented by a flash-assisted gaze tracking system (e.g., Figure 1A and Figure 5 The flash-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in the tracking state, the flash-assisted gaze tracking system uses previous information from previous frames when analyzing the current frame to track the pupil outline and the flash in the current frame. When not in the tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and the flash in the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in the tracking state.

[0206] like Figure 6 As shown, the gaze tracking camera can capture left and right images of the user's left and right eyes. The captured images are then input to the gaze tracking pipeline for processing starting at 610. As indicated by the arrow returning to element 600, the gaze tracking system can continue to capture images of the user's eyes at a rate of, for example, 60 to 120 frames per second. In some embodiments, each group of captured images can be input to the pipeline for processing. However, in some embodiments or under some conditions, not all captured frames are processed by the pipeline.

[0207] At 610, for the currently captured image, if the tracking status is yes, the method proceeds to element 640. At 610, if the tracking status is no, the image is analyzed to detect the user's pupil and glint in the image, as indicated at 620. At 630, if the pupil and glint are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eye.

[0208] At 640, if advancing from element 610, the current frame is analyzed to track pupils and glints based in part on previous information from previous frames. At 640, if advancing from element 630, the tracking state is initialized based on the pupils and glints detected in the current frame. The processing results at element 640 are checked to verify that the results of the tracking or detection can be credible. For example, the results can be checked to determine whether the pupil and a sufficient number of glints for performing gaze estimation are successfully tracked or detected in the current frame. At 650, if the results are not likely to be credible, at element 660, the tracking state is set to no, and the method returns to element 610 to process the next image of the user's eye. At 650, if the results are credible, the method advances to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's gaze point.

[0209] Figure 6 It is intended to be used as an example of an eye tracking technology that can be used for a particular implementation. As recognized by one of ordinary skill in the art, according to various embodiments, other eye tracking technologies currently existing or developed in the future can be used in place of or in combination with the flash-assisted eye tracking technology described herein in a computer system 101 for providing an XR experience to a user.

[0210] In some embodiments, the captured portion of the real-world environment 602 is used to provide an XR experience to the user, such as a mixed reality environment in which one or more virtual objects are superimposed on top of a representation of the real-world environment 602 .

[0211] Thus, the description herein describes some embodiments of a three-dimensional environment (e.g., an XR environment) including representations of real-world objects and representations of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table present in a physical environment, which is captured and displayed in the three-dimensional environment (e.g., actively displayed via a camera and display of a computer system or passively displayed via a transparent or translucent display of a computer system). As previously described, the three-dimensional environment is optionally a mixed reality system, wherein the three-dimensional environment is based on a physical environment captured by one or more sensors of a computer system and displayed via a display generation component. As a mixed reality system, the computer system is optionally capable of selectively displaying parts and / or objects of the physical environment so that the corresponding parts and / or objects of the physical environment appear as if they exist in the three-dimensional environment displayed by the computer system. Similarly, the computer system is optionally capable of displaying virtual objects in a three-dimensional environment by placing virtual objects at corresponding positions in the three-dimensional environment that have corresponding positions in the real world to appear as if the virtual objects exist in the real world (e.g., a physical environment). For example, the computer system optionally displays a vase so that the vase appears as if a real vase is placed on top of a table in the physical environment. In some embodiments, a corresponding position in the three-dimensional environment has a corresponding position in the physical environment. Thus, when a computer system is described as displaying a virtual object at a corresponding position relative to a physical object (e.g., such as a position at or near a user's hand or at or near a physical table), the computer system displays the virtual object at a specific position in the three-dimensional environment so that it appears as if the virtual object is at or near the physical object in the physical environment (e.g., the virtual object is displayed at a position in the three-dimensional environment that corresponds to the position in the physical environment where the virtual object would be displayed if it were a real object at that specific position).

[0212] In some embodiments, real-world objects present in the physical environment that are displayed in the three-dimensional environment (e.g., and / or visible via a display generation component) can interact with virtual objects that only exist in the three-dimensional environment. For example, the three-dimensional environment may include a table and a vase placed on top of the table, where the table is a view (or representation) of a physical table in the physical environment and the vase is a virtual object.

[0213] In a three-dimensional environment (e.g., a real environment, a virtual environment, or an environment comprising a mixture of real objects and virtual objects), an object is sometimes referred to as having depth or simulated depth, or an object is referred to as being visible, displayed, or placed at different depths. In this context, depth refers to a dimension different from height or width. In some embodiments, depth is defined relative to a fixed coordinate set (e.g., where a room or object has a height, depth, and width defined relative to a fixed coordinate set). In some embodiments, depth is defined relative to the user's position or viewpoint, in which case the depth dimension varies based on the position and angle of the user's position and / or the user's viewpoint. In some embodiments where depth is defined relative to the user's position relative to the surface of the environment (e.g., the surface of the floor or ground of the environment), an object farther away from the user along a line extending parallel to the surface is considered to have greater depth in the environment, and / or the depth of an object is measured along an axis extending outward from the user's position and parallel to the surface of the environment (e.g., depth is defined in a cylindrical or substantially cylindrical coordinate system, where the user's position is at the center of a cylinder extending from the user's head toward the user's feet). In some embodiments where depth is defined relative to a user's viewpoint (e.g., relative to a direction of a point in space that determines which portion of an environment is visible via a head-mounted device or other display), objects that are farther away from the user's viewpoint along a line extending parallel to the user's viewpoint are considered to have greater depth in the environment, and / or the depth of an object is measured along an axis extending outward from the user's viewpoint and parallel to the user's viewpoint (e.g., depth is defined in a spherical or substantially spherical coordinate system with the origin of the viewpoint at the center of a sphere extending outward from the user's head). In some embodiments, depth is defined relative to a user interface container (e.g., a window or application in which applications and / or system content are displayed), where the user interface container has a height and / or width, and the depth is a dimension orthogonal to the height and / or width of the user interface container. In some embodiments, where the depth is defined relative to a user interface container, when the container is placed in a three-dimensional environment or is initially displayed (e.g., such that the depth dimension of the container extends outward away from the user or the user's viewpoint), the height and / or width of the container is generally orthogonal or substantially orthogonal to a straight line extending from a user-based position (e.g., the user's viewpoint or the user's position) to the user interface container (e.g., the center of the user interface container or another feature point of the user interface container). In some embodiments, where the depth is defined relative to the user interface container, the depth of an object relative to the user interface container refers to the position of the object along the depth dimension of the user interface container. In some embodiments, multiple different containers may have different depth dimensions (e.g., different depth dimensions extending away from the user or the user's viewpoint in different directions and / or from different starting points).In some embodiments, when depth is defined relative to a user interface container, the direction of the depth dimension remains constant for the user interface container as the position of the user interface container, the user, and / or the user's viewpoint changes (e.g., or when multiple different viewers are viewing the same container in a three-dimensional environment, such as during an in-person collaboration session and / or when multiple participants are in a real-time communication session with shared virtual content that includes the container). In some embodiments, for curved containers (e.g., including containers with curved surfaces or curved content areas), the depth dimension optionally extends into the surface of the curved container. In some cases, z separation (e.g., the separation of two objects in the depth dimension), z height (e.g., the distance of one object from another object in the depth dimension), z position (e.g., the position of an object in the depth dimension), z depth (e.g., the position of an object in the depth dimension), or simulated z dimension (e.g., depth used as a dimension of an object, a dimension of an environment, a direction in space, and / or a direction in simulated space) is used to refer to the concept of depth as described above.

[0214] In some embodiments, the user is optionally able to use one or both hands to interact with virtual objects in a three-dimensional environment as if the virtual objects were real objects in a physical environment. For example, as described above, one or more sensors of a computer system optionally capture one or more hands of a user and display representations of the user's hands in a three-dimensional environment (e.g., in a manner similar to displaying real-world objects in a three-dimensional environment as described above), or in some embodiments, due to the transparency / translucency of a portion of a display generation component that is displaying a user interface, or due to a projection of a user interface onto a transparent / translucent surface or a projection of a user interface onto a user's eyes or into the field of view of a user's eyes, the user's hands can be seen via the display generation component, via the ability to see the physical environment through the user interface. Therefore, in some embodiments, the user's hands are displayed at corresponding locations in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment, which can interact with virtual objects in the three-dimensional environment as if these virtual objects were physical objects in the physical environment. In some embodiments, the computer system can update the display of the representation of the user's hands in the three-dimensional environment in conjunction with the movement of the user's hands in the physical environment.

[0215] In some of the embodiments described below, the computer system is optionally capable of determining an "effective" distance between a physical object in the physical world and a virtual object in a three-dimensional environment, for example, to determine whether the physical object is directly interacting with the virtual object (e.g., whether the hand is touching, grabbing, holding, etc. a virtual object or is within a threshold distance of the virtual object). For example, a hand that directly interacts with a virtual object optionally includes one or more of the following: a finger of a hand pressing a virtual button, a user's hand grabbing a virtual vase, a user's hand closing together and pinching / holding the user interface of an application, and two fingers that perform any other type of interaction described herein. For example, when determining whether a user is interacting with a virtual object and / or how a user is interacting with a virtual object, the computer system optionally determines the distance between the user's hand and the virtual object. In some embodiments, the computer system determines the distance between the user's hand and the virtual object by determining the distance between the position of the hand in the three-dimensional environment and the position of the virtual object of interest in the three-dimensional environment. For example, the one or more hands of the user are located at a specific location in the physical world, and the computer system optionally captures the one or more hands and displays the one or more hands at a specific corresponding location in the three-dimensional environment (e.g., if the hand is a virtual hand instead of a physical hand, the location where the hand will be displayed in the three-dimensional environment). The location of the hand in the three-dimensional environment is optionally compared with the location of the virtual object of interest in the three-dimensional environment to determine the distance between the one or more hands of the user and the virtual object. In some embodiments, the computer system optionally determines the distance between the physical object and the virtual object by comparing the location in the physical world (e.g., instead of comparing the location in the three-dimensional environment). For example, when determining the distance between the one or more hands of the user and the virtual object, the computer system optionally determines the corresponding position of the virtual object in the physical world (e.g., if the virtual object is a physical object instead of a virtual object, the location where the virtual object will be located in the physical world), and then determines the distance between the corresponding physical location and the one or more hands of the user. In some embodiments, the same technology is optionally used to determine the distance between any physical object and any virtual object. Therefore, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, a computer system optionally executes any of the techniques described above to map the position of the physical object to a three-dimensional environment and / or to map the position of the virtual object to the physical environment.

[0216] In some embodiments, the same or similar techniques are used to determine where and what the user's gaze is directed to, and / or where and what the physical stylus held by the user is pointed to. For example, if the user's gaze is directed to a particular location in the physical environment, the computer system optionally determines the corresponding location in the three-dimensional environment (e.g., the virtual location of the gaze), and if a virtual object is located at the corresponding virtual location, the computer system optionally determines that the user's gaze is directed to the virtual object. Similarly, the computer system is optionally able to determine the direction in which the stylus is pointing in the physical environment based on the orientation of the physical stylus. In some embodiments, based on this determination, the computer system determines a corresponding virtual location in the three-dimensional environment that corresponds to the location in the physical environment that the stylus is pointing to, and optionally determines that the stylus is pointing to the corresponding virtual location in the three-dimensional environment.

[0217] Similarly, the embodiments described herein may refer to the position of a user (e.g., a user of a computer system) in a three-dimensional environment and / or the position of a computer system in a three-dimensional environment. In some embodiments, the user of the computer system is holding, wearing, or otherwise located at or near the computer system. Therefore, in some embodiments, the position of the computer system is used as a proxy for the position of the user. In some embodiments, the position of the computer system and / or the user in the physical environment corresponds to the corresponding position in the three-dimensional environment. For example, the position of the computer system will be a position in the physical environment (and its corresponding position in the three-dimensional environment), and if the user stands at this position, facing the corresponding part of the physical environment visible via the display generation component, the user will see from this position in the physical environment in the same positioning, orientation and / or size (e.g., in an absolute sense and / or relative to each other) of the objects displayed in the three-dimensional environment by the display generation component of the computer system or visible in the three-dimensional environment via the display generation component. Similarly, if the virtual objects displayed in the three-dimensional environment are physical objects in the physical environment (e.g., physical objects placed at the same locations in the physical environment as the virtual objects are located in the three-dimensional environment, and physical objects that have the same size and orientation in the physical environment as they do in the three-dimensional environment), then the position of the computer system and / or user is the position from which the user would see the positions of the virtual objects in the physical environment at the same positions, orientations, and / or sizes (e.g., in an absolute sense and / or relative to each other and real-world objects) as the virtual objects displayed in the three-dimensional environment by the display generation components of the computer system.

[0218] In the present disclosure, various input methods are described with respect to interaction with a computer system. When an input device or input method is used to provide an example, and another input device or input method is used to provide another example, it should be understood that each example is compatible with the input device or input method described with respect to another example and optionally utilizes the input device or input method. Similarly, various output methods are described with respect to interaction with a computer system. When an output device or output method is used to provide an example, and another output device or output method is used to provide another example, it should be understood that each example is compatible with the output device or output method described with respect to another example and optionally utilizes the output device or output method. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment by a computer system. When an example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, it should be understood that each example is compatible with the method described with respect to another example and optionally utilizes these methods. Therefore, the present disclosure discloses an embodiment as a combination of features of multiple examples, without exhaustively listing all features of the embodiment in the description of each example embodiment.

[0219] User interface and associated processes

[0220] Attention is now focused on embodiments of a user interface ("UI") and associated processes that may be implemented on a computer system, such as a portable multifunction device or a head mounted device, in communication with a display generating component, one or more input devices, and optionally one or more cameras.

[0221] FIG. 7A to FIG. 7CH and 17A1 to Figure 17PAn illustration of a three-dimensional environment visible via a display generation component (e.g., display generation component 7100 or display generation component 120) of a computer system (e.g., computer system 101) and interactions occurring in the three-dimensional environment caused by user input directed to the three-dimensional environment and / or input received from other computer systems and / or sensors. In some embodiments, input is directed to a virtual object within the three-dimensional environment by a user gaze detected in an area occupied by the virtual object or by a gesture performed at a location in the physical environment corresponding to the area of ​​the virtual object. In some embodiments, input is directed to a virtual object within the three-dimensional environment by a gesture performed (e.g., optionally, at a location in the physical environment that is unrelated to the area of ​​the virtual object in the three-dimensional environment) when the virtual object has input focus (e.g., when the virtual object has been selected by concurrently and / or previously detected gaze input, by concurrently or previously detected pointer input, and / or by concurrently and / or previously detected gesture input). In some embodiments, input is directed to a virtual object within the three-dimensional environment by an input device that has positioned a focus selector object (e.g., a pointer object or a selector object) at the location of the virtual object. In some embodiments, input is directed to a virtual object within a three-dimensional environment via other means (e.g., voice and / or control buttons). In some embodiments, input is directed to a physical object or a representation of a virtual object corresponding to a physical object by user hand movement (e.g., whole hand movement, whole hand movement in a corresponding posture, movement of one part of a user's hand relative to another part of the hand, and / or relative movement between both hands) and / or manipulation relative to a physical object (e.g., touch, swipe, tap, open, move toward, and / or move relative to). In some embodiments, a computer system displays some changes to a three-dimensional environment (e.g., displaying additional virtual content, stopping displaying existing virtual content, and / or transitioning between different immersion levels of displayed visual content) based on input from sensors (e.g., image sensors, temperature sensors, biometric sensors, motion sensors, and / or proximity sensors) and contextual conditions (e.g., location, time, and / or the presence of other people in the environment). In some embodiments, a computer system displays some changes in a three-dimensional environment (e.g., displaying additional virtual content, ceasing to display existing virtual content, and / or transitioning between different immersive levels of displayed visual content) based on input from other computers used by other users sharing a computer-generated environment with a user of the computer system (e.g., in a shared computer-generated experience, in a shared virtual environment, and / or in a shared virtual or augmented reality environment of a communication session).In some embodiments, a computer system displays some changes in a three-dimensional environment (e.g., displaying movement, deformation, and / or changes in visual characteristics of a user interface, virtual surfaces, user interface objects, and / or virtual scenery) based on input from a sensor that detects movement of other people and objects as well as standard movement of a user that may not be considered standard gesture input that triggers an associated operation of the computer system.

[0222] In some embodiments, the three-dimensional environment visible via the display generation component described herein is a virtual three-dimensional environment including virtual objects and content at different virtual locations in the three-dimensional environment without a representation of the physical environment. In some embodiments, the three-dimensional environment is a mixed reality environment that displays virtual objects at different virtual locations in the three-dimensional environment that are constrained by one or more physical aspects of the physical environment (e.g., walls, floors, positioning and orientation of surfaces, gravity direction, time of day, and / or spatial relationships between physical objects). In some embodiments, the three-dimensional environment is an augmented reality environment that includes a representation of the physical environment. In some embodiments, the representation of the physical environment includes corresponding representations of physical objects and surfaces at different locations in the three-dimensional environment, so that the spatial relationship between different physical objects and surfaces in the physical environment is reflected by the spatial relationship between the representations of the physical objects and surfaces in the three-dimensional environment. In some embodiments, when a virtual object is placed relative to the positioning of the representations of the physical objects and surfaces in the three-dimensional environment, the virtual object appears to have a corresponding spatial relationship with the physical objects and surfaces in the physical environment. In some embodiments, a computer system transitions between displaying different types of environments based on user input and / or situational conditions (e.g., transitioning between presenting computer-generated environments or experiences with different levels of immersion, adjusting the relative prominence of audio / visual sensory input from virtual content and from representations of the physical environment).

[0223] In some embodiments, the display generation component includes a see-through portion in which a representation of the physical environment is displayed or visible. In some embodiments, the see-through portion of the display generation component is a transparent or translucent (e.g., see-through) portion of the display generation component that shows at least a portion of the physical environment around the user or within the user's field of view (sometimes referred to as "optical see-through"). For example, the see-through portion is a portion of a head-mounted display or a head-up display that is made translucent (e.g., less than 50%, 40%, 30%, 20%, 15%, 10% or 5% opacity) or transparent, so that the user can see through it to view the real world around the user without removing the head-mounted display or moving away from the head-up display. In some embodiments, when a virtual or mixed reality environment is displayed, the see-through portion gradually transitions from translucent or transparent to completely opaque. In some embodiments, the see-through portion of the display generating component displays a real-time feed of an image or video of at least a portion of the physical environment captured by one or more cameras (e.g., a rear-facing camera of a mobile device or associated with a head-mounted display, or other cameras that feed image data to a computer system) (sometimes referred to as "optical see-through"). In some embodiments, the one or more cameras are pointed at a portion of the physical environment that is directly in front of the user's eyes (e.g., behind the display generating component relative to the user of the display generating component). In some embodiments, the one or more cameras are pointed at a portion of the physical environment that is not directly in front of the user's eyes (e.g., in a different physical environment, or to the side or behind the user).

[0224] In some embodiments, when virtual objects are displayed at a location corresponding to the location of one or more physical objects in a physical environment (e.g., at a location in a virtual reality environment, a mixed reality environment, or an augmented reality environment), at least some of the virtual objects are displayed to replace (e.g., replace the display of) a portion of the camera's real-time view (e.g., a portion of the physical environment captured in the real-time view). In some embodiments, at least some of the virtual objects and content are projected onto a physical surface or blank space in the physical environment and are visible through a see-through portion of a display generation component (e.g., visible as part of the camera view of the physical environment, or visible through a transparent or translucent portion of the display generation component). In some embodiments, at least some of the virtual objects and content are displayed to cover a portion of the display and obstruct at least a portion of the view of the physical environment visible through a transparent or translucent portion of the display generation component.

[0225] In some embodiments, the display generation component displays different views of the three-dimensional environment based on user input or movement that changes the virtual positioning of the viewpoint of the currently displayed view of the three-dimensional environment relative to the three-dimensional environment. In some embodiments, when the three-dimensional environment is a virtual environment, the viewpoint moves based on navigation or motion requests (e.g., aerial gestures and / or gestures performed by movement of one part of the hand relative to another part of the hand) without requiring movement of the user's head, torso, and / or display generation component in the physical environment. In some embodiments, movement of the user's head and / or torso, and / or movement of the display generation component or other positioning sensing elements of the computer system (e.g., due to the user holding the display generation component or wearing an HMD) relative to the physical environment causes a corresponding movement of the viewpoint relative to the three-dimensional environment (e.g., with a corresponding movement direction, movement distance, movement speed, and / or orientation change), thereby causing a corresponding change in the current displayed view of the three-dimensional environment. In some embodiments, when a virtual object has a preset spatial relationship relative to a viewpoint (e.g., anchored or fixed to the viewpoint), movement of the viewpoint relative to the three-dimensional environment will cause movement of the virtual object relative to the three-dimensional environment while maintaining the position of the virtual object in the field of view (e.g., the virtual object is said to be head-locked). In some embodiments, the virtual object is body-locked to the user and moves relative to the three-dimensional environment as the user as a whole moves in the physical environment (e.g., carrying or wearing the display generation components and / or other position sensing components of the computer system), but will not move in the three-dimensional environment in response to individual user head movements (e.g., the display generation components and / or other position sensing components of the computer system rotate around a fixed position of the user in the physical environment). In some embodiments, the virtual object is optionally locked to another part of the user, such as the user's hand or the user's wrist, and moves in the three-dimensional environment in accordance with the movement of that part of the user in the physical environment to maintain a preset spatial relationship between the position of the virtual object and the virtual position of that part of the user in the three-dimensional environment. In some embodiments, the virtual object is locked to a preset portion of the field of view provided by the display generation component and moves in the three-dimensional environment based on the movement of the field of view, regardless of the user's movement that does not cause the field of view to change.

[0226] In some embodiments, such as FIG. 7A to FIG. 7CH as well as FIG. 17A2 to FIG. 17PAs shown, the view of the three-dimensional environment sometimes does not include the representation of the user's hand, arm and / or wrist. In some embodiments, the representation of the user's hand, arm and / or wrist is included in the view of the three-dimensional environment. In some embodiments, the representation of the user's hand, arm and / or wrist is included in the view of the three-dimensional environment as a part of the representation of the physical environment provided via the display generation component. In some embodiments, these representations are not part of the representation of the physical environment, and are captured separately (e.g., by one or more cameras pointing to the user's hand, arm and wrist) and are displayed in the three-dimensional environment independently of the view currently displayed of the three-dimensional environment. In some embodiments, these representations include camera images captured by one or more cameras of the computer system or stylized versions of arms, wrists and / or hands based on information captured by various sensors. In some embodiments, these representations replace the display of a part of the representation of the physical environment, overlay on this part of the representation of the physical environment or block the view of this part of the representation of the physical environment. In some embodiments, when the display generation component does not provide a view of the physical environment and provides a completely virtual environment (e.g., no camera view and no transparent pass-through portion), a real-time visual representation of one or both arms, wrists, and / or hands of the user is optionally still displayed in the virtual environment (e.g., a stylized representation or a segmented camera image). In some embodiments, if no representation of the user's hands is provided in the view of the three-dimensional environment, the position corresponding to the user's hands is optionally indicated in the three-dimensional environment, for example, by changing the appearance of the virtual content at the position in the three-dimensional environment corresponding to the position of the user's hands in the physical environment (e.g., by a change in translucency and / or simulated reflectivity). In some embodiments, the representation of the user's hand or wrist is outside the currently displayed view of the three-dimensional environment, and the virtual position corresponding to the position of the user's hand or wrist in the three-dimensional environment is outside the current field of view provided by the display generation component; and in response to the virtual position corresponding to the position of the user's hand or wrist moving within the current field of view due to the movement of the display generation component, the user's hand or wrist, the user's head, and / or the user as a whole, the representation of the user's hand or wrist is made visible in the view of the three-dimensional environment.

[0227] FIG. 7A to FIG. 7R Examples of displaying window controls for virtual objects (eg, application windows and / or three-dimensional objects) are illustrated. Figure 8 is a flow chart of an exemplary method 800 for conditionally displaying a window control. Fig. 9 is a flow chart of an exemplary method 900 for updating visual properties of a window control in response to user interaction. FIG. 7A to FIG. 7R The user interface in is used to illustrate the process described below, including Figure 8 and Fig. 9 process.

[0228] Fig. 7A A view of a physical environment including a user 7002 interacting with a display generation component 7100 is illustrated. In the examples described below, the user 7002 uses one or both of his or her hands (hand 7020 and hand 7022) to provide input or instructions to the computer system. In some examples described below, the computer system also uses the position or movement of the user's arm (such as the user's left arm 7028 connected to the user's left hand 7020) as part of the input provided by the user to the computer system. The physical environment 7000 includes a physical object 7014 and physical walls 7004 and 7006. The physical environment 7000 also includes a physical floor 7008.

[0229] like FIG. 7A to FIG. 7CH and FIG. 17A2 to FIG. 17P As shown in the example in FIG. 1 , the display generating component 7100 of the computer system 101 is a touch screen held by the user 7002. In some embodiments, the display generating component of the computer system 101 is a head mounted display (e.g., head mounted display 7100a, such as FIG. 100b ) worn on the head of the user 7002. FIG. 7F2 to FIG. 7F3 , Figure 7K1 to Figure 7K2 , Figure 7T2 to Figure 7T3 , Figure 7AD2 to Figure 7AD3 , Figure 7AN2-7AN3 , Figure 7AU2 to Figure 7AU3 , Figure 7BA2 to Figure 7BA3 , Figure 7BD2 to Figure 7BD3 , Figure 7BM2 to Figure 7BN2 and FIG. 17A1 to FIG. 17B1 ) (e.g., in FIG. 7A to FIG. 7CH and FIG. 17A2 to FIG. 17PThe content shown as visible via the display generation component 7100 of the computer system 101 corresponds to the field of view of the user 7002 when wearing the head-mounted display). In some embodiments, the display generation component is a stand-alone display, a projector, or another type of display. In some embodiments, the computer system communicates with one or more input devices, which include cameras or other sensors and input devices that detect movement of a user's hands, movement of the user's entire body, and / or movement of the user's head in the physical environment. In some embodiments, the one or more input devices detect movement and current posture, orientation, and position of the user's hands, face, and / or entire body. For example, in some embodiments, when a user's hand 7020 is within the field of view of one or more sensors of the HMD 7100a (e.g., within the user's field of view), a representation of the user's hand 7020 is displayed in a user interface displayed on the display of the HMD 7100a (e.g., as a transparent representation and / or virtual representation of the user's hand 7020). In some embodiments, when the user's hand 7022 is within the field of view of one or more sensors of the HMD 7100a (e.g., within the user's field of view), a representation of the user's hand 7022' is displayed in a user interface displayed on the display of the HMD 7100a (e.g., as a pass-through representation and / or a virtual representation of the user's hand 7022). In some embodiments, the user's hand 7020 and / or the user's hand 7022 are used to perform one or more gestures (e.g., one or more air gestures), optionally in combination with gaze input. In some embodiments, the one or more gestures performed with the user's hand 7020 and / or 7022 include direct air gesture input, which is based on the positioning of the representation of the user's hand 7020' and / or 7022' displayed within the user interface on the display of the HMD 7100a. For example, direct air gesture input is determined to be directed to a user interface object displayed at a location that intersects with the displayed location of the representation of the user's hand 7020' and / or 7022' in the user interface. In some embodiments, the one or more gestures performed with the user's hand 7020 and / or 7022 include indirect air gesture input, which is based on a virtual object displayed at a location corresponding to the location where the user's attention is currently detected (e.g., and / or optionally not based on the location of the representation of the user's hand 7020' and / or 7022' displayed within the user interface). For example, an indirect air gesture, such as gaze and pinch (e.g., or other gestures performed with the user's hand), is performed relative to the user interface object when the user's attention is detected on the user interface object (e.g., based on gaze or other indication of user attention).

[0230] In some embodiments, user input is detected via a touch-sensitive surface or a touch screen. In some embodiments, one or more input devices include an eye tracking component that detects the position and movement of the user's gaze. In some embodiments, the display generation component and optionally one or more input devices and a computer system are part of a head-mounted device that moves and rotates with the user's head in a physical environment and changes the user's viewpoint in a three-dimensional environment provided by the display generation component. In some embodiments, the display generation component is a head-up display that does not move or rotate with the user's head or the user's entire body, but optionally changes the user's viewpoint in a three-dimensional environment based on the movement of the user's head or body relative to the display generation component. In some embodiments, the display generation component (e.g., a touch screen) is optionally moved and rotated by the user's hand relative to the physical environment or relative to the user's head, and changes the user's viewpoint in a three-dimensional environment based on the movement of the display generation component relative to the user's head or face or relative to the physical environment.

[0231] In some embodiments, the display generation component 7100 is a head mounted display (HMD) 7100a. For example, Figure 7F2 (For example, and Figure 7K1 , Figure 7T2 , Figure 7AD2 , Figure 7AN2 , Figure 7AU2 , Figure 7BA2 , Figure 7BD2 and Figure 7BM2 to Figure 7BN2), the head mounted display 7100a includes one or more displays that display a representation of a portion of the three-dimensional environment 7000' corresponding to the user's perspective, while the HMD typically includes multiple displays, including a display for the right eye and a separate display for the left eye that display slightly different images to generate a user interface with stereoscopic depth, a single image corresponding to the image for a single eye is shown in these figures and the depth information is indicated with other annotations or descriptions of these figures. In some embodiments, the HMD 7100a includes one or more sensors (e.g., one or more inward-facing and / or outward-facing image sensors 314), such as sensor 7101a, sensor 7101b, and / or sensor 7101c, for detecting the user's state, including tracking of the user's face and / or eyes (e.g., using one or more inward-facing sensors 7101a and / or 7101b) and / or tracking of the user's hand, torso, or other movements (e.g., using one or more outward-facing sensors 7101c). In some embodiments, the HMD 7100a includes one or more input devices, such as one or more buttons, a touchpad, a touch screen, a scroll wheel, a rotatable and depressible digital crown, or other input devices, optionally located on the housing of the HMD 7100a. In some embodiments, the input element is a mechanical input element, and in some embodiments, the input element is a solid-state input element that responds to a press input based on detected pressure or intensity. For example, in Figure 7F2 (For example, and Figure 7K1 , Figure 7T2 , Figure 7AD2 , Figure 7AN2 , Figure 7AU2 , Figure 7BA2 , Figure 7BD2 and Figure 7BM2 to Figure 7BN2 ), the HMD 7100a includes one or more of a button 701a, a button 701b, and a digital crown 703 for providing input to the HMD 7100a. It should be understood that additional and / or alternative input devices may be included in the HMD 7100a.

[0232] Figure 7F3 (For example, and Figure 7K2 , Figure 7T3 , Figure 7AD3 , Figure 7AN3 , Figure 7AU3 , Figure 7BA3 and Figure 7BD3 ) illustrates an overhead view of a user 7002 in a physical environment 7000. For example, the user 7002 is wearing an HMD 7100a such that the user's hands 7020 and / or 7022 (e.g., which are optionally used to provide mid-air gestures or other user input) are physically present within the physical environment 7000 behind the display of the HMD 7100a.

[0233] Figure 7F2 (For example, and Figure 7K1 , Figure 7T2 , Figure 7AD2 , Figure 7AN2 , Figure 7AU2 , Figure 7BA2 , Figure 7BD2 and Figure 7BM2 to Figure 7BN2 ) illustrates an alternative display generation component of a computer system, rather than 7A to 7E , Figure 7G to Figure 7J , FIG. 7L to FIG. 7T1 , Figure 7U to Figure 7AD1 , Figure 7AE to Figure 7AN1 , Figure 7AO to Figure 7U 1. Figure 7V to Figure 7BA1 , Figure 7BC to Figure 7BD1 , Figures 7BL to 7N 1 and Figures 7BO to 7CH It should be understood that the display illustrated in this document is 7A to 7E , Figure 7G to Figure 7J , FIG. 7L to FIG. 7T1 , Figure 7U to Figure 7AD1 , Figure 7AE to Figure 7AN1 , Figure 7AO to Figure 7U 1. Figure 7V to Figure 7BA1 , Figure 7BC to Figure 7BD1 , Figures 7BL to 7N 1 and Figures 7BO to 7CH The processes, features and functions described in the display generation component 7100 are also applicable to FIG. 7F2 to FIG. 7F3 , Figure 7K1 to Figure 7K2 , Figure 7T2 to Figure 7T3 , Figure 7AD2 to Figure 7AD3 , Figure 7AN2 to Figure 7AN3 , Figure 7AU2 to Figure 7AU3 , Figure 7BA2 to Figure 7BA3 , Figure 7BD2 to Figure 7BD3 and 7BM2 to Figure 7BN2 Illustrated HMD 7100a.

[0234] like Figure 7BAs shown, a computer system (e.g., display generation component 7100) displays a view of a three-dimensional environment (e.g., environment 7000', a virtual three-dimensional environment, an augmented reality environment, a pass-through view of a physical environment, or a camera view of a physical environment). In some embodiments, the three-dimensional environment is a virtual three-dimensional environment without a representation of the physical environment 7000. In some embodiments, the three-dimensional environment is a mixed reality environment, which is a virtual environment enhanced by sensor data corresponding to the physical environment. In some embodiments, the three-dimensional environment is an augmented reality environment, which includes one or more virtual objects (e.g., application window 702 and / or virtual object 7028) and a representation of at least a portion of the physical environment around the display generation component 7100 (e.g., representations 7004, 7006' of walls in the three-dimensional environment 7000, representations 7008' of floors, and / or representations 7014' of physical objects 7014). In some embodiments, the representation of the physical environment includes a camera view of the physical environment. In some embodiments, the representation of the physical environment includes a view of the physical environment through a transparent or translucent portion of the first display generation component.

[0235] In some embodiments, application window 702 is displayed at a first location in a first view of three-dimensional environment 7000'. In some embodiments, application window 702 is associated with a first application executed on a computer system. For example, application window 702 displays the content of the first application. In some embodiments, application window 702 is displayed as having a first horizontal location, a first vertical location, and a first depth or a perceived distance from the user (e.g., a location defined by an x-axis, a y-axis, and a z-axis) within a first view of three-dimensional environment 7000'. In some embodiments, application window 702 is locked (also referred to herein as anchoring) to the three-dimensional environment so that when the field of view of the three-dimensional environment changes, application window 702 maintains its location within the three-dimensional environment.

[0236] In an embodiment where the display generation component 7100 of the computer system 101 is a head mounted display, the application window 702 will be displayed in a peripheral area of ​​the field of view of the user's eyes when looking at the three-dimensional environment via the display generation component.

[0237] In some embodiments, the user is enabled to move the location of the application window 702 to place it in a different location in the three-dimensional environment 7000' so that the application window 702 becomes locked to the new location in the three-dimensional environment. For example, the grabber 706-1 is a selectable user interface object for the application window 702 that enables the user to reposition the application window 702 within the three-dimensional environment 7000 when selected by the user (e.g., using gaze and / or gestures, such as air gestures). In some embodiments, the grabber 706-1 is displayed along the bottom center edge of the application window 702. In some embodiments, the grabber 706-1 is displayed at different locations relative to the application window 702. In some embodiments, the shape and / or size of the grabber bar changes based on the size of the application window 702. For example, the size of the grabber 706-1 increases and / or decreases as the size of the application window 702 increases and / or decreases. In some embodiments, the application window 702 is a two-dimensional object (e.g., from the user's point of view, the application window 702 appears to be flat).

[0238] In some embodiments, when application window 702 is displayed in a three-dimensional environment, grabber 706-1 is automatically and without user input displayed with application window 702. In some embodiments, grabber 706-1 is displayed only when the user's attention is directed to application window 702, and disappears in response to the user's attention being moved away from application window 702. In some embodiments, grabber 706-1 is displayed in response to detecting that the user's gaze is on the bottom center portion or other predefined portion of application window 702.

[0239] Figure 7B Also illustrated is a virtual object 7028 that optionally does not correspond to a physical object in the physical environment. In some embodiments, the virtual object 7028 is a three-dimensional object, such as a ball. In some embodiments, the virtual object 7028 is associated with a second application that is different from the first application associated with the application window 702. In some embodiments, the first application associated with the application window 702 and / or the second application associated with the virtual object 7028 is a system application of the computer system (e.g., associated with an operating system) or an application associated with a third party executed by the computer system.

[0240] Figure 7CThe computer system is illustrated to detect the user's attention 710-1 (e.g., the user's gaze) directed to the upper left corner or the area around the upper left corner (e.g., within a predefined area overlapping the upper left corner) of the application window 702. In some embodiments, the user's attention 710-1 is detected as a gaze input. In some embodiments, the display generation component 7100 optionally displays an indication (e.g., a cursor or other user interface object) corresponding to the detected user gaze input, so that the indication moves within the display area of ​​the display generation component 7100 according to the movement of the user's gaze. In some embodiments, the user's attention 710-1 is detected as another type of input, such as an air gesture. In some embodiments, the user's attention 710-1 indicates the positioning of a cursor (e.g., or other visual indicator) associated with an input device (e.g., controlled by the user's hand rather than the user's gaze). In some embodiments, based on determining that the user's attention is directed to at least a portion of the application window 702, the application window 702 is displayed in front of the virtual object 7028 (e.g., the application window is displayed as being closer to the user than the virtual object). For example, the computer system optionally visually weakens the virtual object 7028 by dimming the virtual object 7028, pushing the virtual object 7028 back within the three-dimensional environment 7000' (e.g., further away from the viewpoint of the user 7002), and / or reducing the size of the virtual object 7028. For example, the computer system automatically changes the perceived depth of the corresponding object based on the detected user's attention.

[0241] In some embodiments, in response to detecting the user's attention 710-1 directed to the upper left corner of the application window 702, the computer system displays a close enable indication 7030 for closing the application window 702, as shown in FIG. 7D (e.g., FIG. 7D1 to FIG. 7D3). In some embodiments, computer system 101 displays close enable indication 7030 after detecting that the user's attention 710-1 has been maintained for a threshold amount of time and / or other attention criteria have been met. For example, in response to detecting a gaze input directed to the upper left corner that has not been maintained for a threshold amount of time (e.g., the user gazes elsewhere in the three-dimensional environment before the threshold amount of time is met), the computer system does not display close enable indication 7030. In some embodiments, close enable indication 7030 is displayed as a user interface element different from application window 702. For example, close enable indication 7030 is separated from application window 702 by a non-zero distance, so that a portion of three-dimensional environment 7000' between application window 702 and close enable indication 7030 is visible. In some embodiments, close enable indication 7030 is displayed above application window 702, near the corner of application window 702. In some embodiments, when close enable indication 7030 is displayed, the computer system optionally maintains the display of grabber 706-1. In some embodiments, when the off enable indication 7030 is displayed, the computer system stops display of the gripper 706-1.

[0242] In some embodiments, such as Fig.7D1 As illustrated, a title bar 716a is displayed below the application window 702, optionally in response to detecting that the user's attention is directed to the bottom portion of the application window 702. In some embodiments, the user interface object 705 is associated with the grabber 706-1 and / or the title bar 716a (e.g., it has the same or similar functionality as the title bar 716, as described in reference to FIG. Figure 7S In some embodiments, user interface object 705 is displayed simultaneously with title bar 716a and / or grabber 706-1 (e.g., to the left and / or right of the title bar and / or grabber). In some embodiments, user interface object 705 corresponds to a minimized version of a close icon and / or minimized version of one or more other controls (e.g., a control menu or another type of control object). In some embodiments, user interface object 705 is displayed before the user's attention is detected directed to the bottom portion of application window 702 (e.g., and optionally, continues to be displayed after the user's attention is diverted from the bottom portion of application window 702).

[0243] In some embodiments, in response to detecting that the user's attention 710-1a is directed to the user interface object 705, the user interface object 705 is updated (e.g., from a minimized state or a reduced state) to display a close enable indication 7030-2, such as Fig.7D2In some embodiments, close enable indication 7030-2 has the same or similar functionality as close enable indication 7030, but close enable indication 7030-2 is optionally displayed at a different location relative to application window 702 than close enable indication 7030 is.

[0244] In some embodiments, when close enable indication 7030 (or close enable indication 7030-2) is displayed, the computer system optionally detects pointing to close enable indication 7030 (or close enable indication 7030-2, such as Fig.7D2 In some embodiments, in response to detecting user input directed toward close affordance 7030 or 7030-2, the computer system stops display of application window 702. Thus, the user is enabled to close application window 702 by selecting close affordance 7030 or 7030-2.

[0245] Figure 7D3 7F (e.g., FIG. 7F ) illustrates that computer system 101 detects user attention 710-2 directed to the lower right corner of application window 702. In some embodiments, in response to detecting user attention 710-2 directed to the lower right corner of application window 702, the computer system displays a resize affordance 708-1 corresponding to the lower right corner (e.g., FIG. 7F (e.g., FIG. 7F )). Figure 7F1 and Figure 7F2 )The resizing indicator shown represents 708-1).

[0246] In some embodiments, before displaying the resize enable representation 708-1, the computer system enables the user to access the functionality associated with the resize enable representation 708-1 and / or perform operations associated with the resize enable representation. For example, without displaying the resize enable representation 708-1, the user is enabled to perform a gesture or other user input to adjust the size of the application window 702 when the user's attention is directed to the lower right corner of the application window 702 (for example, and in response to detecting the gesture or other user input, the size of the application window 702 is adjusted according to the gesture or other user input). In some embodiments, the resize enable representation 708-1 is displayed during and / or after the user performs the gesture. In some embodiments, the resize enable representation 708-1 is displayed in response to detecting that the user's attention 710-2 meets the attention standard. For example, the user maintains the user's gaze at the lower right corner for a threshold amount of time (for example, 1 second, 2 seconds, 5 seconds, or another amount of time). In some embodiments, the resize enable representation 708-1 is displayed based on the size and / or shape of the application window 702. For example, in some embodiments, the size of resizing enable representation 708-1 is based on the size of application window 702. In some embodiments, resizing enable representation 708-1 is displayed as having an L-shape around the corner of application window 702 (e.g., to extend along a portion of the bottom edge and a portion of the right edge of application window 702), where the outline of the L-shape follows the outline of the corner of application window 702.

[0247] In some embodiments, the user's attention 710-2 satisfies an attention criterion including a criterion that is satisfied when the user's attention 710-2 is directed to a first area having a first size corresponding to a corresponding portion of the application window 702. For example, the user's attention 710-2 is directed to an area having a first size centered on the lower right corner of the application window 702.

[0248] In some embodiments, in response to detecting the user's attention 710-2 directed to the lower right corner of the application window 702, the computer system displays an animated transition from the display grabber 706-1 and to the resize affordance 708-1. For example, the animated transition includes the display grabber 706-1 gradually shifting to the right, such as Fig. 7EFor example, until the gripper is replaced by the resize-enabled representation 708-1 displayed in the lower right corner of the application window 702. In some embodiments, the gripper 706-1 is displayed as a single bar of the resize-enabled representation 708-1 deformed into an L-shape. In some embodiments, the animated transition includes optionally reducing the size of the gripper 706-1 and / or fading the gripper 706-1 without shifting the positioning of the gripper 706-1 until the gripper 706-1 is no longer displayed, and optionally increasing the size of the resize-enabled representation 708-1 and / or fading in the resize-enabled representation 708-1 while or after reducing the size of the gripper 706-1. For example, the animated transition removes the gripper 706-1 and initiates the display of the resize-enabled representation 708-1 at the corner where the user's gaze is detected. In some embodiments, the gripper 706-1 is not displayed when the resize-enabled representation 708-1 is displayed.

[0249] In some embodiments, resize enable representation 708-1 is displayed based on determining that the user's attention 710-2 is directed to the lower right corner of the application window for at least a threshold amount of time. For example, display of resize enable representation 708-1 is delayed until the threshold amount of time is reached (e.g., and optionally the user is enabled to resize application window 702 before resize enable representation 708-1 is displayed).

[0250] In some embodiments, the animated transition between displaying the grabber 706-1 and displaying the resize enable representation 708-1 is an example of an animation displayed for displaying object management controls, which include resize enable representations, grabbers, close enable representations, title bars, and / or other enable representations that are dynamically displayed in response to detecting that the user's attention is directed to a corresponding portion of the application window 702 (e.g., or other virtual objects). For example, the enable representations described herein respond to detecting the user's attention (such as, the user's gaze) so that the enable representations are displayed based on determining that the user's attention is directed to a portion of the displayed area that corresponds to the enable representation (e.g., indicating that the user intends to interact with the enable representation). Therefore, in some embodiments, in response to detecting that the user's attention is directed to a corresponding portion of the application window 702, an animation of the corresponding enable representation of the corresponding portion of the application window 702 is initiated, and the user is enabled to perform operations associated with the enable representation, regardless of whether the animation is completed or not. For example, while an animation is in progress (e.g., or before initiating the animation), as long as the user's attention has remained at the corresponding location corresponding to the enable representation for a threshold amount of time, the user is enabled to select the corresponding enable representation or otherwise perform a corresponding operation associated with the corresponding enable representation (e.g., even before the corresponding enable representation is displayed). For example, an animation is initiated to display the corresponding enable representation after the user's attention has remained and a threshold amount of time has passed, but the user is enabled to interact with the enable representation before the enable representation is displayed (e.g., by directing the user's attention to the location corresponding to the location of the enable representation).

[0251] In some embodiments, the application window 702 (e.g., or other virtual objects, such as three-dimensional virtual objects) is divided into multiple areas, such as a left edge area, a left corner area, a bottom area, a right corner area, and a right edge area. In some embodiments, each of the multiple areas optionally includes an area outside the application window 702 and / or close to the application window. For example, the left corner area includes an area extending beyond the left corner of the application window 702. In some embodiments, the corresponding affordance is enabled to appear in any of these areas (e.g., the same affordance appears in any of these areas, or based on the area, different affordances appear in different areas). For example, based on which corner area of ​​these corner areas the user's attention is currently directed to, the resize affordance 708-1 appears in the lower left corner area of ​​the application window 702, the lower right corner area of ​​the application window 702, the upper left corner area of ​​the application window 702, and / or the upper right corner area of ​​the application window 702.

[0252] In some embodiments, the system determines a current state of each of the multiple regions and optionally performs an operation (e.g., and / or enables an operation to be performed in response to user input) based on the current state (e.g., and / or a change to the current state of the corresponding region). For example, possible states include: the enable indication is invisible and does not allow user interaction, the enable indication is invisible but does allow interaction to perform an operation associated with the enable indication, the enable indication is displayed but the user's attention does not remain in the region (e.g., the user's attention is detected as being directed to the region for a time less than a threshold amount of time), the enable indication is displayed and the user's attention remains in the region (e.g., the user's attention is detected as being directed to the region for at least a threshold amount of time), and the enable indication is displayed and selected (e.g., pressed or otherwise interacted with).

[0253] In some embodiments, the title bar and / or other enable representations are visible only in one area of ​​the multiple areas (e.g., not visible or available in other areas of the multiple areas). For example, in a first area of ​​the multiple areas, a user is enabled to perform an operation associated with a first corresponding enable representation (e.g., even if the first corresponding enable representation is inactive or not displayed), and / or a user is enabled to direct the user's attention to the first corresponding enable representation in the first area and / or select the first corresponding enable representation, while in a second area of ​​the multiple areas, the first corresponding enable representation is hidden and / or disabled (e.g., so that the user cannot select the first corresponding enable representation or interact with the first corresponding enable representation). In some embodiments, in response to detecting user input directed to a first area including a first corresponding enable representation (e.g., the user's attention directed to the first area) (e.g., the user's gaze is directed to the corresponding area for at least a threshold amount of time), the computer system provides visual feedback (e.g., a change in opacity, blurring, and / or other visual feedback) in the area where the enable representation is currently displayed.

[0254] In some embodiments, in response to detecting that the user's attention is directed to a third area (e.g., an area other than the first area) of the multiple areas, where the first corresponding enable representation is not displayed in the third area, the computer system displays the following: an animation showing a second corresponding enable representation (e.g., an enable representation that is the same as or different from the first corresponding enable representation) of the third area to which the user is currently directing the user's attention. For example, as referenced Fig. 7E to FIG. 7F (eg, Figure 7F1 , Figure 7F2 and Figure 7F3 ) as described above, the grab bar 706-1 is animated into a resizing indicator 708-1 in response to detecting that the user's attention is directed to the lower right corner area.

[0255] In some embodiments, when the user's attention is detected pointing to the lower right corner of the application window 702, the resize enable representation 708-1 continues to be displayed. In some embodiments, in response to detecting that the user's attention is no longer directed to the lower right corner of the application window 702, the resize enable representation 708-1 is no longer displayed, and the grabber 706-1 is optionally redisplayed. In some embodiments, detecting that the user's attention is no longer directed to the lower right corner of the application window 702 includes determining that the user's attention is directed outside a second area of ​​a second size (e.g., different from a first area of ​​a first size, the first area corresponding to the corresponding portion of the application window 702 used to determine that the user's attention meets the attention standard). For example, the second area of ​​the second size is a larger area than the first area of ​​the first size. In some embodiments, the second area completely covers the first area. Therefore, detecting that the user's attention is no longer directed to the lower right corner is based on whether the user's attention has moved to an area larger than the first area to which the user's attention is directed, which is used to determine that the user's attention meets the attention standard (e.g., and the resize enable representation is displayed based on determining that the user's attention meets the attention standard).

[0256] In some embodiments, in response to detecting that the user's attention is directed to the lower left corner, a resize-enabled representation (e.g., similar to resize-enabled representation 708-1) is displayed near the lower left corner. For example, a mirror image of resize-enabled representation 708-1 is displayed in the lower left corner. Therefore, based on detecting which of the bottom corners the user's attention is directed to, the computer system displays a corresponding resize-enabled representation at the corresponding corner (e.g., the lower right corner and / or the lower left corner).

[0257] In some embodiments, in response to detecting that the user's attention is directed to another area of ​​the application window 702 other than the lower right corner and / or the lower left corner, the computer system displays a corresponding enable representation and / or abandons displaying a corresponding enable representation that does not correspond to the current location to which the user's attention is directed. For example, detecting that the user's attention is directed to the upper left corner of the application window 702 causes the computer system to stop resizing the display of the enable representation 708-1 and display the close enable representation 7030, and optionally display (e.g., or redisplay) the grabber 706-1. It should be understood that different enable representations are associated with corresponding portions of the application window 702 (and / or virtual object 7028), so that the user invokes display of the corresponding enable representation by directing the user's attention to the corresponding portion of the application window 702 associated with the corresponding enable representation. Although the examples described herein associate the bottom corner of application window 702 with a resize affordance and the upper left corner of application window 702 with a close affordance, it should be understood that these corners may be assigned to different types of affordances based on the application windows (e.g., different applications may associate different affordances with these corners). For example, some application windows and / or virtual objects cannot be resized, and resize affordances are not displayed in response to a user looking at a corner of an application window and / or virtual object. In some embodiments, one or more application windows and / or virtual objects cannot be repositioned within a three-dimensional object, and a grabber is not displayed for the one or more application windows and / or virtual objects.

[0258] FIG. 7F (eg, Figure 7F1 , Figure 7F2 and Figure 7F3 ) illustrates that computer system 101 detects user attention 710-4 directed toward resize enable representation 708-1. In some embodiments, in response to detecting that the user's attention is directed toward resize enable representation 708-1 (e.g., detecting that the user is looking at resize enable representation 708-1), the computer system updates one or more visual attributes of resize enable representation 708-1 to indicate that the computer system detected the user's attention directed toward resize enable representation 708-1. For example, Figure 7G Resizing enable representation 708-2 (e.g., an updated version of resizing enable representation 708-1) is illustrated as being displayed in a color different from the color of resizing enable representation 708-1 in FIG. 7F. In some embodiments, updating the one or more visual attributes of resizing enable representation 708-1 includes displaying (updated) resizing enable representation 708-2 in a different size, color, and / or transparency when the user's attention 710-4 is detected as being directed toward resizing enable representation 708-2.

[0259] In some embodiments, after updating the one or more visual attributes of resizable affordance 708-1, and optionally before detecting additional user input (e.g., a direct air gesture such as an air tap or air pinch at the location where the user is interacting with it, an indirect air gesture such as an air pinch when the user's attention or the user's gaze is directed to the location where the user is interacting with it, a tap input, a gaze input, a drag input, and / or another type of selection input) selecting the (updated) resizable affordance 708-2, the computer system detects that the user's attention is directed to another portion of the three-dimensional environment that does not correspond to resizable affordance 708-2. For example, Figure 7H As illustrated, the user's attention 710-5 shifts to the left side of resizing affordance 708-3 (e.g., which is similar to resizing affordance 708-2, but has different properties as described below). In some embodiments, in response to detecting that the user's attention is no longer directed toward resizing affordance 708-2 (e.g., and / or detecting that the user's attention is directed toward another portion of the three-dimensional environment), the computer system reverses the update of one or more visual properties of the resizing affordance, such as Figure 7H For example, instead of maintaining the updated visual attributes of resize-enabling representation 708-2, resize-enabling representation 708-3 is displayed with the same visual attributes as resize-enabling representation 708-1 (e.g., displayed with the same appearance as resize-enabling representation 708-1 in FIG. 7F ).

[0260] In some embodiments, in response to detecting that the user's attention is not directed toward resizing enable indication 708-3, the computer system optionally maintains updates to the one or more visual attributes of resizing enable indication 708-2 (e.g., resizing enable indication 708-3 has the same appearance as resizing enable indication 708-2), depending on where the user's attention is directed (e.g., whether the user's attention is directed outside the proximity range of resizing enable indication 708-3). For example, if the user moves their gaze away from application window 702, resizing enable indication 708-3 ceases to be displayed, and if the user looks at another part of application window 702 (e.g., but not at resizing enable indication 708-3), the appearance of resizing enable indication 708-3 is optionally maintained (e.g., it has a different size and / or color than resizing enable indication 708-2). In some embodiments, resizing enable indication 708-3 continues to be displayed, but is displayed with a different visual appearance (e.g., the color and / or size is changed, and it is different from resizing enable indication 708-2). Figure 7GThus, depending on where the user's gaze is detected, resizable affordance 708-3 ceases to be displayed or is maintained, optionally with different visual attributes.

[0261] In some embodiments, computer system 101 detects that the user's attention is redirected to resize-enabled representation 708-3 (e.g., after being redirected therefrom) before a threshold amount of time has elapsed. For example, the user has quickly looked away from resize-enabled representation 708-3 and then returned to looking at resize-enabled representation 708-3, and the computer system maintains display of resize-enabled representation 708-3 (e.g., for a threshold amount of time).

[0262] In some embodiments, computer system 101 detects that the user's attention is directed to another portion of the three-dimensional environment that does not correspond to resize-enabled representation 708-3 for a threshold amount of time, and after the threshold amount of time has passed, the computer system stops displaying resize-enabled representation 708-3 and optionally redisplays grabber 706-1, such as Figure 7J In some embodiments, the computer system stops display of resize-enabled representation 708-3 by displaying an animated transition (such as reducing the size of resize-enabled representation 708-3 and / or fading resize-enabled representation 708-3 until it is no longer displayed). In some embodiments, the animated transition includes shifting the resize-enabled representation downward and / or around the corresponding corner of application window 702 (optionally in a direction toward grabber 706-1). For example, in Fig.7I , the resizing affordance 708-4 is shifted downward and to the left, as if evolving and / or morphing into Figure 7J In some embodiments, in response to ceasing display of resize affordance 708-3, grabber 706-1 is optionally automatically redisplayed without requiring additional user input (e.g., without requiring the user to gaze at the bottom center edge of application window 702). In some embodiments, in response to detecting that user attention 710-6 is directed toward the bottom center edge of application window 702 (e.g., including an area along the bottom edge of application window 702 outside of application window 702, such as Fig.7I In some embodiments, in response to detecting the user's attention 710-6 directed to the bottom center edge of the application window 702, the computer system 101 displays an animated transition of the mobile resize affordance 708-4 until it is displayed as the grabber 706-1.

[0263] Figure 7JDetecting the user's attention 710-7 directed to the lower right corner of the application window 702 is illustrated. In some embodiments, in response to detecting the user's attention 710-7, the computer system 101 displays (e.g., or redisplays) the resize enable representation 708-1, and in response to detecting that the user's attention is directed to the resize enable representation 708-1, updates the resize enable representation in a first manner (e.g., updates it to the resize enable representation 708-3), such as by changing the color of the resize enable representation (e.g., as described above with reference to Figures 7F to 7F). Figure 7G described above).

[0264] In some embodiments, after the resizable enable representation is updated in a first manner, when the resizable enable representation 708-3 is displayed with the updated one or more visual attributes (e.g., as in Figure 7H ), computer system 101 detects user input directed to resize affordance 708-3 (e.g., a direct air gesture such as an air tap or air pinch at a location where the user is interacting with it, an indirect air gesture such as an air pinch when the user's attention or the user's gaze is directed at the location where the user is interacting with it, a tap input, a gaze input, a drag input, and / or another type of user input) (e.g., as indicated by FIG. 7K (e.g., Figure 7K1 , Figure 7K2 and Figure 7K3 ) in the image). In some embodiments, the user input directed to resizing enable representation 708-3 is an air gesture detected when the user's gaze is detected as directed to resizing enable representation 708-3, such as a pinch gesture or a tap input. In some embodiments, in response to the user input directed to resizing enable representation 708-3, the computer system updates the display of resizing enable representation 708-3 in a second manner (e.g., which is different from the first manner). For example, in response to the pinch gesture, the computer system changes the size of resizing enable representation 708-3 to a (e.g., smaller) size of resizing enable representation 708-5, optionally while maintaining the updated color of the resizing enable representation, as shown in FIG. 7K (e.g., Figure 7K1 , Figure 7K2 and Figure 7K3). Thus, when a user interacts with a resizable enable representation, the computer system provides two levels of visual feedback to the user. For example, the computer system first changes the color of the resizable enable representation to indicate that the computer system detects that the user's attention is on the resizable enable representation, and upon detecting further user input interacting with the resizable enable representation (e.g., direct air gestures such as an air tap or air pinch at the location where the user is interacting with it, indirect air gestures such as an air pinch when the user's attention or the user's gaze is directed to the location where the user is interacting with it, tap input, gaze input, drag input, and / or another type of user input), the computer system changes the size of the resizable enable representation to indicate that the resizable enable representation has been selected by the user input.

[0265] In some embodiments, changing the size of resize-enabled representation 708-3 to the size of resize-enabled representation 708-5 includes changing the width or thickness of the resize-enabled representation without changing the length. For example, resize-enabled representation 708-3 and resize-enabled representation 708-5 differ in a first dimension (e.g., width) but are the same in other dimensions (e.g., thickness).

[0266] In some embodiments, the computer system detects FIG. 7K (e.g., Figure 7K1 , Figure 7K2 and Figure 7K3 ) in the three-dimensional environment (e.g., user input performed via the user's hand 7020, such as an air gesture). In some embodiments, in response to detecting that the user's attention 710-5 has moved away from the resizing enable representation 708-5, after detecting the user input directed to the resizing enable representation 708-5, even when the user's attention is directed to the portion of the three-dimensional environment that does not correspond to the size of the resizing enable representation 708-5, the update of the visual attributes is optionally maintained (e.g., continuing to display the resizing enable representation 708-5 with the updated size and updated color).

[0267] In some embodiments, after detecting the user input in Figure 7K, in response to detecting that the user's attention 710-5 has moved away from the resizing enable representation 708-5, the computer system continues the user's interaction with the resizing enable representation 708-5 (e.g., even when the user's attention 710-5 is not directed to the resizing enable representation 708-5). For example, after the user's attention shifts outside the first area having the first size and / or the second area having the second size, the user input detected via the user's hand 7020 is enabled to continue to interact with the resizing enable representation 708-5 (e.g., enabling the user to continue to resize the window by dragging the resizing enable representation 708-5 in one or more directions).

[0268] In some embodiments, detected user input via the user's hand 7020 indicates movement of the user's hand 7020 that causes movement of the resizing enable representation 708-5. For example, when resizing enable representation 708-5 is selected (e.g., in response to a gaze and air pinch gesture or other selection input), the user input continues by the user optionally moving the user's hand 7020 in a corresponding direction while maintaining the pinch gesture and / or moving the user's hand by a corresponding amount (e.g., the user performs a drag gesture or an air drag gesture). For example, the user performs a pinch and drag gesture while gazing at resizing enable representation 708-5. In some embodiments, in response to detecting a user's drag gesture directed to resizing enable representation 708-5, computer system 101 moves the resizing enable representation by an amount of movement and / or direction of movement corresponding to the user's drag gesture, and adjusts the size of application window 702, such as Figure 7L For example, the size of application window 702 is adjusted according to the direction and / or amount of movement of the user input detected via the user's hand 7020. For example, the user's hand 7020 moves up and to the left, and in response, the corner of the application window 702 moves up and to the left, thereby reducing the size of the application window 702 by a corresponding amount. In some embodiments, movement of the user's hand 7020 in the opposite direction (e.g., down and to the right) causes the size of the application window 702 to increase by an amount corresponding to the amount of movement of the user's hand 7020.

[0269] In some embodiments, resizing application window 702 includes maintaining the positioning of one or more edges of application window 702 within the three-dimensional environment. Figure 7L, the top edge and left edge of application window 702 remain in the same position within the three-dimensional environment before and after resizing application window 702. In some embodiments, the edges that maintain their respective positions are edges that are opposite (e.g., are not part of) the corner of the resizing affordance. For example, in FIG. 7K (e.g., Figure 7K1 and Figure 7K3 ), the resizing indicator 708-5 is displayed in the lower right corner so that the top and left edges of application window 702 are maintained when the size of application window 702 is adjusted, and the lower right corner moves inward toward the top and left edges of application window 702 to reduce the size of application window 702.

[0270] In some embodiments, resizing application window 702 includes maintaining the center of application window 702 at the same location before, after, and / or during resizing of application window 702. For example, when the size of application window 702 is reduced, multiple (or optionally, all) of the edges of application window 702 move inward (e.g., uniformly and by the same distance) toward the center of application window 702 to reduce the size of application window 702. Similarly, when the size of application window 702 is increased, multiple (or optionally, all) of the edges of application window 702 move outward away from the center of application window 702 while maintaining the center of application window 702 at the same location.

[0271] In some embodiments, when a user provides user input pointing to resize-enabled representation 708-5, the computer system increases or decreases the size of resize-enabled representation 708-5 to indicate that resize-enabled representation 708-5 is currently selected by the user. In some embodiments, resize-enabled representation 708-5 has a first size when currently selected by user input, and resize-enabled representation 708-5 has a different size (e.g., a second size different from the first size) when application window 702 is being resized. For example, resize-enabled representation 708-6 is displayed at a size that is based on (e.g., proportional to) the size of application window 702. For example, Figure 7H Compared to the application window 702 in Figure 7L The size of the application window 702 in is reduced, and Figure 7H708-6 is reduced in size (optionally by an amount proportional to the amount of reduction in the size of application window 702) compared to resize-enabled representation 708-3 in FIG. 708-6. Thus, the size of resize-enabled representation 708-5 changes based on (i) the input currently being selected by the user and / or changes by an amount based on (ii) the current size of application window 702. Thus, the size of resize-enabled representation 708-5 changes as the user resizes application window 702.

[0272] In some embodiments, such as Figure 7L As illustrated, while the user's attention 710-9 continues to be detected as being directed toward resize-enabled representation 708-6, resize-enabled representation 708-6 is displayed in a first updated state (e.g., in a second color), but in the absence of user input currently selecting resize-enabled representation 708-6, resize-enabled representation 708-6 is not displayed in its second updated state (e.g., in a different size). For example, the computer system displays resize-enabled representation 708-6 in gray to provide visual feedback that the computer system detects that the user's gaze is directed toward resize-enabled representation 708-6 (e.g., however, if the user's gaze is not directed toward the resize-enabled representation, resize-enabled representation 708-6 is displayed in white, as described above).

[0273] Figure 7M Detecting the user's attention 710-10 directed to a virtual object 7028 is illustrated. In some embodiments, the application window 702 continues to be displayed in the three-dimensional environment even when the user's attention is not directed to the application window 702. In some embodiments, even when the user's attention is not directed to the application window 702, a grabber 706-2 for the application window 702 is optionally displayed. In some embodiments, when the user's attention is not directed to the application window 702, the grabber 706-2 is optionally stopped from being displayed, and in response to detecting that the user's attention is directed to the application window 702 (e.g., or to a corresponding portion of the application window 702, such as the bottom center portion), the grabber 706-2 is displayed.

[0274] In some embodiments, after the size of application window 702 is reduced, the size of grabber 706-2 is reduced relative to the size of grabber 706-1. For example, grabber 706-1 is displayed at a size proportional to application window 702, and the size is updated as the size of application window 702 changes. In some embodiments, the size of application window 702 depends on the perceived distance from the user (e.g., and / or the user's viewpoint). For example, if the positioning of application window 702 moves away from the user, the size of application window 702 is reduced, and if the positioning of application window 702 (e.g., and its associated controls, such as grabber 706-1) moves toward the user, the size of application window 702 (e.g., and its associated controls, such as grabber 706-1) increases based on the positioning being closer to the user.

[0275] In some embodiments, in response to detecting that the user's attention 710-10 is directed toward the virtual object 7028, a base plate 7029 is displayed below the virtual object 7028. In some embodiments, the base plate 7029 includes a flat surface that appears to be substantially parallel to the floor 7008'. For example, the base plate 7029 is displayed as a surface, optionally a floating surface, on which the virtual object 7028 is located in a three-dimensional environment. In some embodiments, the base plate 7029 is displayed for a three-dimensional virtual object (such as, the virtual object 7028), while a two-dimensional object (such as, the application window 702) is displayed without the base plate. In some embodiments, the size of the base plate 7029 is based on the size of the virtual object 7028. In some embodiments, the base plate 7029 is displayed while the user's attention 710-10 continues to be directed toward the virtual object 7028 and / or one or more controls for the virtual object 7028 (e.g., the grabber 712-1, the resize enable indication 714-1, and / or the close enable indication 717), and is optionally no longer displayed in response to detecting that the user's attention has moved away from the virtual object 7028 and / or the one or more controls for the virtual object 7028.

[0276] Figure 7N A gripper 712-1 for moving a position of a virtual object 7028 is illustrated. In some embodiments, the gripper 712-1 is displayed as a user interface object distinct from the base plate 7029 and is displayed along an edge of the base plate 7029 (e.g., the nearest edge of the base plate 7029). In some embodiments, the gripper 712-1 is automatically displayed simultaneously with the display of the base plate 7029 without the need for additional user input (e.g., in response to detecting that the user's attention is directed toward the virtual object 7028). In some embodiments, in response to detecting that the user's attention is directed toward a bottom portion of the virtual object 7028 (such as, toward a portion of the virtual object 7028 that is adjacent to the gripper 712-1), the gripper 712-1 is automatically displayed simultaneously with the display of the base plate 7029. Figure 7N The user focuses on the nearest edge of the base plate 7029 in an area close to where it is displayed in the three-dimensional environment) and displays the gripper 712-1. For example, the user focuses on the nearest edge of the base plate 7029 and the gripper 712-1 is displayed. In some embodiments, the user selects the gripper 712-1 using a selection input (e.g., a direct air gesture such as an air tap or air pinch at a location where the user is interacting with it, an indirect air gesture such as an air pinch when the user's attention or the user's gaze is directed at the location where the user is interacting with it, a tap input, a pinch input, or other selection input) and moves the virtual object 7028 within the three-dimensional environment (e.g., via a drag gesture, an air drag gesture, or other movement input by the user).

[0277] Figure 7N The computer system 101 is also illustrated as detecting the user's attention 710-11 directed toward the corner of the bottom plate 7029. In some embodiments, in response to detecting that the user's attention 710-11 is directed toward the lower left corner of the bottom plate 7029, the computer system 101 displays a resize enable indication 714-1 in the lower left corner, such as Fig.7O In some embodiments, resize enable representation 714-1 is displayed as an L-shape that extends along a portion of the edge of base plate 7029 and has an outline that matches the outline of the corner of base plate 7029. In some embodiments, the user gazes at the lower left corner or the lower right corner of base plate 7029, and in response, a corresponding resize enable representation is displayed at the corresponding corner of base plate 7029. In some embodiments, an animated transition is displayed along the edge of base plate 7029, and the animated transition includes stopping the display of grabber 712-1 and initiating the display of resize enable representation 714-1 (for example, including the above reference Fig. 7E to any of the animated transitions between grabber 706-1 and resize affordance 708-1 described in FIG. 7F).

[0278] Fig.7O The example illustrates that computer system 101 detects user attention 710-12 directed to resize-enabled representation 708-1, and in response to detecting user attention 710-12, computer system 101 changes the color of resize-enabled representation 714-1 and / or updates one or more other visual attributes of resize-enabled representation 708-1. In some embodiments, in response to detecting user input (such as a mid-air pinch gesture) selecting resize-enabled representation 708-1, the computer system changes the size of resize-enabled representation 708-1 (e.g., to indicate that it has been selected by the user) and / or updates one or more other visual attributes of resize-enabled representation 708-1.

[0279] Figure 7PThe computer system 101 is illustrated as detecting a user input indicating a direction and amount of movement based on movement of the user's hand 7020 (e.g., a direct air gesture such as an air tap or air pinch at a location where the user is interacting with it, an indirect air gesture such as an air pinch when the user's attention or the user's gaze is directed at the location where the user is interacting with it, a tap input, a gaze input, a drag input, or other types of user input) and changing the size of the virtual object 7028 according to the user input (e.g., according to the movement and / or direction of movement of the user's hand 7020). In some embodiments, when the user is interacting with the resize enable representation 714-2, the size of the resize enable representation 714-2 changes to indicate that the computer system detects the user interaction (e.g., gaze input and / or air gesture). For example, while a user is interacting with resizing enable indication 714-2 (such as performing an air-drag gesture on resizing enable indication 714-2 while the user's attention 710-13 is directed toward resizing enable indication 714-2), the size of resizing enable indication 714-2 is optionally increased (e.g., by moving the user's hand 7020 while resizing enable indication 714-2 is selected).

[0280] In some embodiments, although the size of resizing enable representation 714-2 increases while the user is interacting with resizing enable representation 714-2, because the user is reducing the size of virtual object 7028, the overall size of resizing enable representation 714-2 appears to decrease (e.g., decreases in accordance with the decrease in the size of virtual object 7028). For example, resizing enable representation 714-2 is displayed at a size proportional to virtual object 7028, so that when the size of virtual object 7028 decreases, the size of resizing enable representation 714-2 also decreases. For example, the amount by which the size of resizing enable representation 714-2 decreases is less than the amount by which the size of the virtual object decreases because the size of resizing enable representation 714-2 increases while the user is interacting with resizing enable representation 714-2.

[0281] Figure 7Q The illustration shows that after user input pointing to the resize-enabling indication 714-2 is no longer detected, the resize-enabling indication 714-2 is stopped from being displayed, and optionally, the gripper 712-2 is redisplayed below the base 7029 of the virtual object 7028. In some embodiments, the gripper 712-2 is displayed at a size based on the resized virtual object 7028 (e.g., the size of the gripper 712-2 decreases as the size of the virtual object 7028 decreases). Figure 7Q It is also illustrated that the computer system 101 detects that the user's attention 710 - 14 is directed toward the upper left corner of the virtual object 7028 .

[0282] Figure 7R In response to detecting the user's attention 710-14 directed to the upper left corner of the virtual object 7028, the computer system 101 displays a close enable representation 717 for the virtual object 7028. In some embodiments, the close icon 717 is displayed at a different portion of the virtual object 7028 (e.g., in accordance with the virtual object 7028 being a three-dimensional object). For example, in some embodiments, the close enable representation 717 is displayed below the base plate 7029 of the virtual object 7028. In some embodiments, when the close enable representation 717 is displayed, the computer system detects a user input (e.g., via the user's hand 7020), such as a tap gesture (e.g., a touch gesture, an air gesture, or other selection input) and / or a pinch gesture (e.g., an air pinch gesture) directed to the close enable representation 717 when the user is looking at the close enable representation 717 (which indicates that the user's attention 710-15 is directed to the close enable representation 717) (e.g., when the close enable representation 717 is in a ready state, as described above).

[0283] In response to a user input selecting to close the enable indication 717, the computer system stops displaying the virtual object 7028 in the three-dimensional environment, such as Figure 7S For example, the computer system closes the virtual object 7028, including optionally closing an application associated with the virtual object 7028.

[0284] about FIG. 7A to FIG. 7S For additional description, refer to the following Figure 8 and Fig. 9 Methods 800 and 900 are described.

[0285] FIG. 7S to FIG. 7A D (for example, Figure 7AD1 , Figure 7AD2 and Figure 7AD3 ) illustrates an example of displaying a title bar that expands in response to detecting a user's attention directed toward the title bar. Fig.10 is a flow chart of an exemplary method 1000 for displaying a title bar adjacent to an application window that provides additional control options to a user. FIG. 7S to FIG. 7A The user interface in D is used to illustrate the process described below, including Fig.10 process.

[0286] Figure 7SDetecting the user's attention 710-16 directed to the application window 702, such as a gaze input, is illustrated. In some embodiments, the application window 702 is displayed simultaneously with the title bar 716. In some embodiments, the title bar 716 displays the name or other indication of the application associated with the application window 702, such as an application icon (e.g., "App 1"). In some embodiments, the title bar 716 indicates the corresponding content currently displayed in the application window 702. For example, the title bar 716 includes the name of the document displayed in the application window 702 and / or the name of the website displayed in the application window 702. In some embodiments, multiple tabs are displayed in the tab bar (optionally including the title bar 716), each tab corresponding to different content, wherein the content associated with each tab is available for display in the application window 702. For example, the user is enabled to switch between tabs or otherwise navigate to display different content (e.g., different documents, different web pages, or other content) of the same application and / or other applications within the application window 702. In some embodiments, an indication of other available tabs is displayed in a tab bar next to the title bar 716, so that selecting another tab in the tab bar switches the content displayed in the application window 702. In some embodiments, the currently selected tab is displayed as the current title bar of the content currently displayed in the application window 702. In some embodiments, the currently selected tab is displayed with a different visual appearance than the other tabs in the tab bar. For example, the currently selected tab bar is displayed with an application icon, a different level of translucency, a different size, and / or a different color to indicate that it is the currently active tab.

[0287] In some embodiments, the title bar 716 is displayed as a distinct user interface object having a non-zero distance between the application window 702 and the title bar 716. In some embodiments, the title bar 716 is displayed when the computer system detects that the user's attention 710-16 is directed toward the application window 702. In some embodiments, the title bar 716 is displayed even if the user's attention is not detected as being directed toward the application window 702 (e.g., in the Figure 7M to Figure 7R In some embodiments, if the user's attention is not detected as being directed toward application window 702 (e.g., and / or at this time), title bar 716 (and optionally, application window 702 and / or other controls of application window 702) is visually weakened (e.g., dimmed, reduced, or otherwise weakened).

[0288] Figure 7T to Figure 7AH An example of conditionally displaying a privacy indicator is illustrated. Fig.11 7T to 7T are flow diagrams of an exemplary method 1100 for maintaining a privacy indicator for an application window as the application window moves in a display area. Figure 7AH The user interface in is used to illustrate the process described below, including Fig.11 process.

[0289] FIG. 7T (for example, Figure 7T1 , Figure 7T2 and Figure 7T3 ) illustrates detecting that the user's attention 710-17 continues to be directed to the application window 702. In some embodiments, when the user's attention 710-17 is directed to the application window 702, the application associated with the application window 702 begins to access, use and / or collect sensor data from one or more sensors of the computer system. For example, the application associated with the application window 702 accesses one or more of the microphone, camera, location, or other sensors of the computer system that provide the application with access to the user's sensitive or private data. In some embodiments, in response to detecting that the application is accessing one or more of the sensors that provide access to the user's sensitive data, the computer system displays a privacy indicator 718-1 above the application window 702, while optionally also maintaining the display of the title bar 716.

[0290] In some embodiments, the privacy indicator 718-1 is displayed with a first set of attributes that indicate which of the one or more sensors is being accessed by the application associated with the application window 702. For example, the privacy indicator 718-1 is displayed with a corresponding color corresponding to the sensor type (e.g., a red indicator indicates that the camera is being accessed, an orange indicator indicates that the microphone is being accessed, and / or a blue indicator indicates that the location data is being accessed). It should be understood that different visual attributes and / or colors may be assigned to specific sensors to indicate which of the sensors are currently being accessed by the application associated with the application window 702.

[0291] In some embodiments, privacy indicator 718-1 is displayed even if the user is not currently directing the user's attention to application window 702. Thus, the computer system indicates to the user whether the user is currently interacting with or focusing on application window 702 when an application is using one or more sensors of the computer system to access sensitive data.

[0292] Figure 7UThe computer system 101 is illustrated as detecting that the user's attention 710-18 is directed to the privacy indicator 718-1. In some embodiments, in response to detecting that the user's attention 710-18 is directed to the privacy indicator 718-1, the computer system displays additional information about the sensor that the application associated with the application window 702 is accessing, such as Figure 7V As illustrated. For example, privacy indicator 718-1 is expanded into expanded privacy indicator 718-2, which includes a text indication that the application associated with application window 702 is accessing the microphone (e.g., and / or an icon representing the sensor). In some embodiments, expanded privacy indicator 718-2 optionally provides the user with a selectable option to disable access to the sensor so that the application can no longer use the sensor to collect sensitive data. In some embodiments, based on determining that the application is currently accessing two or more sensors, expanded privacy indicator 718-2 lists or otherwise indicates each of the sensors that the application associated with application window 702 is accessing.

[0293] Figure 7V The example shows a grabber 706-2 that detects the user's attention 710-19 directed toward the application window 702. In some embodiments, as Figure 7W As illustrated, in response to detecting that the user's attention 710-19 is directed toward the gripper 706-2 (e.g., at Figure 7V ), the computer system displays grabber 706-3 in a color different from the color of grabber 706-2 (e.g., updates the display of grabber 706-2 to grabber 706-3 and / or replaces the display of grabber 706-2 with grabber 706-3) to indicate that the computer system detects that the user's attention is directed to grabber 706-2 (e.g., and directed to the currently displayed grabber 706-3), which is similar to reference Figures 7F to 7F. Figure 7G The resizing indicator 708-2 may represent a change in the visual attributes of the display.

[0294] Figure 7XDetection of user input of selecting gripper 706-4 (e.g., which is an updated display of gripper 706-3 and / or replaces gripper 706-3) via a user's hand 7020 is illustrated. In some embodiments, the user input is an air tap gesture, an air pinch gesture, or another selection gesture (e.g., an air gesture) detected when the user is looking at gripper 706-4. In some embodiments, in response to detecting user input selecting gripper 706-4, the user input continues by moving the user's hand 7020 to drag gripper 706-4 from its corresponding position (e.g., via an air drag gesture) to another position within the three-dimensional environment. In some embodiments, the user is enabled to move gripper 706-4 and associated application window 702 in three dimensions, including changing the position of application window 702 in a horizontal direction (e.g., left and / or right), in a vertical direction (e.g., up and / or down), and / or in depth (e.g., forward and / or backward) in a three-dimensional environment.

[0295] Figure 7Y The example illustrates that in response to a user moving the user's hand 7020 (e.g., from left to right and / or optionally away from the user's body) when grabber 706-5 (e.g., which is similar to grabber 706-4, but represents a grabber as the user continues to interact with the grabber) is selected (e.g., and the user's attention 710-22 is detected as being directed toward grabber 706-5), the application window 702 moves in the three-dimensional environment based on the position and / or amount of movement of the user's hand 7020. For example, if the user moves the user's hand to the left, the application window 702 moves to the left based on the movement of the user's hand. In some embodiments, as the user moves the application window 702 in the three-dimensional environment, the title bar 716 and / or the privacy indicator 718-1 continue to be displayed at the same respective positions relative to the application window 702.

[0296] Figure 7Z The example illustrates displaying the gripper 706-6 without updating the one or more visual attributes (e.g., the color of the gripper 706-6 that was the same as before the user input and / or attention was detected directed to the gripper 706-2) in response to detecting the end of the user input, such as the user releasing the pinch gesture or the mid-air pinch gesture directed to the gripper 706-5 (e.g., the user moving the user's hand 7020), stopping moving the user's hand 7020, or lifting off (e.g., the lifting off of the user input directed to the gripper 706-5). Figure 7V In some embodiments, gripper 706-6 is the same as gripper 706-2 (e.g., computer system 101 redisplays gripper 706-2). Figure 7ZAlso illustrated is that the application window 702 has been repositioned (eg, using grabber 706 - 5 ) to a different position in the three-dimensional environment.

[0297] Figure 7Z It is also illustrated that even while application window 702 is being moved (e.g., and after application window 702 has been moved), title bar 716 and privacy indicator 718-1 continue to be displayed at their same respective locations relative to application window 702. For example, title bar 716 continues to be displayed above application window 702, in the center of application window 702, and privacy indicator 718-1 is displayed to the right of title bar 716 and above application window 702. It should be understood that alternative arrangements of title bar 716 and / or privacy indicator 718-1 relative to application window 702 may be implemented (e.g., to the right and / or left of application window 702), but the respective locations of the title bar and / or the respective locations of the privacy indicator remain the same even when application window 702 is moved to a different location within the three-dimensional environment. Additionally, in som...

Claims

1. A method, comprising: At a computer system in communication with a first display generating component and one or more input devices: displaying, via the first display generating component, a first object in a first view of a three-dimensional environment, wherein the first object comprises at least a first portion of the first object and a second portion of the first object; detecting, via the one or more input devices, a first gaze input that satisfies a first criterion while the first object is displayed, wherein the first criterion requires that the first gaze input be directed toward the first portion of the first object in order to satisfy the first criterion; In response to detecting that the first gaze input satisfies the first criterion, displaying a first control element corresponding to a first operation associated with the first object, wherein the first control element is not displayed before detecting that the first gaze input satisfies the first criterion; while the first control element is displayed, detecting, via the one or more input devices, a first user input directed to the first control element; as well as The first operation is performed relative to the first object in response to detecting the first user input directed to the first control element. 2 . The method of claim 1 , wherein displaying the first object in the first view of the three-dimensional environment comprises displaying an application window of a first application in the first view of the three-dimensional environment.

3. The method according to any one of claims 1 to 2, comprising: while the first object is displayed, detecting, via the one or more input devices, a third gaze input directed to the second portion of the first object, the second portion being different from the first portion of the first object, the third gaze input not satisfying the first criterion; as well as In response to detecting that the third gaze input does not satisfy the first criterion, display of the first control element corresponding to the first operation associated with the first object is abandoned.

4. The method of any one of claims 1 to 3, wherein the first control element is a first resize affordance; and The method comprises: while displaying the first resizable affordance, detecting, via the one or more input devices, a second user input directed to the first resizable affordance; as well as In response to detecting the second user input directed toward the first resize affordance, the first object is resized.

5. The method according to claim 4, wherein: Detecting the second user input directed toward the first resizable affordance includes detecting a direction of movement of the second user input directed toward the first resizable affordance; and Resizing the first object in response to detecting the second user input directed to the first resize affordance includes: According to determining that the moving direction of the second user input is a first direction, increasing the size of the first object; as well as Based on determining that the movement direction of the second user input is a second direction different from the first direction, the size of the first object is reduced.

6. The method according to any one of claims 4 to 5, wherein: Detecting the second user input directed toward the first resizable affordance includes detecting an amount of movement of the second user input directed toward the first resizable affordance; and Resizing the first object in response to detecting the second user input directed to the first resize affordance includes: According to determining that the movement amount input by the second user is the first movement amount, changing the size of the first object to a first size selected based on the first movement amount input by the second user; as well as Based on determining that the movement amount of the second user input is a second movement amount different from the first movement amount, the size of the first object is changed to a second size different from the first size, the second size being selected based on the second movement amount of the second user input.

7. The method according to any one of claims 4 to 6, wherein: Detecting the first user input directed to the first control element includes detecting a first mid-air gesture directed to the first control element; and The first object is resized in response to detecting the first air gesture directed to the first control element.

8. The method of any one of claims 4 to 7, wherein the first portion of the first object comprises a first corner of the first object, and the first criterion requires that the first gaze input is directed to the first corner of the first object in order to satisfy the first criterion.

9. A method according to any one of claims 1 to 8, wherein the first part of the first object includes a first subpart of the first object and a second subpart of the first object, wherein the first subpart of the first object and the second subpart of the first object are separated by a third subpart of the first object that is not included in the first part of the first object, and the first criterion requires that the first gaze input points to at least one of the first subpart and the second subpart of the first object in order to satisfy the first criterion.

10. The method according to claim 9, wherein: Detecting the first user input while the first control element is displayed includes detecting the first user input directed to the first sub-portion of the first object or the second sub-portion of the first object; and In response to detecting the first user input directed to the first control element, performing the first operation relative to the first object includes: in response to determining that the first user input is directed to the first sub-portion of the first object, changing the size of the first object to a first size while maintaining the location of the center of the first object; as well as Based on determining that the first user input is directed to the second sub-portion of the first object, the size of the first object is changed to a second size while maintaining the location of the center of the first object in the three-dimensional environment.

11. The method according to any one of claims 9 to 10, wherein: The first portion of the first object corresponds to a first edge of the first object and does not correspond to a second edge of the first object; and In response to detecting the first user input directed to the first control element, performing the first operation relative to the first object includes: The first object is resized by moving the first edge of the first object while maintaining the position of the second edge of the first object in the three-dimensional environment.

12. The method according to any one of claims 1 to 11, comprising: while the first object is displayed, detecting, via the one or more input devices, a fourth gaze input directed to a corresponding portion of the first object; In response to detecting the fourth gaze input: displaying the first control element based on determining that the fourth gaze input is directed to the first portion of the first object and satisfies the first criterion relative to the first portion of the first object; displaying a second control element corresponding to a second operation associated with the first object in response to determining that the fourth gaze input is directed to a third portion of the first object that is different from the first portion of the first object, wherein the second control element is not displayed before detecting that the fourth gaze input is directed to the third portion of the first object; while the second control element is displayed, detecting a third user input directed to the second control element; as well as The second operation is performed with respect to the first object in response to detecting the third user input directed to the second control element.

13. The method of any one of claims 1 to 3, wherein the first control element is a shutdown affordance; and The method comprises: while displaying the close enable indication, detecting, via the one or more input devices, a fourth user input directed toward the close enable indication; as well as In response to detecting the fourth user input directed toward the close affordance, closing the first object includes ceasing display of the first object in the first view of the three-dimensional environment.

14. The method according to any one of claims 1 and 12 to 13, comprising: detecting a fifth gaze input directed toward the first object while displaying the first object and the second object in the first view of the three-dimensional environment; In response to detecting that the fifth gaze input is directed toward the first object: displaying a first off affordance for the first object based on determining that the fifth gaze input is directed toward a first subportion of the first object and satisfies the first criterion relative to the first subportion of the first object; as well as displaying a first resize affordance for the first object based on determining that the fifth gaze input is directed to a second subportion of the first object and satisfies the first criterion relative to the second subportion of the first object; while displaying the first object and the second object in the first view of the three-dimensional environment, detecting a sixth gaze input directed toward the second object; In response to detecting that the sixth gaze input is directed toward the second object: displaying a second off affordance for the second object based on determining that the sixth gaze input is directed toward a first subportion of the second object and satisfies the first criterion relative to the first subportion of the second object; as well as Based on determining that the sixth gaze input is directed to a second subportion of the second object and satisfies the first criterion relative to the second subportion of the second object, display of a second resizable affordance for the second object is abandoned.

15. The method according to any one of claims 1 and 12 to 13, comprising: detecting a seventh gaze input directed toward the first object while displaying the first object and a third object in the first view of the three-dimensional environment; In response to detecting that the seventh gaze input is directed toward the first object: displaying a first off affordance for the first object based on determining that the seventh gaze input is directed toward a first subportion of the first object and satisfies the first criterion relative to the first subportion of the first object; as well as displaying a first movement affordance for the first object based on determining that the seventh gaze input is directed toward a third subportion of the first object and satisfies the first criterion relative to the third subportion of the first object; detecting, while displaying the first object and the third object in the first view of the three-dimensional environment, an eighth gaze input directed toward the third object; In response to detecting that the eighth gaze input is directed toward the third object: displaying a third off affordance for the third object based on determining that the eighth gaze input is directed toward a first subportion of the third object and satisfies the first criterion relative to the first subportion of the third object; as well as Based on determining that the eighth gaze input is directed to a second subportion of the third object and satisfies the first criterion relative to the second subportion of the third object, display of a second movement affordance for the third object is abandoned.

16. The method according to any one of claims 1 to 3 and 9 to 15, comprising: while displaying the first object, displaying a first move affordance for repositioning the first object in the first view of the three-dimensional environment, detecting a ninth gaze input directed to a respective portion of the first object corresponding to the first resize affordance; In response to detecting the ninth gaze input directed to the corresponding portion of the first object corresponding to the first resize affordance: ceasing display of said first movement affordance; as well as The first resize-enabling representation is displayed, and the first resize-enabling representation is used to resize the first object at or near the corresponding portion of the first object corresponding to the first resize-enabling representation.

17. The method according to any one of claims 1 to 16, comprising: while displaying the first object, displaying a second move affordance for repositioning the first object in the first view of the three-dimensional environment, detecting a tenth gaze input directed to a respective portion of the first object corresponding to the second resize affordance; In response to detecting the tenth gaze input directed to the corresponding portion of the first object corresponding to the second resize enable representation, displaying an animated transition between displaying the second move enable representation and displaying the second resize enable representation, including: moving the second move-enabled representation toward a position corresponding to the second resize-enabled representation; and The second resizable-enabled representation is displayed at the position corresponding to the second resizable-enabled representation.

18. A method according to any one of claims 16 to 17, comprising: while displaying a corresponding movement affordance for moving the first object in the three-dimensional environment, detecting a fifth user input directed toward the corresponding movement affordance for moving the first object; as well as In response to detecting the fifth user input directed toward the corresponding mobile enable representation, the display of the first object is updated from being displayed at a first object location in the first view of the three-dimensional environment to being displayed at a second object location in the first view of the three-dimensional environment that is different from the first object location.

19. The method of claim 18, wherein updating the display of the first object from being displayed at the first object location in the first view of the three-dimensional environment to being displayed at the second object location comprises updating the location of the first object in three different dimensions in the three-dimensional environment.

20. The method according to any one of claims 1 to 11, comprising: while displaying the first object, displaying a third movement affordance for repositioning the first object in the first view of the three-dimensional environment, and detecting an eleventh gaze input directed toward a corresponding portion of the first object; In response to detecting the eleventh gaze input directed to the corresponding portion of the first object: Based on determining that the corresponding portion of the first object corresponds to a first type of control element for the first object, display of the third mobile enable representation is stopped and a corresponding instance of the first type of control element is displayed.

21. The method according to claim 20, comprising: In response to detecting the eleventh gaze input directed to the corresponding portion of the first object: Based on determining that the corresponding portion of the first object corresponds to a second type of control element for the first object that is different from the first type of control element, display of the third mobile enable representation is maintained and a corresponding instance of the second type of control element is displayed.

22. The method of any one of claims 1 to 21, wherein in response to detecting that the first gaze input satisfies the first criterion relative to the first portion of the first object, displaying the first control element corresponding to the first operation associated with the first object comprises: Based on determining that the first object is displayed as a two-dimensional object in the three-dimensional environment, displaying the first control element at a first location having a first spatial relationship with the first object 23. The method of claim 22, wherein in response to detecting that the first gaze input satisfies the first criterion relative to the first portion of the first object, displaying the first control element corresponding to the first operation associated with the first object comprises: Based on determining that the first object is displayed as a three-dimensional object in the three-dimensional environment, the first control element is displayed at a second location having a second spatial relationship with the first object, wherein the first spatial relationship is different from the second spatial relationship.

24. The method according to claim 23, comprising: In response to detecting that the first portion of the first gaze input relative to the first object satisfies the first criterion: Based on determining that the first object is displayed as a three-dimensional object in the three-dimensional environment, a second object is displayed in the three-dimensional environment via the first display generation component, wherein the second object is displayed as a three-dimensional application object having a third spatial relationship with the first object.

25. The method according to claim 24, comprising: A second control element is displayed below the second object.

26. A method according to any one of claims 1 to 25, comprising: Before displaying the first object in the first view of the three-dimensional environment: detecting a sixth user input corresponding to a request to display the first object in the first view of the three-dimensional environment; In response to detecting the sixth user input corresponding to the request to display the first object in the first view of the three-dimensional environment: simultaneously displaying the first object and a first set of control elements for the first object in the first view of the three-dimensional environment, the first set of control elements including a first control element; as well as After simultaneously displaying the first object and the first set of control elements for the first object for a threshold amount of time, display of the first set of control elements for the first object is ceased while display of the first object is maintained.

27. The method of claim 26, wherein: The first set of control elements includes a means for stopping display of a closed affordance representation of the first object in the three-dimensional environment.

28. The method of any one of claims 26 to 27, wherein after simultaneously displaying the first object and the first set of control elements for the first object for the threshold amount of time, while maintaining display of the first object, ceasing display of the first set of control elements for the first object, comprises: moving the first set of control elements toward the first object while changing one or more visual attributes of the first set of control elements; as well as After moving the first group of control elements and changing the one or more visual attributes of the first group of control elements, display of the first group of control elements is stopped.

29. The method according to any one of claims 1 to 3 and 13 to 28, wherein: while displaying the first object, displaying a fourth movement affordance for repositioning the first object in the first view of the three-dimensional environment; while displaying the fourth movement affordance with the first object, detecting a twelfth gaze input directed toward a respective portion of the three-dimensional environment corresponding to the first object; as well as In response to detecting the twelfth gaze input directed toward the corresponding portion of the three-dimensional environment corresponding to the first object: Based on determining that the twelfth gaze input corresponds to a request to display an off affordance for the first object, a corresponding off affordance is displayed adjacent to the fourth move affordance.

30. The method of claim 29, further comprising: Prior to detecting the twelfth gaze input directed toward the corresponding portion of the three-dimensional environment corresponding to the first object, a first preview of the corresponding closed enable representation is displayed adjacent to the fourth mobile enable representation, wherein determining that the twelfth gaze input corresponds to a request to display a closed enable representation for the first object includes determining that the twelfth gaze input is directed toward the first preview of the corresponding closed enable representation.

31. A method according to any one of claims 29 to 30, comprising: detecting, via the one or more input devices, a seventh user input directed to the corresponding close enable indication while the corresponding close enable indication is displayed; as well as In response to detecting the seventh user input directed to the corresponding close enable representation, closing the first object based on determining that the seventh user input meets the selection criteria, including stopping display of the first object in the first view of the three-dimensional environment.

32. A method according to any one of claims 1 to 31, comprising: detecting, via the one or more input devices, a thirteenth gaze input moving relative to a respective portion of the first object corresponding to a respective control element of the first object while the first object is displayed in the first view of the three-dimensional environment; In response to detecting the thirteenth gaze input moving relative to the corresponding part of the first object corresponding to the corresponding control element of the first object, determining whether the user attention is directed to the corresponding part of the first object corresponding to the corresponding control element of the first object includes: determining that user attention is directed to the corresponding portion of the first object corresponding to the corresponding control element based on determining that the thirteenth gaze input moved relative to the corresponding portion of the first object has moved into a first corresponding portion of the first object corresponding to the corresponding control element of the first object, wherein the first corresponding portion of the first object corresponding to the corresponding control element has a first size; determining that user attention remains on the corresponding portion of the first object corresponding to the corresponding control element based on determining that the thirteenth gaze input moved relative to the corresponding portion of the first object has moved from within the first corresponding portion of the first object to an area outside the first corresponding portion of the first object and within a second corresponding portion of the first object corresponding to the corresponding control element of the first object, wherein the second corresponding portion of the first object corresponding to the corresponding control element has a second size greater than the first size; and Based on determining that the thirteenth gaze input moved relative to the corresponding part of the first object has moved from within the first corresponding part of the first object to an area outside the second corresponding part of the first object corresponding to the corresponding control element of the first object, it is determined that the user's attention has left the corresponding part of the first object corresponding to the corresponding control element.

33. A method according to any one of claims 1 to 32, comprising: while displaying the first object in the first view of the three-dimensional environment, detecting an eighth user input when user attention is directed toward the first portion of the first object; as well as In response to detecting the eighth user input while user attention is directed to the first portion of the first object, initiating performance of one or more operations corresponding to the first portion of the first object according to the eighth user input; detecting that user attention has moved away from the first portion of the first object while performing the one or more operations corresponding to the first portion of the first object; as well as After detecting that the user's attention has moved away from the first part of the first object, and based on determining that the one or more operations corresponding to the first part of the first object are in progress, continue to perform the one or more operations corresponding to the first part of the first object.

34. A method according to any one of claims 1 to 33, wherein: the first criterion requiring that the first gaze input be maintained on the first portion of the first object for at least a first threshold amount of time in order to satisfy the first criterion; and, Displaying the first control element includes displaying the first control element after the first gaze input is directed at the first portion of the first object for at least the first threshold amount of time.

35. The method according to any one of claims 1 to 12, 14 to 26 and 29 to 34, comprising: detecting a start of a first portion of the first user input when the first gaze input is directed to the first portion of the first object and before displaying the first control element; In response to detecting the start of the first portion of the first user input while the first gaze input is directed to the first portion of the first object and before displaying the first control element: Based on determining that an operation execution criterion is met, before displaying the first control element, initiating execution of the first operation based on the first part of the first user input, wherein: detecting the first user input directed to the first control element comprises detecting a second portion of the first user input after detecting the first portion of the first user input and after displaying the first control element; and Performing the first operation relative to the first object in response to detecting the first user input directed to the first control element includes continuing to perform the first operation according to the second portion of the first user input.

36. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a first display generating component and one or more input devices, the one or more programs comprising instructions for executing the method according to any one of claims 1 to 35.

37. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 1 to 35.

38. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: A component for carrying out the method according to any one of claims 1 to 35.

39. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a first display generating component and one or more input devices, the one or more programs comprising instructions for: displaying, via the first display generating component, a first object in a first view of a three-dimensional environment, wherein the first object comprises at least a first portion of the first object and a second portion of the first object; detecting, via the one or more input devices, a first gaze input that satisfies a first criterion while the first object is displayed, wherein the first criterion requires that the first gaze input be directed toward the first portion of the first object in order to satisfy the first criterion; In response to detecting that the first gaze input satisfies the first criterion, displaying a first control element corresponding to a first operation associated with the first object, wherein the first control element is not displayed before detecting that the first gaze input satisfies the first criterion; while the first control element is displayed, detecting, via the one or more input devices, a first user input directed to the first control element; as well as The first operation is performed relative to the first object in response to detecting the first user input directed to the first control element.

40. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the first display generating component, a first object in a first view of a three-dimensional environment, wherein the first object comprises at least a first portion of the first object and a second portion of the first object; detecting, via the one or more input devices, a first gaze input that satisfies a first criterion while the first object is displayed, wherein the first criterion requires that the first gaze input be directed toward the first portion of the first object in order to satisfy the first criterion; In response to detecting that the first gaze input satisfies the first criterion, displaying a first control element corresponding to a first operation associated with the first object, wherein the first control element is not displayed before detecting that the first gaze input satisfies the first criterion; while the first control element is displayed, detecting, via the one or more input devices, a first user input directed to the first control element; as well as The first operation is performed relative to the first object in response to detecting the first user input directed to the first control element.

41. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: means for displaying, via the first display generating component, a first object in a first view of a three-dimensional environment, wherein the first object comprises at least a first portion of the first object and a second portion of the first object; means enabled when the first object is displayed for detecting, via the one or more input devices, a first gaze input that satisfies a first criterion, wherein the first criterion requires that the first gaze input be directed toward the first portion of the first object in order to satisfy the first criterion; means for displaying a first control element corresponding to a first operation associated with the first object, enabled in response to detecting that the first gaze input satisfies the first criterion, wherein the first control element is not displayed prior to detecting that the first gaze input satisfies the first criterion; means, enabled when the first control element is displayed, for detecting, via the one or more input devices, a first user input directed to the first control element; and Means for performing the first operation relative to the first object is enabled in response to detecting the first user input directed to the first control element.

42. A method comprising: At a computer system in communication with a first display generating component and one or more input devices displaying, via the first display generating component, in a first view of a three-dimensional environment, a first user interface object and a first control element associated with performing a first operation with respect to the first user interface object, wherein the first control element is spaced apart from the first user interface object in the first view of the three-dimensional environment, and wherein the first control element is displayed with a first appearance; detecting, via the one or more input devices, a first gaze input directed toward the first control element while the first control element is displayed in the first appearance; in response to detecting the first gaze input directed toward the first control element, updating an appearance of the first control element from the first appearance to a second appearance different from the first appearance; detecting, via the one or more input devices, a first user input directed to the first control element while the first control element is displayed in the second appearance; as well as In response to detecting the first user input directed to the first control element: updating the appearance of the first control element from the second appearance to a third appearance based on determining that the first user input satisfies a first criterion, The third appearance is different from the first appearance and the second appearance and indicates that additional movement associated with the first user input will cause the first operation associated with the first control element to be performed.

43. The method according to claim 42, comprising: detecting a second user input while the first control element is displayed in the third appearance, the second user input comprising an additional movement associated with the first user input directed to the first control element; as well as In response to detecting the second user input, performing the first operation relative to the first user interface object in accordance with the additional movement of the second user input.

44. A method according to any one of claims 42 to 43, comprising: When the first control element is displayed in the second appearance, detecting, via the one or more input devices, that the first gaze input is no longer directed toward the first control element, and In response to detecting that the first gaze input is no longer directed toward the first control element when the first control element is displayed in the second appearance, restoring the appearance of the first control element from the second appearance to the first appearance.

45. A method according to any one of claims 42 to 44, comprising: When the first control element is displayed in the third appearance, detecting, via the one or more input devices, that the first gaze input is no longer directed toward the first control element; as well as In response to detecting that the first gaze input is no longer directed toward the first control element when the first control element is displayed in the third appearance, display of the first control element in the third appearance is maintained.

46. ​​A method according to any one of claims 42 to 45, wherein the first user interface object is a first application window displayed in the first view of the three-dimensional environment, and the first control element is associated with performing the first operation relative to the first application window.

47. The method of any one of claims 42 to 46, wherein updating the appearance of the first control element from the first appearance to the second appearance comprises updating a color of the first control element from a first color to a second color different from the first color.

48. The method of any one of claims 42 to 47, wherein updating the appearance of the first control element from the second appearance to the third appearance comprises updating a size of the first control element from a first size to a second size different from the first size.

49. The method of claim 48, wherein updating the size of the first control element from the first size to the second size comprises reducing the size of the first control element from the first size to the second size that is smaller than the first size.

50. The method of any one of claims 48 to 49, wherein reducing the size of the first control element from the first size to the second size that is smaller than the first size comprises: A first dimension of the first control element is reduced by a first amount, and a second dimension of the first control element is reduced by a second amount different than the first amount.

51. A method according to any one of claims 42 to 50, wherein the first control element is an object movement control for moving the first user interface object within the three-dimensional environment.

52. The method of claim 51, comprising: while displaying the object move control in the third appearance, detecting a third user input, the third user input comprising a first additional movement associated with the first user input directed to the object move control; as well as In response to detecting the third user input directed to the object movement control, the first user interface object is moved in the three-dimensional environment according to the first additional movement of the third user input.

53. A method according to any one of claims 42 to 50, wherein the first control element is an object resize control for changing the size of the first user interface object within the three-dimensional environment.

54. The method of claim 53, comprising: While the object resize control is displayed in the third appearance, detecting a fourth user input, the fourth user input comprising a second additional movement of the first user input directed to the object resize control; and In response to detecting the fourth user input directed to the object resize control, changing the size of the first user interface object in accordance with the second additional movement of the fourth user input.

55. A method according to any one of claims 42 to 50, wherein the first control element is an object close control for stopping display of the first user interface object in the three-dimensional environment.

56. A method according to any one of claims 42 to 55, comprising: detecting, via the one or more input devices, a second gaze input directed toward a portion of the three-dimensional environment while the first control element is displayed with a first corresponding appearance; as well as In response to detecting the second gaze input directed toward the portion of the three-dimensional environment: updating the appearance of the first control element from the first corresponding appearance to a second corresponding appearance different from the first corresponding appearance based on determining that the second gaze input is directed to a first position relative to the first control element in the three-dimensional environment; as well as Based on determining that the second gaze input is directed to a second location in the three-dimensional environment relative to the first control element, display of the first control element from the three-dimensional environment is stopped.

57. The method of claim 56, comprising: In response to detecting the second gaze input directed toward the portion of the three-dimensional environment: Based on determining that the second gaze input points to the second position relative to the first control element in the three-dimensional environment, a second control element associated with performing a second operation relative to the first user interface object is displayed in the first view of the three-dimensional environment, wherein the second control element is separated from the first user interface object in the first view of the three-dimensional environment.

58. A method according to any one of claims 42 to 57, wherein: The first view of the three-dimensional environment corresponds to a first viewpoint of a user of the computer system, and Displaying the first control element in the first appearance includes: Based on determining that the first user interface object is displayed at a first location within the three-dimensional environment at a first distance from the first viewpoint of the user, displaying the first control element at a first simulated size corresponding to the first distance; and Based on determining that the first user interface object is displayed at a second location within the three-dimensional environment a second distance from the first viewpoint of the user, the first control element is displayed at a second simulated size corresponding to the second distance.

59. A method according to any one of claims 42 to 58, wherein: displaying the first control element at a first size when the first user interface object is displayed at a first location within the three-dimensional environment at a first distance from a first viewpoint of the user; and The method further comprises: detecting movement of the first user interface object from the first location to a second location a second distance from the first viewpoint of the user, wherein the second distance is greater than the first distance; as well as In response to detecting the movement of the first user interface object from the first location to the second location that is further away from the first viewpoint of the user than the first location: The first control element is displayed at or near the second distance from the first viewpoint of the user at an increased simulated size compared to when the first control element is displayed at or near the first distance from the first viewpoint of the user.

60. The method of any one of claims 42 to 59, wherein: displaying the first control element at a first size when the first user interface object is displayed at a first location within the three-dimensional environment at a first distance from a first viewpoint of the user; and The method further comprises: detecting a third location of the first user interface object from the first location to a third distance from the first viewpoint of the user, wherein the third distance is less than the first distance; as well as In response to detecting the movement of the first user interface object from the first location to the third location that is closer to the first viewpoint of the user than the first location: The first control element is displayed at or about the third distance from the first viewpoint of the user at a reduced simulated size compared to when the first control element is displayed at or about the first distance from the first viewpoint of the user.

61. The method of any one of claims 42 to 57, wherein: The first view of the three-dimensional environment corresponds to a first viewpoint of the user, and Displaying the first user interface object includes: Based on determining that the first user interface object is displayed at a first location within the three-dimensional environment at a first distance from the first viewpoint of the user, displaying the first user interface object at a third simulated size corresponding to the first distance; and Based on determining that the first user interface object is displayed at a second location within the three-dimensional environment a second distance from the first viewpoint of the user, the first control element is displayed at a second simulated size corresponding to the second distance.

62. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 42 to 61.

63. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 42 to 61.

64. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: A component for carrying out the method according to any one of claims 42 to 61.

65. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a first display generating component and one or more input devices, the one or more programs comprising instructions for: displaying, via the first display generating component, in a first view of a three-dimensional environment, a first user interface object and a first control element associated with performing a first operation with respect to the first user interface object, wherein the first control element is spaced apart from the first user interface object in the first view of the three-dimensional environment, and wherein the first control element is displayed with a first appearance; detecting, via the one or more input devices, a first gaze input directed toward the first control element while the first control element is displayed in the first appearance; in response to detecting the first gaze input directed toward the first control element, updating an appearance of the first control element from the first appearance to a second appearance different from the first appearance; detecting, via the one or more input devices, a first user input directed to the first control element while the first control element is displayed in the second appearance; as well as In response to detecting the first user input directed to the first control element: updating the appearance of the first control element from the second appearance to a third appearance based on determining that the first user input satisfies a first criterion, The third appearance is different from the first appearance and the second appearance and indicates that additional movement associated with the first user input will cause the first operation associated with the first control element to be performed.

66. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the first display generating component, in a first view of a three-dimensional environment, a first user interface object and a first control element associated with performing a first operation with respect to the first user interface object, wherein the first control element is spaced apart from the first user interface object in the first view of the three-dimensional environment, and wherein the first control element is displayed with a first appearance; detecting, via the one or more input devices, a first gaze input directed toward the first control element while the first control element is displayed in the first appearance; in response to detecting the first gaze input directed toward the first control element, updating an appearance of the first control element from the first appearance to a second appearance different from the first appearance; detecting, via the one or more input devices, a first user input directed to the first control element while the first control element is displayed in the second appearance; as well as In response to detecting the first user input directed to the first control element: Based on determining that the first user input satisfies a first criterion, the appearance of the first control element is updated from the second appearance to a third appearance, the third appearance being different from the first appearance and the second appearance and indicating that additional movement associated with the first user input will cause the first operation associated with the first control element to be performed.

67. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: means for displaying, via the first display generating component, in a first view of a three-dimensional environment, a first user interface object and a first control element associated with performing a first operation with respect to the first user interface object, wherein the first control element is spaced apart from the first user interface object in the first view of the three-dimensional environment, and wherein the first control element is displayed with a first appearance; means for detecting, via the one or more input devices, a first gaze input directed toward the first control element, enabled when the first control element is displayed in the first appearance; means enabled in response to detecting the first gaze input directed toward the first control element for updating an appearance of the first control element from the first appearance to a second appearance different from the first appearance; means for detecting, via the one or more input devices, a first user input directed to the first control element, enabled when the first control element is displayed in the second appearance; and means for enabling, in response to detecting the first user input directed to the first control element: Based on determining that the first user input satisfies a first criterion, the appearance of the first control element is updated from the second appearance to a third appearance, the third appearance being different from the first appearance and the second appearance and indicating that additional movement associated with the first user input will cause the first operation associated with the first control element to be performed.

68. A method comprising: At a computer system in communication with a first display generating component and one or more input devices displaying, via the first display generating component, a first application window and a first title bar of the first application window simultaneously, wherein the first title bar of the first application window is separated from the first application window on a first side of the first application window and displays a corresponding identifier of the first application window; When displaying the first application window together with the first title bar separated from the first application window, detecting, via the one or more input devices, that the user's attention is directed toward the first title bar; as well as In response to detecting that the user's attention is directed to the first title bar: Based on determining that the user's attention meets a first criterion relative to the first title bar, the first title bar is expanded to display one or more first selectable controls for interacting with a first application corresponding to the first application window, wherein the one or more first selectable controls are not displayed before the first title bar is expanded.

69. The method of claim 68, comprising: displaying, via the first display generating component, a second application window and a second title bar of the second application window, wherein the second title bar of the second application window is separated from the second application window on a first side of the second application window and displays a corresponding identifier of the second application window; When the second application window is displayed together with the second title bar separated from the second application window, detecting, via the one or more input devices, that the user's attention is directed to the second title bar of the second application window; as well as In response to detecting that the user's attention is directed to the second title bar of the second application window: Based on determining that the user's attention relative to the second title bar of the second application window satisfies the first criterion, the second title bar is expanded to display one or more second selectable controls for interacting with a second application corresponding to the second application window, wherein the one or more second selectable controls are not displayed before the second title bar is expanded.

70. The method according to any one of claims 68 to 69, comprising: In response to detecting that the user's attention is directed to the first title bar of the first application window: Based on determining that the user's attention relative to the first title bar of the first application window does not satisfy the first criterion, abandoning extending the first title bar.

71. The method of any one of claims 68 to 70, wherein in response to detecting that the user's attention is directed toward the first title bar and based on determining that the user's attention satisfies the first criterion relative to the first title bar, expanding the first title bar to display the one or more first selectable controls comprises: extending the first title bar to overlap at least a portion of the first application window; as well as The one or more first selectable controls are displayed over the portion of the first application window.

72. A method according to any one of claims 68 to 71, wherein the first title bar is one of a set of one or more application control affordances displayed as separate from the first application window on a respective side of the first application window, and the method comprises: When the first application window and the set of one or more application control affordances separated from the first application window are displayed together, detecting, via the one or more input devices, that the user's attention is directed toward a corresponding application control affordance in the set of one or more application control affordances; as well as In response to detecting that the user's attention is directed toward the corresponding application control affordance for the first application window: Based on determining that the user's attention relative to the corresponding application control enable indication of the first application window satisfies the first criterion, expanding the corresponding application control enable indication to display additional information and / or controls that were not displayed before detecting that the user's attention relative to the corresponding application control enable indication satisfies the first criterion; as well as Based on determining that the user's attention relative to the corresponding application control enable representation of the first application window does not meet the first standard, abandoning expanding the corresponding application control enable representation.

73. A method according to claim 72, wherein the corresponding application control enable representation includes information indicating the content displayed in the first application window.

74. The method of any one of claims 72 to 73, wherein the set of application control affordances includes a first set of one or more privacy indicators for the first application, wherein the first set of one or more privacy indicators has been displayed based on a determination that the first application is accessing one or more sensors of the computer system; and The method comprises: detecting that the first application associated with the first application window no longer has access to the one or more sensors of the computer system while displaying the first application window of the first application with the first set of one or more privacy indicators; as well as Based on determining that the first application no longer has access to the one or more sensors of the computer system, ceasing to display the first set of privacy indicators with the first application window.

75. A method according to claim 74, wherein the first set of one or more privacy indicators includes a first privacy indicator, which has been displayed based on determining that the first application is collecting audio information through at least one of the one or more sensors of the computer system.

76. A method according to any one of claims 74 to 75, wherein the first set of one or more privacy indicators includes a second privacy indicator, which has been displayed based on determining that the first application is collecting location information through at least one of the one or more sensors of the computer system.

77. The method of any one of claims 72 to 76, wherein the set of application control affordances includes a first sharing indicator for the first application, wherein the first sharing indicator has been displayed in a first appearance based on a determination that the first application is sharing content with another device; and The method comprises: detecting that the first application is no longer sharing content with another device while the first application window and the first sharing indicator having the first appearance are displayed simultaneously; as well as Based on determining that the first application is no longer sharing content with another device, ceasing to display the first sharing indicator in the first appearance.

78. The method of claim 77, wherein displaying the first sharing indicator comprises displaying the first sharing indicator proximate the first title bar on the first side of the first application window.

79. The method according to any one of claims 77 to 78, comprising: detecting that the first application is sharing content with another device; In response to detecting that the first application is sharing content with another device: Based on determining that the other device is associated with a first other user, displaying a first identifier corresponding to the first other user in the first sharing indicator; and Based on determining that the other device is associated with a second other user different from the first other user, a second identifier corresponding to the second other user is displayed in the first sharing indicator, the second identifier being different from the first identifier.

80. The method of any one of claims 77 to 79, comprising: detecting that the first application is no longer sharing content with another device; as well as In response to detecting that the first application is no longer sharing content with another device, the first sharing indicator is displayed in a second appearance different from the first appearance to indicate that the first application is not sharing content with another device.

81. The method according to any one of claims 77 to 80, comprising, detecting that the first application is sharing content with another device; and In response to detecting that the first application is sharing content with another device, the first sharing indicator is displayed in the first appearance to indicate that the first application is sharing content with another device.

82. The method of any one of claims 72 to 81, wherein: The first application window is displayed as a two-dimensional object in a three-dimensional environment, and a corresponding application control affordance associated with the first application in the set of application control affordances is displayed on a first side of the two-dimensional object; and The method comprises: displaying, via the first display generating component, a first three-dimensional object and a second set of application control affordances associated with the first three-dimensional object in the three-dimensional environment, wherein: The first three-dimensional object is associated with a third application, the second set of application control affordances associated with the first three-dimensional object being separated from the first three-dimensional object on a second side of the first three-dimensional object, The corresponding application control affordance associated with the first application in the set of application control affordances is offset in a first direction from the first side of the two-dimensional object, and The second set of application control indicators can represent an offset from the second side of the first three-dimensional object in a second direction different from the first direction.

83. The method of any one of claims 68 to 82, comprising: Simultaneously displaying, via the first display generating component, the first application window and a third application control enable representation associated with the first application window, wherein the third application control enable representation is separate from the first application window and on a third side of the first application window; detecting, via the one or more input devices, user input directed to the third application control affordance while the first application window is displayed together with the third application control affordance; as well as In response to detecting the user input directed to the third application control enable representation associated with the first application window, a third application window is displayed based on determining that the user input satisfies selection criteria relative to the third application control enable representation.

84. The method of any one of claims 68 to 83, comprising: In response to detecting that the user's attention is directed to the first title bar and based on determining that the user's attention satisfies the first criterion relative to the first title bar, the first title bar is expanded to display an application icon of the first application associated with the first application window.

85. The method of any one of claims 68 to 84, comprising: The first application window and the first title bar separate from the first application window and one or more selectable controls on the first side of the first application window separate from the first application window and the title bar are simultaneously displayed.

86. A method according to claim 85, wherein the one or more selectable controls displayed on the first side of the first application window separately from the first application window and the first title bar include a first selectable control for enabling or disabling display of a gaze cursor.

87. A method according to any one of claims 85 to 86, wherein the one or more selectable controls displayed on the first side of the first application window separately from the first application window and the title bar include a second selectable control for changing the orientation mode of the first application window between a portrait mode and a landscape mode.

88. The method of any one of claims 85 to 87, wherein the one or more selectable controls displayed on the first side of the first application window separately from the first application window and the first title bar include a set of third selectable controls, wherein respective third selectable controls in the set of third selectable controls correspond to respective sets of window contents of the first application; detecting, via the one or more input devices, a first user input directed to a respective one of the first title bar and the respective third selectable control in the set of third selectable controls; In response to detecting that the first user input is directed to the first title bar and the corresponding one of the corresponding third selectable controls in the set of third selectable controls: Based on determining that the first user input satisfies activation criteria with respect to the first title bar, expanding the first title bar to display one or more selectable controls; and Based on determining that the first user input satisfies the activation criteria with respect to the corresponding third selectable control in the set of third selectable controls, displaying the corresponding set of window contents of the first application corresponding to the corresponding third selectable control.

89. The method of claim 88, comprising: In response to detecting that the first user input is directed to the first title bar and the corresponding one of the corresponding third selectable controls in the set of third selectable controls: Based on determining that the first user input satisfies the activation criteria with respect to the corresponding third selectable control in the set of third selectable controls, a visual appearance of the corresponding third selectable control is changed to indicate that the corresponding third selectable control is selected.

90. The method according to any one of claims 88 to 89, comprising: When the corresponding third selectable control is displayed and the corresponding set of window contents of the first application corresponding to the corresponding third selectable control is displayed, detecting that the user's attention is directed to the corresponding third selectable control; and In response to detecting that the user's attention is directed to the corresponding third selectable control, the corresponding third selectable control is expanded to display one or more selectable controls.

91. The method of any one of claims 88 to 90, comprising: When the first application window is displayed together with the first title bar and the set of third selectable controls, detecting that the user's attention is directed away from the first application window, the first title bar, and the set of third selectable controls; and In response to detecting that the user's attention is directed away from the first application window, the first title bar, and the set of third selectable controls: Display of the first title bar and the set of third selectable controls is visually muted relative to content displayed outside of the first application window, the first title bar, and the set of third selectable controls.

92. The method of any one of claims 88 to 91, comprising: detecting that the first application is accessing one or more sensors of the computer system while displaying the first application window together with the first title bar separate from the first application window and the set of third selectable controls; as well as In response to detecting that the first application is accessing one or more sensors of the computer system, displaying an indication that the first application is accessing the one or more sensors of the computer system, the indication having a first spatial relationship with the first application window and a second spatial relationship with the set of third selectable controls of the first application window.

93. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generating component and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 68 to 92.

94. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 68 to 92.

95. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: A component for carrying out the method according to any one of claims 68 to 92.

96. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a first display generating component and one or more input devices, the one or more programs comprising instructions for: displaying, via the first display generating component, a first application window and a first title bar of the first application window simultaneously, wherein the first title bar of the first application window is separated from the first application window on a first side of the first application window and displays a corresponding identifier of the first application window; When displaying the first application window together with the first title bar separated from the first application window, detecting, via the one or more input devices, that the user's attention is directed toward the first title bar; as well as In response to detecting that the user's attention is directed to the first title bar: Based on determining that the user's attention meets a first criterion relative to the first title bar, the first title bar is expanded to display one or more first selectable controls for interacting with a first application corresponding to the first application window, wherein the one or more first selectable controls are not displayed before the first title bar is expanded.

97. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the first display generating component, a first application window and a first title bar of the first application window simultaneously, wherein the first title bar of the first application window is separated from the first application window on a first side of the first application window and displays a corresponding identifier of the first application window; When displaying the first application window together with the first title bar separated from the first application window, detecting, via the one or more input devices, that the user's attention is directed toward the first title bar; as well as In response to detecting that the user's attention is directed to the first title bar: Based on determining that the user's attention meets a first criterion relative to the first title bar, the first title bar is expanded to display one or more first selectable controls for interacting with a first application corresponding to the first application window, wherein the one or more first selectable controls are not displayed before the first title bar is expanded.

98. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: means for simultaneously displaying, via the first display generating component, a first application window and a first title bar of the first application window, wherein the first title bar of the first application window is separated from the first application window on a first side of the first application window and displays a corresponding identifier of the first application window; means for detecting, via the one or more input devices, a user's attention directed toward the first title bar, enabled when the first application window is displayed together with the first title bar separate from the first application window; and In response to detecting that the user's attention is directed toward the first title bar, means for: Based on determining that the user's attention meets a first criterion relative to the first title bar, the first title bar is expanded to display one or more first selectable controls for interacting with a first application corresponding to the first application window, wherein the one or more first selectable controls are not displayed before the first title bar is expanded.

99. A method comprising: At a computer system in communication with a first display generating component having a first display area and one or more input devices: displaying, via the first display generating component, a first application window of a first application at a first window location in the first display area; Based on determining that the first application is accessing one or more sensors of the computer system, displaying a first indicator in the first display area at a first indicator location having a first spatial relationship to the first application window as an indication that the first application is accessing the one or more sensors of the computer system; while displaying the first indicator at the first indicator location in the first display area having the first spatial relationship with the first application window, detecting a first user input corresponding to a request to move the first application window of the first application to a second window location in the first display area, the second window location being different from the first window location; as well as In response to detecting the first user input corresponding to the request to move the first application window of the first application from the first window location to the second window location in the first display area: displaying the first application window of the first application at the second window location in the first display area; as well as Based on determining that the first application is accessing the one or more sensors of the computer system, the first indicator is displayed in the first display area at a second indicator location that is different from the first indicator location in the first display area, wherein the second indicator location in the first display area has the first spatial relationship with the first application window displayed at the second window location.

100. The method of claim 99, wherein displaying the first indicator in the first display area at the first indicator location having the first spatial relationship with the first application window comprises displaying the first indicator as separate from the first application window.

101. The method according to any one of claims 99 to 100, comprising: displaying, via the first display generating component, a second application window of a second application at a third window location in the first display area; Based on determining that the second application is accessing the one or more sensors of the computer system, a second indicator is displayed at a third indicator location in the first display area having a second spatial relationship with the second application window as an indication that the second application is accessing the one or more sensors of the computer system.

102. The method of any one of claims 99 to 101, wherein displaying the first indicator in the first display area at the first indicator location having the first spatial relationship with the first application window as an indication that the first application is accessing the one or more sensors of the computer system comprises: displaying the first indicator in a first appearance based on determining that the first application is accessing a first sensor of the one or more sensors of the computer system; as well as Based on determining that the first application is accessing a second sensor of the one or more sensors of the computer system that is different from the first sensor, the first indicator is displayed with a second appearance that is different from the first appearance.

103. The method of any one of claims 99 to 102, wherein displaying the first indicator in the first display area at the first indicator location having the first spatial relationship to the first application window as an indication that the first application is accessing the one or more sensors of the computer system comprises: displaying the first indicator in a first appearance when the first application is accessing a first sensor of the one or more sensors of the computer system; detecting that the first application is accessing a second sensor of the computer system that is different from the first sensor; In response to detecting that the first application is accessing the second sensor of the one or more sensors, display of the first indicator is updated from the first appearance to display of the first indicator in a second appearance different from the first appearance.

104. The method according to any one of claims 99 to 103, comprising: detecting that the first application does not have access to the one or more sensors of the computer system when the first indicator is displayed in the first display area at a corresponding indicator location having the first spatial relationship with the first application window; as well as In response to detecting that the first application does not have access to the one or more sensors of the computer system, display of the first indicator at the corresponding indicator location in the first display area having the first spatial relationship with the first application window is ceased.

105. The method of any one of claims 99 to 103, comprising: detecting a gaze input directed to the first indicator located in the first display area at the corresponding indicator location having the first spatial relationship with the first application window while the first indicator is displayed in the first display area at the corresponding indicator location having the first spatial relationship with the first application window; as well as In response to detecting the gaze input directed toward the first indicator, the first indicator is expanded.

106. The method of claim 105, comprising: In response to detecting the gaze input directed to the first indicator, information indicating the one or more sensors being accessed by the first application is displayed after the first indicator has been expanded.

107. The method of any one of claims 99 to 106, comprising: Displaying, via the first display generating component, a first view of a three-dimensional environment, wherein displaying the first indicator at a corresponding window location having the first spatial relationship with the first application window of the first application in the first display area comprises: The first application window and the first indicator of the first application are displayed in the first view of the three-dimensional environment via the first display generation component, wherein the first application window and the first indicator have the first spatial relationship in the three-dimensional environment.

108. The method of any one of claims 99 to 103, comprising: A title bar of the first application window is displayed, wherein the first indicator displayed in the first display area at a corresponding indicator location having the first spatial relationship with the first application window has a second spatial relationship with the title bar.

109. The method of any one of claims 99 to 103, comprising: detecting, via the one or more input devices, that a user's attention is directed toward a portion of the first display area outside of the first application window when the first indicator is displayed in the first display area at a corresponding indicator location having the first spatial relationship with the first application window; as well as In response to detecting that the user's attention is directed to the portion of the first display area outside of the first application window, the first indicator is visually emphasized relative to the first application window in the first display area.

110. The method of claim 109, wherein visually emphasizing the first indicator relative to the first application window in the first display area comprises: reducing a visibility metric of the first application window by a first amount; as well as The visibility metric of the first indicator is maintained or reduced by a second amount, the second amount being less than the first amount.

111. The method according to any one of claims 109 to 110, comprising: displaying one or more user interface objects associated with the first application window when the first indicator is displayed in the first display area at the corresponding indicator location having the first spatial relationship with the first application window; as well as In response to detecting that the user's attention is directed to the portion of the first display area outside of the first application window, the first indicator is visually emphasized relative to the one or more user interface objects associated with the first application window in the first display area.

112. The method according to any one of claims 99 to 111, comprising: Detecting that the user's attention is directed to a first predetermined position in the first display area; In response to detecting that the user's attention is directed toward the first predetermined position in the first display area: Based on determining that one or more applications are accessing the one or more sensors of the computer system, a third indicator is displayed at a fourth indicator location in the first display area as an indication that one or more applications are accessing the one or more sensors of the computer system.

113. The method according to claim 112, comprising: detecting that the user's attention satisfies a first criterion relative to the first predetermined position in the first display area when the third indicator is displayed at the fourth indicator location in the first display area, wherein the first criterion requires that the user's attention has been directed to the first predetermined position for more than a threshold amount of time in order to satisfy the first criterion; and In response to detecting that the user's attention satisfies the first criterion: displaying a user interface element including one or more selectable controls for accessing system functions of the computer system, wherein based on determining that the one or more applications are accessing the one or more sensors of the computer system, The third indicator is displayed simultaneously with the user interface element.

114. The method of claim 113, wherein the user interface element including the one or more selectable controls includes state information corresponding to the computer system, the state information being updated as a state of the computer system changes.

115. The method of any one of claims 112 to 114, wherein displaying the third indicator at the fourth indicator location in the first display area comprises: displaying the third indicator in a first appearance of the third indicator in response to determining that the one or more applications are accessing a first sensor of the one or more sensors of the computer system; as well as Based on determining that the one or more applications are accessing a second sensor of the computer system that is different from the first sensor, the third indicator is displayed with a second appearance of the third indicator that is different from the first appearance of the third indicator.

116. The method according to any one of claims 112 to 115, comprising: detecting that the user's attention satisfies a second criterion relative to the third indicator in the first display area when the third indicator is displayed at the fourth indicator location in the first display area, wherein the second criterion requires that the user's attention has been directed to the third indicator for more than a first threshold amount of time in order to satisfy the second criterion; and In response to detecting that the user's attention satisfies the second criterion with respect to the third indicator in the first display area: Information related to the one or more sensors being accessed by the one or more applications is displayed.

117. The method of claim 116, wherein displaying the information related to the one or more sensors being accessed by the one or more applications comprises: displaying a first indication of a first sensor of the one or more sensors being accessed by the one or more applications; and The method comprises: detecting, while displaying in the first display area the first indication of the first sensor being accessed by the one or more applications, that the user's attention satisfies a third criterion relative to the first indication of the first sensor in the first display area, wherein the third criterion requires that the user's attention has been directed toward the first indication of the first sensor for more than a second threshold amount of time in order to satisfy the third criterion; and In response to detecting that the first indication of the user's attention in the first display area relative to the first sensor satisfies the third criterion: Information related to a first application of the one or more applications that is accessing the first sensor of the one or more sensors is displayed.

118. The method according to any one of claims 99 to 117, comprising: displaying, via the first display generating component, a plurality of application windows simultaneously in the first display area, wherein the plurality of application windows are associated with different applications in a plurality of applications; as well as Based on determining that two or more applications from the plurality of applications are accessing the one or more sensors of the computer system, displaying corresponding indicators to identify the two or more applications from the plurality of applications includes: Respective indicators are displayed in a fourth spatial relationship to respective application windows of the two or more applications as an indication that the two or more applications are accessing the one or more sensors of the computer system.

119. The method of claim 118, comprising: Based on determining that one or more corresponding applications among the multiple applications do not have access to the one or more sensors of the computer system, foregoing displaying the corresponding indicator without identifying the one or more corresponding applications among the multiple applications that do not have access to the one or more sensors of the computer system.

120. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generating component and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 99 to 119.

121. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 99 to 119.

122. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: A component for carrying out the method according to any one of claims 99 to 119.

123. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a first display generating component having a first display area and one or more input devices, the one or more programs comprising instructions for: displaying, via the first display generating component, a first application window of a first application at a first window location in the first display area; Based on determining that the first application is accessing one or more sensors of the computer system, displaying a first indicator in the first display area at a first indicator location having a first spatial relationship to the first application window as an indication that the first application is accessing the one or more sensors of the computer system; while displaying the first indicator at the first indicator location in the first display area having the first spatial relationship with the first application window, detecting a first user input corresponding to a request to move the first application window of the first application to a second window location in the first display area, the second window location being different from the first window location; as well as In response to detecting the first user input corresponding to the request to move the first application window of the first application from the first window location to the second window location in the first display area: displaying the first application window of the first application at the second window location in the first display area; as well as Based on determining that the first application is accessing the one or more sensors of the computer system, the first indicator is displayed in the first display area at a second indicator location different from the first indicator location in the first display area, wherein the second indicator location in the display area has the first spatial relationship with the first application window displayed at the second window location.

124. A computer system in communication with a first display generating component having a first display area and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the first display generating component, a first application window of a first application at a first window location in the first display area; Based on determining that the first application is accessing one or more sensors of the computer system, displaying a first indicator in the first display area at a first indicator location having a first spatial relationship to the first application window as an indication that the first application is accessing the one or more sensors of the computer system; while displaying the first indicator at the first indicator location in the first display area having the first spatial relationship with the first application window, detecting a first user input corresponding to a request to move the first application window of the first application to a second window location in the first display area, the second window location being different from the first window location; as well as In response to detecting the first user input corresponding to the request to move the first application window of the first application from the first window location to the second window location in the first display area: displaying the first application window of the first application at the second window location in the first display area; as well as Based on determining that the first application is accessing the one or more sensors of the computer system, the first indicator is displayed in the first display area at a second indicator location that is different from the first indicator location in the first display area, wherein the second indicator location in the first display area has the first spatial relationship with the first application window displayed at the second window location.

125. A computer system in communication with a first display generating component having a first display area and one or more input devices, the computer system comprising: means for displaying, via the first display generating component, a first application window of a first application at a first window location in the first display area; means enabled in response to determining that the first application is accessing one or more sensors of the computer system for: displaying a first indicator in the first display area at a first indicator location having a first spatial relationship to the first application window as an indication that the first application is accessing the one or more sensors of the computer system; means enabled when the first indicator is displayed in the first display area at the first indicator location having the first spatial relationship with the first application window, for: detecting a first user input corresponding to a request to move the first application window of the first application to a second window location in the first display area, the second window location being different from the first window location; and means for enabling, in response to detecting the first user input corresponding to the request to move the first application window of the first application from the first window location to the second window location in the first display area, displaying the first application window of the first application at the second window location in the first display area; as well as Based on determining that the first application is accessing the one or more sensors of the computer system, the first indicator is displayed in the first display area at a second indicator location that is different from the first indicator location in the first display area, wherein the second indicator location in the first display area has the first spatial relationship with the first application window displayed at the second window location.

126. A method comprising: At a computer system in communication with a first display generating component and one or more input devices: Displaying a user interface, wherein displaying the user interface comprises simultaneously displaying a content area having first content, a first user interface object, and a second user interface object in the user interface, wherein: The corresponding content in the content area is constrained to have an appearance in which the corresponding parameter is within a first range of values; the first user interface object is displayed with an appearance in which the corresponding parameter has a value outside the first range of values; and the second user interface object is displayed with an appearance in which the corresponding parameter has a value outside the first range of values; When the first content, the first user interface object, and the second user interface object are displayed simultaneously, updating the user interface includes: While the corresponding content in the content area continues to be constrained to have an appearance in which the corresponding parameter is within the first value range, changing the first content to second content; updating the appearance of the first user interface object and continuing to display the first user interface object with the appearance in which the corresponding parameter has a value outside of the first range of values; and The appearance of the second user interface object is updated, and the second user interface object continues to be displayed with the appearance in which the corresponding parameter has a value outside of the first range of values.

127. The method of claim 126, wherein: The first value range is a first brightness value range; When the content area having the first content, the first user interface object, and the second user interface object are displayed simultaneously in the user interface: The first user interface object is displayed with an appearance in which the corresponding parameter has a first brightness value outside the first range of brightness values; and The second user interface object is displayed with an appearance in which the corresponding parameter has a second brightness value outside the first brightness value range; and After updating said UI: The first user interface object is displayed with an appearance in which the corresponding parameter has a third brightness value outside of the first range of brightness values; and The second user interface object is displayed with an appearance in which the corresponding parameter has a fourth brightness value outside of the first range of brightness values.

128. The method of claim 126, wherein: The first value range is a first color value range; When the content area having the first content, the first user interface object, and the second user interface object are displayed simultaneously in the user interface: The first user interface object is displayed with an appearance in which the corresponding parameter has a first color value outside of the first range of color values; and the second user interface object being displayed with an appearance in which the corresponding parameter has a second color value outside of the first color value range; After updating said UI: The first user interface object is displayed with an appearance in which the corresponding parameter has a third color value outside of the first range of color values; and The second user interface object is displayed with an appearance in which the corresponding parameter has a fourth color value outside of the first range of color values.

129. The method of any one of claims 126 to 128, wherein: The user interface is an application user interface for an application of the computer system; and The first range of values ​​is selected by the application.

130. The method of any one of claims 126 to 129, comprising: detecting gaze input directed toward the first user interface object; as well as In response to detecting the gaze input directed to the first user interface object, a system user interface for accessing functionality of the computer system is displayed.

131. The method of any one of claims 126 to 130, comprising: detecting a first user input directed to the first user interface object; as well as In response to detecting the first user input directed to the first user interface object, a first operation is performed with respect to the user interface.

132. The method of any one of claims 126 to 131, comprising: prior to simultaneously displaying the content area having the first content, the first user interface object, and the second user interface object, displaying the user interface without displaying at least one of the first user interface object and the second user interface object; detecting a gaze input satisfying a first criterion while displaying the user interface without displaying the at least one of the first user interface object and the second user interface object; as well as In response to detecting the gaze input that satisfies the first criterion, updating the user interface to simultaneously display the content area having the first content, the first user interface object, and the second user interface object.

133. The method of any one of claims 126 to 132, wherein: At least one of the first user interface object and the second user interface object is a closed affordance corresponding to a first window displayed via the first display generating component; The method comprises: detecting a second user input directed to the at least one of the first user interface object and the second user interface object while the content area and the at least one of the first user interface object and the second user interface object are displayed simultaneously; as well as In response to detecting the second user input directed to the at least one of the first user interface object and the second user interface object, ceasing to display the first window via the first display generating component.

134. A method according to any one of claims 126 to 133, wherein at least one of the first user interface object and the second user interface object includes text displayed as an overlay on a representation of a physical environment.

135. A method according to claim 134, wherein the corresponding content in the content area includes the representation of the physical environment, and wherein the representation of the physical environment is constrained to have an appearance in which the corresponding parameter is within the first value range.

136. A method according to any one of claims 134 to 135, wherein the text includes a title of an object that is visible via the first display generating component simultaneously with the text and the representation of the physical environment.

137. A method according to any one of claims 134 to 136, wherein the text comprises descriptive text of media content displayed via the first display generating component as an overlay on the representation of the physical environment.

138. A method according to any one of claims 126 to 137, wherein the user interface includes a view of a three-dimensional environment, and at least one of the first user interface object and the second user interface object is displayed as an overlay on the content area as a virtual augmentation of the corresponding content in the content area.

139. A method according to any one of claims 126 to 138, wherein at least one of the first user interface object and the second user interface object is a visual indicator that indicates the location of the user's gaze.

140. The method of claim 139, wherein: The user interface includes two or more control elements for performing functions within the user interface; and The method comprises: detecting that the gaze of the user is directed to a corresponding control element of the two or more control elements; and In response to detecting that the gaze of the user is directed to the corresponding control element of the two or more control elements: Based on determining that the gaze of the user is directed toward a first control element of the two or more control elements, displaying the visual indicator on at least a portion of the first control element; and Based on determining that the gaze of the user is directed toward a second control element of the two or more control elements, the visual indicator is displayed on at least a portion of the second control element.

141. The method of any one of claims 139 to 140, wherein the user interface comprises application content of a first application; and The method comprises: detecting that the gaze of the user is directed toward the application content of the first application; as well as In response to detecting that the gaze of the user is directed toward the application content of the first application: displaying the visual indicator on a first portion of the application content based on determining that the gaze of the user is directed toward the first portion of the application content; as well as Based on determining that the gaze of the user is directed toward a second portion of the application content, displaying the visual indicator on the second portion of the application content.

142. The method of any one of claims 139 to 141, wherein displaying the visual indicator indicating the location of the gaze of the user comprises: A virtual lighting effect is applied to original content displayed at the location of the visual indicator, wherein the original content is constrained to have an appearance in which the corresponding parameter is within the first range of values.

143. The method of any one of claims 139 to 141, wherein displaying the visual indicator indicating the location of the gaze of the user comprises: A feathering effect is applied to at least a portion of a boundary between the visual indicator and original content displayed at the location of the visual indicator, wherein the original content is constrained to have an appearance in which the corresponding parameter is within the first range of values.

144. The method of any one of claims 139 to 141, wherein displaying the visual indicator indicating the location of the gaze of the user comprises: The visual indicator is made partially transparent to reveal at least some visual characteristics of original content displayed at the location of the visual indicator, wherein the original content is constrained to have an appearance in which the corresponding parameter is within the first range of values.

145. The method of any one of claims 126 to 144, wherein: When the content area having the second content, the first user interface object, and the second user interface object are displayed simultaneously in the user interface, further updating the user interface comprises: While the corresponding content in the content area continues to be constrained to have an appearance in which the corresponding parameter is within the first value range, changing the second content to a third content; updating the appearance of at least one of the first user interface object and the second user interface object in response to determining that the third content is to be displayed with an appearance in which the corresponding parameter is within a first sub-range of the first value range, and continuing to display the at least one of the first user interface object and the second user interface object with an appearance in which the corresponding parameter has a value outside of the first value range; and Based on determining that the third content is to be displayed with an appearance in which the corresponding parameter is within a second sub-range of the first value range that is different from the first sub-range of the first value range, the appearance of at least one of the first user interface object and the second user interface object is updated, and the first user interface object and at least one of the second user interface objects are displayed with an appearance in which the corresponding parameter is constrained to have a value within the first value range.

146. A method according to any one of claims 126 to 145, wherein the first user interface object is displayed as having one or more first values ​​of the corresponding parameter outside the first value range in one or more first parts of the first user interface object, and the first user interface object is displayed as having one or more second values ​​of the corresponding parameter within the first value range in one or more second parts of the first user interface object.

147. A method according to claim 146, wherein the one or more first portions of the first user interface object include a first appearance change caused by a simulated lighting effect applied to the first user interface object.

148. A method according to any one of claims 146 to 147, wherein the one or more first portions of the first user interface object include a second appearance change caused by a visual indication of the location of the user's gaze on the one or more first portions of the first user interface object.

149. A method according to any one of claims 146 to 148, wherein the one or more first portions of the first user interface object include control elements for performing a first operation relative to the first user interface object.

150. The method of claim 149, comprising: A first view of a three-dimensional environment is displayed, wherein displaying the first view of the three-dimensional environment includes simultaneously displaying the content area having the first content, the first user interface object, and the second user interface object in the three-dimensional environment.

151. The method of any one of claims 126 to 150, comprising: detecting a state change of a second user interface when the first user interface object is displayed with the appearance in which the corresponding parameter has a corresponding value outside the first value range, wherein the state change of the second user interface causes the second user interface to cover a corresponding portion of the first user interface object; In response to detecting the change in state of the second user interface, displaying the corresponding portion of the first user interface object with an appearance in which the corresponding parameter has a different value than the corresponding value outside of the first range of values.

152. A method according to claim 151, wherein the state change of the second user interface causes the second user interface to, in addition to displaying the corresponding portion of the first user interface object with the appearance in which the corresponding parameter has the corresponding value outside the first value range, at least partially cover the corresponding content in the content area that is constrained to have an appearance in which the corresponding parameter is within the first value range.

153. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generating component and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 126 to 152.

154. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 126 to 152.

155. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: A component for carrying out the method according to any one of claims 126 to 152.

156. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a first display generating component and one or more input devices, the one or more programs comprising instructions for: Displaying a user interface, wherein displaying the user interface comprises simultaneously displaying a content area having first content, a first user interface object, and a second user interface object in the user interface, wherein: The corresponding content in the content area is constrained to have an appearance in which the corresponding parameter is within a first range of values; the first user interface object is displayed with an appearance in which the corresponding parameter has a value outside the first range of values; and the second user interface object is displayed with an appearance in which the corresponding parameter has a value outside the first range of values; When the first content, the first user interface object, and the second user interface object are displayed simultaneously, updating the user interface includes: While the corresponding content in the content area continues to be constrained to have an appearance in which the corresponding parameter is within the first value range, changing the first content to second content; updating the appearance of the first user interface object and continuing to display the first user interface object with the appearance in which the corresponding parameter has a value outside of the first range of values; and The appearance of the second user interface object is updated, and the second user interface object continues to be displayed with the appearance in which the corresponding parameter has a value outside of the first range of values.

157. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: Displaying a user interface, wherein displaying the user interface comprises simultaneously displaying a content area having first content, a first user interface object, and a second user interface object in the user interface, wherein: The corresponding content in the content area is constrained to have an appearance in which the corresponding parameter is within a first range of values; the first user interface object is displayed with an appearance in which the corresponding parameter has a value outside the first range of values; and the second user interface object is displayed with an appearance in which the corresponding parameter has a value outside the first range of values; When the first content, the first user interface object, and the second user interface object are displayed simultaneously, updating the user interface includes: While the corresponding content in the content area continues to be constrained to have an appearance in which the corresponding parameter is within the first value range, changing the first content to second content; updating the appearance of the first user interface object and continuing to display the first user interface object with the appearance in which the corresponding parameter has a value outside of the first range of values; and The appearance of the second user interface object is updated, and the second user interface object continues to be displayed with the appearance in which the corresponding parameter has a value outside of the first range of values.

158. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: Means for displaying a user interface, wherein displaying the user interface comprises simultaneously displaying a content area having first content, a first user interface object, and a second user interface object in the user interface, wherein: The corresponding content in the content area is constrained to have an appearance in which the corresponding parameter is within a first range of values; the first user interface object is displayed with an appearance in which the corresponding parameter has a value outside the first range of values; and the second user interface object is displayed with an appearance in which the corresponding parameter has a value outside the first range of values; A component for updating the user interface enabled when the first content, the first user interface object, and the second user interface object are displayed simultaneously, the component comprising: means for: while the corresponding content in the content area continues to be constrained to have an appearance in which the corresponding parameter is within the first range of values, changing the first content to second content; means for updating the appearance of the first user interface object and continuing to display the first user interface object with the appearance wherein the corresponding parameter has a value outside of the first range of values; and Means for updating the appearance of the second user interface object and continuing to display the second user interface object with the appearance wherein the corresponding parameter has a value outside of the first range of values.

159. A method comprising: At a computer system in communication with a first display generating component and one or more input devices: displaying, via the first display generating component, a first view of the three-dimensional environment corresponding to a first viewpoint of a user; while displaying, via the first display generating component, the first view of the three-dimensional environment corresponding to the first viewpoint of the user, detecting a first event corresponding to a request to display a first virtual object in the first view of the three-dimensional environment; in response to detecting the first event corresponding to a request to display the first virtual object in the first view of the three-dimensional environment, displaying the first virtual object in the first view of the three-dimensional environment at a first location in the three-dimensional environment, wherein the first virtual object is displayed with a first object management user interface corresponding to the first virtual object, and wherein the first object management user interface has a first appearance relative to the first virtual object at the first location in the three-dimensional environment; detecting, via the one or more input devices, a first user input corresponding to a request to move the first virtual object in the three-dimensional environment; as well as In response to detecting the first user input corresponding to a request to move the first virtual object in the three-dimensional environment: The first virtual object is displayed in the first view of the three-dimensional environment at a second position in the three-dimensional environment that is different from the first position, wherein the first virtual object is displayed simultaneously with the first object management user interface at the second position in the three-dimensional environment, and wherein the first object management user interface has a second appearance relative to the first virtual object that is different from the first appearance.

160. The method of claim 159, wherein the first user input is detected while the first virtual object and the first object management user interface are displayed simultaneously, wherein the first object management user interface has the first appearance relative to the first virtual object.

161. The method of any one of claims 159 to 160, wherein: displaying the first virtual object at the first location with the first object management user interface having the first appearance relative to the first virtual object comprises displaying the first virtual object at a first orientation relative to the first viewpoint of the user and displaying the first object management user interface at a second orientation relative to the first virtual object; and Displaying the first virtual object at the second location together with the first object management user interface having the second appearance relative to the first virtual object includes displaying the first virtual object in a third orientation different from the first orientation relative to the first viewpoint of the user and displaying the first object management user interface in a fourth orientation different from the second orientation relative to the first virtual object.

162. The method of any one of claims 159 to 161, wherein: displaying the first object management user interface having the first appearance relative to the first virtual object comprises displaying the first object management user interface in a corresponding orientation facing the first viewpoint; and Displaying the first object management user interface having the second appearance relative to the first virtual object includes: when the first object management user interface is displayed with the first virtual object displayed at the second position in the three-dimensional environment, displaying the first object management user interface with the corresponding orientation facing the first viewpoint.

163. The method of any one of claims 159 to 162, wherein: displaying the first virtual object at the first location with the first object management user interface having the first appearance relative to the first virtual object includes displaying the first virtual object at a first size and displaying the first object management user interface at a second size; and Displaying the first virtual object at the second location with the first object management user interface having the second appearance relative to the first virtual object includes displaying the first virtual object at the first size and displaying the first object management user interface at a third size different from the second size.

164. The method of claim 163, wherein: Displaying the first virtual object in the first size and displaying the first object management user interface in the third size includes: In response to determining that the second location is further from the first viewpoint than the first location, increasing the size of the first object management user interface; and Based on determining that the second location is closer to the first viewpoint than the first location, the size of the first object management user interface is reduced.

165. The method of any one of claims 163 to 164, wherein displaying the first virtual object at the first size at the second location comprises: Based on determining that the second location is further away from the first viewpoint than the first location, displaying the first virtual object at a reduced display size without changing the size of the first virtual object relative to the three-dimensional environment; and Based on determining that the second location is closer to the first viewpoint than the first location, the first virtual object is displayed at an increased display size without changing the size of the first virtual object relative to the three-dimensional environment.

166. The method of any one of claims 159 to 164, comprising: In response to detecting the first user input corresponding to a request to move the first virtual object in the three-dimensional environment, the first object management user interface is moved through a plurality of intermediate positions according to a current position of the first virtual object between the first position and the second position, and the appearance of the first object management user interface is updated through a plurality of intermediate appearances relative to the first virtual object according to the current position of the first virtual object between the first position and the second position.

167. The method of any one of claims 159 to 164, comprising: In response to detecting the first user input corresponding to a request to move the first virtual object in the three-dimensional environment: based on determining that the first virtual object is moving in response to the first user input, reducing the visual prominence of the first object management user interface relative to the first virtual object when the first virtual object is moving; as well as Based on determining that the first virtual object has maintained its position for at least a threshold amount of time, restoring the visual prominence of the first object management user interface relative to the first virtual object while the first virtual object maintains its position.

168. The method of any one of claims 159 to 167, comprising: before detecting the first user input, displaying the first virtual object at the first location in the three-dimensional environment without concurrently displaying the first object management user interface; When the first virtual object is displayed at the first location without simultaneously displaying the first object management user interface, detecting that the user's attention is directed toward a first portion of the first virtual object; as well as In response to detecting that the user's attention is directed to the first portion of the first virtual object, based on determining that the user's attention meets a first criterion relative to the first portion of the first virtual object, the first object management user interface is displayed together with the first virtual object at the first position in the three-dimensional environment.

169. The method of any one of claims 159 to 168, wherein displaying the first virtual object with the first object management user interface comprises one or more of: displaying a corresponding backplane user interface object, the corresponding backplane user interface object providing a reference surface on which the first virtual object is placed, displaying a corresponding movement-enabled representation that, when dragged, causes the computer system to move the first virtual object along with the corresponding movement-enabled representation, displaying a corresponding close enable indication, which, when activated, causes the computer system to close the first virtual object and stop displaying the first virtual object and its associated object management user interface in the three-dimensional environment, displaying a corresponding resizable affordance that, when dragged, causes the computer system to change the size of the first virtual object relative to the three-dimensional environment, displaying a corresponding title bar, wherein the corresponding title bar displays a title of the first virtual object, and A corresponding object menu is displayed that, when selected, displays a plurality of selectable options corresponding to different operations that can be performed with respect to the first virtual object.

170. The method of claim 169, wherein: The first object management user interface includes the corresponding backplane user interface object; and Displaying the corresponding bottom plate user interface object includes displaying a simulated reflection and / or a simulated shadow of the first virtual object on the corresponding bottom plate user interface object according to a spatial relationship between the first virtual object and the corresponding bottom plate user interface object.

171. The method of claim 169, wherein: The first object management user interface includes the corresponding resizing affordance; When the first virtual object is displayed at the first location, the corresponding resizing affordance is displayed at a first depth relative to the first viewpoint in the three-dimensional environment; and When the first virtual object is displayed at the second location, the corresponding resized indication can be represented as being displayed at a second depth different from the first depth relative to the first viewpoint in the three-dimensional environment.

172. The method of any one of claims 159 to 171, comprising: while displaying the first virtual object at a corresponding location in the three-dimensional environment together with the first object management user interface having a corresponding appearance relative to the first virtual object, detecting a first movement of the user's current viewpoint from the first viewpoint to a second viewpoint different from the first viewpoint; and In response to detecting the first movement of the current viewpoint of the user from the first viewpoint to the second viewpoint, displaying in a second view of the three-dimensional environment corresponding to the second viewpoint of the user, includes displaying the first virtual object at the corresponding position in the three-dimensional environment together with the first object management user interface having the corresponding appearance relative to the first virtual object.

173. The method of claim 172, comprising: detecting a second user input corresponding to a request to move the first virtual object in the three-dimensional environment while displaying the first virtual object together with the first object management user interface having the corresponding appearance relative to the first virtual object at a corresponding location in the second view of the three-dimensional environment corresponding to the second viewpoint of the user; and In response to detecting the second user input corresponding to a request to move the first virtual object in the three-dimensional environment: The first virtual object is displayed at a third location in the three-dimensional environment that is different from the corresponding location in the second view of the three-dimensional environment, wherein the first virtual object is displayed at the third location together with the first object management user interface having a third appearance relative to the first virtual object that is different from the corresponding appearance.

174. The method of any one of claims 159 to 173, comprising: while displaying the first virtual object, detecting, via the one or more input devices, a first gaze input that satisfies a first criterion, wherein the first criterion requires that the first gaze input be directed toward a first portion of the three-dimensional environment in order to satisfy the first criterion; and In response to detecting that the first gaze input satisfies the first criterion, displaying an off enable indication for the first virtual object, wherein the off enable indication is not displayed before detecting that the first gaze input satisfies the first criterion.

175. The method of claim 174, comprising: prior to displaying the off affordance in response to detecting the first gaze input satisfying the first criterion, displaying the first virtual object concurrently with a preview of the move affordance and the off affordance for the first virtual object; and In response to detecting that the first gaze input satisfies the first criterion, display of the preview of the off affordance is replaced with display of the off affordance.

176. The method of any one of claims 174 to 175, comprising: detecting, via the one or more input devices, user input directed toward the close affordance while the close affordance is displayed with the first virtual object; as well as In response to detecting the user input directed toward the close enable indication, display of the first virtual object is stopped based on determining that the user input directed toward the close enable indication meets a close criteria.

177. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a first display generating component and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 159 to 176.

178. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 159 to 176.

179. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: A component for carrying out the method according to any one of claims 159 to 176.

180. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a first display generating component and one or more input devices, the one or more programs comprising instructions for: displaying, via the first display generating component, a first view of the three-dimensional environment corresponding to a first viewpoint of a user; while displaying, via the first display generating component, the first view of the three-dimensional environment corresponding to the first viewpoint of the user, detecting a first event corresponding to a request to display a first virtual object in the first view of the three-dimensional environment; in response to detecting the first event corresponding to a request to display the first virtual object in the first view of the three-dimensional environment, displaying the first virtual object in the first view of the three-dimensional environment at a first location in the three-dimensional environment, wherein the first virtual object is displayed with a first object management user interface corresponding to the first virtual object, and wherein the first object management user interface has a first appearance relative to the first virtual object at the first location in the three-dimensional environment; detecting, via the one or more input devices, a first user input corresponding to a request to move the first virtual object in the three-dimensional environment; as well as In response to detecting the first user input corresponding to the request to move the first virtual object in the three-dimensional environment: The first virtual object is displayed in the first view of the three-dimensional environment at a second position in the three-dimensional environment that is different from the first position, wherein the first virtual object is displayed simultaneously with the first object management user interface at the second position in the three-dimensional environment, and wherein the first object management user interface has a second appearance relative to the first virtual object that is different from the first appearance.

181. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the first display generating component, a first view of the three-dimensional environment corresponding to a first viewpoint of a user; while displaying, via the first display generating component, the first view of the three-dimensional environment corresponding to the first viewpoint of the user, detecting a first event corresponding to a request to display a first virtual object in the first view of the three-dimensional environment; in response to detecting the first event corresponding to a request to display the first virtual object in the first view of the three-dimensional environment, displaying the first virtual object in the first view of the three-dimensional environment at a first location in the three-dimensional environment, wherein the first virtual object is displayed with a first object management user interface corresponding to the first virtual object, and wherein the first object management user interface has a first appearance relative to the first virtual object at the first location in the three-dimensional environment; detecting, via the one or more input devices, a first user input corresponding to a request to move the first virtual object in the three-dimensional environment; as well as In response to detecting the first user input corresponding to a request to move the first virtual object in the three-dimensional environment: The first virtual object is displayed in the first view of the three-dimensional environment at a second position in the three-dimensional environment that is different from the first position, wherein the first virtual object is displayed simultaneously with the first object management user interface at the second position in the three-dimensional environment, and wherein the first object management user interface has a second appearance relative to the first virtual object that is different from the first appearance.

182. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: means for displaying, via the first display generating component, a first view of the three-dimensional environment corresponding to a first viewpoint of a user; means for detecting a first event corresponding to a request to display a first virtual object in the first view of the three-dimensional environment, enabled when displaying the first view of the three-dimensional environment corresponding to the first viewpoint of the user via the first display generating component; means for displaying the first virtual object at a first location in the three-dimensional environment in the first view of the three-dimensional environment, enabled in response to detecting the first event corresponding to a request to display the first virtual object in the first view of the three-dimensional environment, wherein the first virtual object is displayed with a first object management user interface corresponding to the first virtual object, and wherein the first object management user interface has a first appearance relative to the first virtual object at the first location in the three-dimensional environment; means for detecting, via the one or more input devices, a first user input corresponding to a request to move the first virtual object in the three-dimensional environment; and In response to detecting the first user input corresponding to a request to move the first virtual object in the three-dimensional environment, means for: The first virtual object is displayed in the first view of the three-dimensional environment at a second position in the three-dimensional environment that is different from the first position, wherein the first virtual object is displayed simultaneously with the first object management user interface at the second position in the three-dimensional environment, and wherein the first object management user interface has a second appearance relative to the first virtual object that is different from the first appearance.

183. A method comprising: At a computer system in communication with one or more display generating components and one or more input devices: When simultaneously displaying a user interface of a first application and a close enable indication associated with the user interface of the first application via the one or more display generating components, detecting a first input directed to the close enable indication; as well as In response to detecting the first input: Based on determining that the first input is a first type of input, a first option for closing an application other than the first application is displayed.

184. The method of claim 183, wherein the first input is detected while the user interface of the first application and one or more user interfaces for one or more applications other than the first application are displayed simultaneously, and the method comprises: detecting a second input directed to the first option for closing an application other than the first application; as well as In response to detecting the second input, ceasing to display the one or more user interfaces for the one or more applications other than the first application.

185. The method of claim 184, wherein determining that the first input is the first type of input comprises determining that the first input comprises a gesture.

186. The method of any one of claims 184 to 185, comprising: In response to detecting the first input: Based on determining that the first input is a second type of input different from the first type of input, closing the first application includes stopping displaying the user interface of the first application.

187. The method of claim 186, wherein: The criteria for detecting the first type of input include a requirement to satisfy a corresponding set of one or more criteria including a time-based criteria in order to detect the first type of input, and The criteria for detecting the second type of input does not include the requirement of satisfying the corresponding set of one or more criteria including the time-based criteria in order to detect the second type of input.

188. The method of any one of claims 183 to 187, comprising: In response to detecting the first input: Based on determining that the first input is the first type of input, a second option for closing only the first application is displayed.

189. The method of any one of claims 183 to 188, comprising: detecting an input corresponding to a request to close one or more user interfaces of one or more applications; as well as In response to detecting the input corresponding to a request to close the one or more user interfaces of the one or more application programs: closing the one or more user interfaces of the one or more applications; as well as Based on determining that the corresponding criteria are met, a main menu user interface is displayed at a first position.

190. A method according to any one of claims 183 to 189, wherein the close enable indication associated with the user interface of the first application is displayed within a threshold distance of the user interface of the first application.

191. A method according to any one of claims 183 to 190, wherein the close indicator associated with the user interface of the first application is represented by an application grabber user interface element displayed near the user interface of the first application.

192. The method of any one of claims 183 to 191, comprising: detecting that the user's attention is directed toward the closing affordance; as well as In response to detecting that the user's attention is directed toward the close enable indication, increasing the size of the close enable indication.

193. The method of any one of claims 183 to 192, comprising: In response to detecting the first input: Based on determining that the first input is a second type of input that is different from the first type of input, a second option for closing a plurality of application user interfaces associated with the first application is displayed.

194. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more display generating components and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 183 to 193.

195. A computer system in communication with one or more display generating components and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 183 to 193.

196. A computer system in communication with one or more display generating components and one or more input devices, the computer system comprising: A component for carrying out the method according to any one of claims 183 to 193.

197. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for: While simultaneously displaying a user interface of a first application and a close affordance associated with the user interface of the first application via the one or more display generating components, detecting a first input directed to the close affordance; and In response to detecting the first input: Based on determining that the first input is a first type of input, a first option for closing an application other than the first application is displayed.

198. A computer system in communication with one or more display generating components and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: When simultaneously displaying a user interface of a first application and a close enable indication associated with the user interface of the first application via the one or more display generating components, detecting a first input directed to the close enable indication; as well as In response to detecting the first input: Based on determining that the first input is a first type of input, a first option for closing an application other than the first application is displayed.

199. A computer system in communication with one or more display generating components and one or more input devices, the computer system comprising: means for detecting a first input directed to a close affordance associated with a user interface of a first application when the user interface of the first application and the close affordance are simultaneously displayed via the one or more display generating components; and In response to detecting the first input, means is enabled, the means comprising: Means for displaying a first option for closing an application other than the first application is enabled based on determining that the first input is an input of a first type.

200. A method comprising: At a computer system in communication with a first display generating component and one or more input devices: displaying, via the first display generating component, a first object at a first location in a first view of a three-dimensional environment; when displaying the first object at the first location in the first view of the three-dimensional environment via the first display generating component, displaying a first set of one or more control objects, wherein respective control objects in the first set of one or more control objects correspond to respective operations applicable to the first object; detecting, via the one or more input devices, a first user input corresponding to a request to move the first object in the three-dimensional environment; as well as In response to detecting the first user input corresponding to a request to move the first object in the three-dimensional environment: moving the first object from the first location to a second location; as well as When the first object is moved from the first position to the second position, at least one control object of the first group of one or more control objects corresponding to the respective operation applicable to the first object is visually weakened relative to the first object.

201. The method of claim 200, wherein the at least one control object of the first set of one or more control objects that is visually weakened during movement of the first object from the first position to the second position comprises an off affordance, and the method comprises: detecting a second user input selecting the close affordance while the close affordance is displayed with the first object; as well as In response to detecting the second user input selecting the off affordance, ceasing to display the first object in the three-dimensional environment.

202. The method of claim 200, wherein the at least one control object of the first set of one or more control objects that is visually weakened during movement of the first object from the first position to the second position comprises a resize affordance, and the method comprises: detecting a third user input of selecting and dragging the resizable affordance while the resizable affordance is displayed with the first object; as well as In response to detecting the third user input of selecting and dragging the resizing affordance, the first object is resized according to the third user input.

203. The method of any one of claims 200 to 202, wherein: Visually weakening the at least one control object of the first group of one or more control objects relative to the first object includes visually weakening a first corresponding control object corresponding to a first corresponding operation that can be applied to the first object and a second corresponding control object corresponding to a second corresponding operation that can be applied to the first object.

204. The method of any one of claims 200 to 203, wherein visually weakening the at least one control object in the first set of one or more control objects relative to the first object during the movement of the first object comprises: A fourth corresponding control object corresponding to a fourth corresponding operation applicable to the first object is visually weakened, while a third corresponding control object corresponding to a third corresponding operation applicable to the first object is not visually weakened.

205. The method according to any one of claims 200 to 204, comprising: detecting a fourth user input selecting a fifth corresponding control object in the first set of one or more control objects while displaying the first object together with the first object; as well as In response to detecting the fourth user input selecting the fifth corresponding control object in the first group of one or more control objects, displaying the multiple control options based on determining that the fifth corresponding control object is associated with multiple control options for performing corresponding operations relative to the first object.

206. The method of any one of claims 200 to 205, wherein visually weakening the at least one control object in the first set of one or more control objects when the first object is moved comprises: A sixth corresponding control object of the at least one control object of the first group of one or more control objects is displayed, the sixth corresponding control object having a smaller size relative to a size of the sixth corresponding control object before the first object begins to move.

207. The method of any one of claims 200 to 206, wherein visually weakening the at least one control object in the first set of one or more control objects when the first object is moved comprises: During at least a portion of the movement of the first object, a seventh corresponding control object of the at least one control object in the first group of one or more control objects ceases to be displayed.

208. A method according to any one of claims 200 to 207, wherein detecting the first user input corresponding to a request to move the first object in the three-dimensional environment includes detecting an air gesture directed to the first object, followed by movement of the air gesture.

209. The method according to any one of claims 200 to 208, comprising: When a first control object of the first group of one or more control objects is displayed together with the first object, wherein the first control object corresponds to a first operation applicable to the first object, detecting that user attention is directed to a corresponding portion of the first object corresponding to a second operation applicable to the first object that is different from the first operation; as well as In combination with detecting that the user's attention is directed toward the corresponding part of the first object corresponding to the second operation that can be applied to the first object, based on determining that the user's attention meets the attention standard, the first control object corresponding to the first operation that can be applied to the first object is visually weakened relative to the first object.

210. The method of claim 209, comprising: In response to detecting that the user's attention is directed to the corresponding part of the first object corresponding to the second operation that can be applied to the first object, based on determining that the user's attention meets the attention standard, a second control object different from the first control object and corresponding to the second operation that can be applied to the first object is displayed.

211. The method according to any one of claims 200 to 210, comprising: When a third control object in the first group of one or more control objects is displayed together with the first object, wherein the third control object corresponds to a third operation applicable to the first object, detecting that the current user attention is directed to the third control object; as well as In conjunction with detecting that the current user attention is directed toward the third control object, based on determining that the current user attention satisfies a second attention criterion, the third control object corresponding to the third operation applicable to the first object is visually emphasized relative to the first object.

212. The method according to claim 211, comprising: In combination with detecting that the current user attention is directed to the third control object, based on determining that the current user attention satisfies the second attention standard, a fourth control object in the first group of one or more control objects is visually weakened relative to the first object, wherein the fourth control object corresponds to a fourth operation that is different from the third operation and can be applied to the first object.

213. The method of claim 212, wherein visually weakening the fourth control object relative to the first object comprises reducing the size of the third control object relative to the first object.

214. The method of any one of claims 212 to 213, wherein: Visually weakening the fourth control object relative to the first object includes reducing a spatial extent of the fourth control object such that the fourth control object does not intersect a reaction area of ​​the third control object.

215. The method according to any one of claims 200 to 214, comprising: detecting that the corresponding user attention is directed to a fifth control object in the first group of one or more control objects; as well as In response to detecting that the corresponding user attention is directed to the fifth control object: visually de-emphasize the fifth control object relative to the first object based on determining that the corresponding user attention satisfies a third attention criterion, wherein the third attention criterion requires that the corresponding user attention has been directed toward the fifth control object for at least a threshold amount of time; as well as Based on determining that the corresponding user attention does not meet the third attention criterion, visually emphasizing the fifth control object relative to the first object is abandoned.

216. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a first display generating component and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 200 to 215.

217. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 200 to 215.

218. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: A component for carrying out the method according to any one of claims 200 to 215.

219. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a first display generating component and one or more input devices, the one or more programs comprising instructions for: displaying, via the first display generating component, a first object at a first location in a first view of a three-dimensional environment; when displaying the first object at the first location in the first view of the three-dimensional environment via the first display generating component, displaying a first set of one or more control objects, wherein respective control objects in the first set of one or more control objects correspond to respective operations applicable to the first object; detecting, via the one or more input devices, a first user input corresponding to a request to move the first object in the three-dimensional environment; as well as In response to detecting the first user input corresponding to a request to move the first object in the three-dimensional environment: moving the first object from the first location to a second location; as well as When the first object is moved from the first position to the second position, at least one control object of the first group of one or more control objects corresponding to the respective operation applicable to the first object is visually weakened relative to the first object.

220. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the first display generating component, a first object at a first location in a first view of a three-dimensional environment; when displaying the first object at the first location in the first view of the three-dimensional environment via the first display generating component, displaying a first set of one or more control objects, wherein respective control objects in the first set of one or more control objects correspond to respective operations applicable to the first object; detecting, via the one or more input devices, a first user input corresponding to a request to move the first object in the three-dimensional environment; as well as In response to detecting the first user input corresponding to a request to move the first object in the three-dimensional environment: moving the first object from the first location to a second location; as well as When the first object is moved from the first position to the second position, at least one control object of the first group of one or more control objects corresponding to the respective operation applicable to the first object is visually weakened relative to the first object.

221. A computer system in communication with a first display generating component and one or more input devices, the computer system comprising: means for displaying, via the first display generating component, a first object at a first location in a first view of a three-dimensional environment; means for displaying a first set of one or more control objects enabled when displaying the first object at the first location in the first view of the three-dimensional environment via the first display generating component, wherein respective control objects in the first set of one or more control objects correspond to respective operations applicable to the first object; means for detecting, via the one or more input devices, a first user input corresponding to a request to move the first object in the three-dimensional environment; and In response to detecting the first user input corresponding to a request to move the first object in the three-dimensional environment, means for: moving the first object from the first location to a second location; as well as When the first object is moved from the first position to the second position, at least one control object of the first group of one or more control objects corresponding to the respective operation applicable to the first object is visually weakened relative to the first object.

222. A method comprising: At a computer system in communication with a display generating component and one or more input devices: while displaying, via the display generation component, a first application user interface at a first location in a three-dimensional environment, detecting, at a first time, via the one or more input devices, a first input corresponding to a request to close the first application user interface; as well as In response to detecting the first input corresponding to a request to close the first application user interface: Closing the first application user interface includes stopping displaying the first application user interface in the three-dimensional environment; as well as Based on determining that the corresponding criteria are met, a main menu user interface is displayed at a corresponding main menu location determined based on the first position of the first application user interface in the three-dimensional environment.

223. The method of claim 222, wherein displaying the main menu user interface at the corresponding main menu location determined based on the first position of the first application user interface in the three-dimensional environment comprises: Based on determining that the first application user interface is displayed at the first application location in the three-dimensional environment, displaying the main menu user interface at the first main menu location in the three-dimensional environment; as well as Based on determining that the first application user interface is displayed at a second application location in the three-dimensional environment that is different from the first application location, the main menu user interface is displayed at a second main menu location in the three-dimensional environment that is different from the first main menu location in the three-dimensional environment.

224. The method according to any one of claims 222 to 223, further comprising: In response to detecting the first input corresponding to the request to close the first application user interface and based on determining that the corresponding criteria are not met due to one or more other application user interfaces being open in the three-dimensional environment, abandoning displaying the main menu user interface in the three-dimensional environment.

225. The method of any one of claims 222 to 224, further comprising: determining the corresponding main menu location based on corresponding positions of the plurality of application user interfaces including the first application user interface in accordance with determining that the plurality of application user interfaces including the first application user interface cease to be displayed within a first time threshold from the first time; as well as Based on determining that the plurality of application user interfaces have not ceased to be displayed within the first time threshold from the first time, the corresponding main menu location is determined based on the first position of the first application user interface.

226. The method of claim 225, wherein displaying the main menu user interface at the corresponding main menu location comprises displaying the main menu user interface at the corresponding main menu location when a plurality of application user interfaces including a first plurality of application user interfaces and a second plurality of application user interfaces are closed within a threshold amount of time, wherein: the corresponding main menu location being based on the location of the first plurality of application user interfaces based on determining that the user's attention was more recently directed to the first plurality of application user interfaces than to other application user interfaces outside the first plurality of application user interfaces; as well as Based on determining that the user's attention was more recently directed to the second plurality of application user interfaces than to other application user interfaces outside the second plurality of application user interfaces, the corresponding main menu positioning is based on the positioning of the second plurality of application user interfaces, wherein the corresponding main menu positioning based on the positioning of the second plurality of application user interfaces is different from the corresponding main menu positioning based on the positioning of the first plurality of application user interfaces.

227. The method of claim 225, wherein displaying the main menu user interface at the corresponding main menu location comprises: determining the respective main menu locations based on the respective positions of the plurality of application user interfaces in accordance with determining that the plurality of application user interfaces ceased to be displayed within the first time threshold from the first time and the first application user interface was the last application closed to cease to be displayed within a third time threshold since closing any other application user interface within the plurality of application user interfaces; as well as Based on determining that the multiple application user interfaces stop displaying within the first time threshold starting from the first time, and the first application user interface is the last closed application that stops displaying when the third time threshold has been exceeded since closing any one of the other application user interfaces within the multiple application user interfaces, the corresponding main menu position is determined based on the first position, wherein the corresponding main menu position based on the first position is different from the corresponding main menu position based on the corresponding positions of the multiple application user interfaces.

228. The method of claim 225, wherein determining the respective main menu locations based on the respective positions of the plurality of application user interfaces comprises: determining the corresponding main menu location based on a set of application user interface locations including the corresponding location of the corresponding application user interface in accordance with determining that the corresponding application user interface at the corresponding location is within a first distance threshold from other application user interfaces of the plurality of application user interfaces; as well as Based on determining that the corresponding application user interface at the corresponding position is not within the first distance threshold from the other application user interfaces among the multiple application user interfaces, the corresponding main menu position is determined based on a set of one or more application user interface positions of the corresponding position that does not include the corresponding application user interface.

229. The method according to any one of claims 222 to 228, comprising: When the main menu user interface is displayed, detecting a user input directed to the main menu user interface; as well as In response to detecting the user input directed to the main menu user interface, initiating execution of system operations including one or more of: displaying an application user interface for a software application, displaying a user interface for initiating a communication session with a communication session participant, and displaying a computer-generated three-dimensional environment in the three-dimensional environment.

230. The method of any one of claims 222 to 229, comprising: detecting a second input corresponding to a request to invoke the main menu user interface; as well as In response to detecting the second input, the main menu user interface is displayed at a location independent of the first location of the first application user interface.

231. The method of claim 230, wherein displaying the main menu user interface at the location independent of the first location of the first application user interface comprises: displaying the main menu user interface at a first main menu position in the three-dimensional environment based on determining that the second input is detected when the user's current attention is directed to a first target position in the three-dimensional environment; as well as Based on determining that the second input is detected when the user's current attention is directed toward a second target position in the three-dimensional environment, wherein the second target position is different from the first target position, the main menu user interface is displayed at a second main menu position in the three-dimensional environment, wherein the second main menu position is different from the first main menu position.

232. A method according to claim 231, wherein displaying the main menu user interface at the first main menu position in the three-dimensional environment includes displaying the main menu user interface at a corresponding distance from the user's viewpoint, and displaying the main menu user interface at the second main menu position in the three-dimensional environment includes displaying the main menu user interface at the corresponding distance from the user's viewpoint.

233. The method of any one of claims 222 to 232, wherein in response to detecting the first input, displaying the main menu user interface at the corresponding main menu location determined based on the first position of the first application user interface in the three-dimensional environment based on determining that one or more other corresponding application user interfaces satisfy clustering criteria relative to the first application user interface in the three-dimensional environment comprises: The main menu user interface is displayed at a characteristic location associated with the first application user interface and the one or more other corresponding application user interfaces.

234. The method of any one of claims 222 to 233, further comprising: detecting a third input; as well as In response to detecting the third input: displaying the main menu user interface in the current viewport of the three-dimensional environment in accordance with determining that the main menu user interface is not within a portion of the three-dimensional environment that is included in a current viewport of the three-dimensional environment when the third input is detected; as well as Based on determining that when the third input is detected, the main menu user interface ceases to be displayed within the portion of the three dimensional environment included in the current viewport of the three dimensional environment.

235. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generating component and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 222 to 234.

236. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 222 to 234.

237. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: A component for performing the method according to any one of claims 222 to 234.

238. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, the one or more programs comprising instructions for: while displaying, via the display generation component, a first application user interface at a first location in a three-dimensional environment, detecting, at a first time, via the one or more input devices, a first input corresponding to a request to close the first application user interface; as well as In response to detecting the first input corresponding to a request to close the first application user interface: Closing the first application user interface includes stopping displaying the first application user interface in the three-dimensional environment; as well as Based on determining that the corresponding criteria are met, a main menu user interface is displayed at a corresponding main menu location determined based on the first position of the first application user interface in the three-dimensional environment.

239. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: while displaying, via the display generation component, a first application user interface at a first location in a three-dimensional environment, detecting, at a first time, via the one or more input devices, a first input corresponding to a request to close the first application user interface; as well as In response to detecting the first input corresponding to a request to close the first application user interface: Closing the first application user interface includes stopping displaying the first application user interface in the three-dimensional environment; as well as Based on determining that the corresponding criteria are met, a main menu user interface is displayed at a corresponding main menu location determined based on the first position of the first application user interface in the three-dimensional environment.

240. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: means enabled when displaying a first application user interface at a first location in a three-dimensional environment via the display generation component, for detecting, at a first time, via the one or more input devices, a first input corresponding to a request to close the first application user interface; and In response to detecting the first input corresponding to a request to close the first application user interface, the component is enabled, the component comprising: means for: closing the first application user interface, comprising ceasing to display the first application user interface in the three-dimensional environment; and Means, enabled based on determining that a corresponding criterion is satisfied, for displaying a main menu user interface at a corresponding main menu location determined based on the first position of the first application user interface in the three-dimensional environment.

241. A method comprising: At a computer system in communication with a display generating component and one or more input devices: displaying, via the display generation component, a first user interface object, wherein the first user interface object includes first content; while displaying, via the display generating component, the first user interface object including the first content, detecting, via the one or more input devices, a first user input directed to the first user interface object; In response to detecting the first user input directed to the first user interface object: based on determining that the first user input corresponds to a request to resize the first user interface object, resizing the first user interface object based on the first user input, Adjusting the size of the first user interface object according to the first user input includes: One or more temporary resize operations, including: based on determining that a characteristic refresh rate of the first content within the first user interface object is a first refresh rate, before updating the first content within the first user interface object based on a first updated size of the first user interface object specified by the first user input, scaling the first user interface object having the first content by a first scaling amount; based on determining that the characteristic refresh rate of the first content within the first user interface object is a second refresh rate different from the first refresh rate, scaling the first user interface object having the first content by a second scaling amount different from the first scaling amount before updating the first content within the first user interface object based on the first updated size of the first user interface object specified by the first user input; After the one or more temporary resize operations, The first user interface object is displayed at the first updated size specified by the first user input, and the first content within the first user interface object is updated according to the first updated size specified by the first user input.

242. The method of claim 241, wherein the one or more temporary resizing operations include: While displaying the first user interface object and the first content at respective scaled sizes, and before displaying the first user interface object at the first updated size and with the updated content corresponding to the first updated size of the first user interface object: Based on determining that the characteristic refresh rate of the first content within the first user interface object is a third refresh rate, before updating the first content within the first user interface object based on the first updated size of the first user interface object specified by the first user input, scaling the first user interface object having the first content by a third scaling amount that is different from the first scaling amount and the second scaling amount.

243. A method according to any one of claims 241 to 242, wherein displaying the first user interface object includes displaying a first application window corresponding to a first application and displaying first application content within the first application window.

244. The method of claim 243, wherein: scaling the first user interface object having the first content includes stretching or shrinking the first content as a whole according to a difference between a currently displayed size and a currently requested size specified by the first user input, and Updating the first content within the first user interface object includes changing the first content in one or more aspects other than scaling of the first content as a whole.

245. The method according to any one of claims 241 to 244, comprising: displaying one or more accessory objects corresponding to the first user interface object concurrently with the first user interface object, wherein the one or more accessory objects maintain their respective spatial relationships with the first user interface object when the first user interface object is displayed at a first location and at a second location different from the first location; as well as In response to detecting the first user input directed to the first user interface object, based on determining that the first user interface object corresponds to a request to resize the first user interface object: When scaling the first user interface object having the first content, abandoning scaling the one or more accessory objects; as well as When the first user interface object is displayed having the first updated size and the updated first content, the one or more accessory objects are displayed in their respective spatial relationship to the first user interface object displayed at the first updated size.

246. The method of any one of claims 241 to 245, comprising: In response to detecting the first user input directed to the first user interface object: Based on determining that the first user input corresponds to a resizing operation of the first user interface object, a visual prominence of the first user interface object is reduced when the first user interface object is resized based on the first user input.

247. The method according to claim 246, comprising: detecting termination of the first user input after reducing the visual prominence of the first user interface object having the first content in accordance with determining that the first user input corresponds to a resize operation of the first user interface object; as well as In response to detecting the termination of the first user input, the visual prominence of the first user interface object is increased.

248. The method of any one of claims 241 to 247, wherein adjusting the size of the first user interface object based on the first user input comprises: Based on determining that the characteristic refresh rate of the first content within the first user interface object is higher than a first threshold refresh rate, abandoning scaling the first user interface object having the first content before displaying the first user interface object as being at the first updated size and having the updated first content corresponding to the first updated size of the first user interface object.

249. The method of any one of claims 241 to 248, wherein adjusting the size of the first user interface object based on the first user input comprises: Based on determining that the characteristic refresh rate of the first content within the first user interface object is lower than a second threshold refresh rate, the first user interface object having the first content is scaled to the first updated size before displaying the first user interface object as being at the first updated size and having the updated first content corresponding to the first updated size of the first user interface object.

250. The method of any one of claims 241 to 249, wherein adjusting the size of the first user interface object based on the first user input comprises: Based on determining that the characteristic refresh rate of the first content within the first user interface object is within a first refresh rate range, before displaying the first user interface object as being at the first updated size and having the updated first content corresponding to the first updated size of the first user interface object, scaling the first user interface object having the first content by a corresponding scaling amount, wherein the corresponding scaling amount is selected from a scaling amount range based on a current value of the characteristic refresh rate of the first content within the first refresh rate range.

251. The method of any one of claims 241 to 250, comprising: calculating two or more candidate refresh rates for the first content using two or more different metrics; as well as A lowest refresh rate is selected from the two or more candidate refresh rates as the characteristic refresh rate of the first content.

252. The method of any one of claims 241 to 251, comprising: A first candidate refresh rate for the first content is determined based on recorded refresh rates obtained during one or more previous resize operations performed on the first user interface object or another object similar to the first user interface object.

253. The method according to any one of claims 241 to 252, comprising: A second candidate refresh rate for the first content is determined based on respective rates of one or more updates to the first content that have been performed on the first content during a current resizing operation of the first user interface object.

254. The method of any one of claims 241 to 253, comprising: A third candidate refresh rate for the first content is determined based on a last update to the first content that has been performed on the first content during a current resizing operation of the first user interface object.

255. The method of any one of claims 241 to 254, comprising: In response to detecting the first user input, a resizable affordance is displayed at a location at or near the movable portion of the first user interface object.

256. The method of any one of claims 241 to 255, comprising: In response to detecting the first user input directed to the first user interface object: providing an indication that a corresponding resizing operation was initiated by the beginning portion of the first user input based on determining that the beginning portion of the first user input is directed to a first location outside of the first user interface object and within a first area outside of the first user interface object; based on determining that the beginning portion of the first user input is directed to a second location outside the first user interface object and within the first area that is not outside the first user interface object, forgoing providing the indication that the corresponding resizing operation was initiated by the beginning portion of the first user input; providing the indication that the corresponding resizing operation was initiated by the starting portion of the first user input based on determining that the starting portion of the first user input is directed to a third location within the first user interface object and within a second area in the first user interface object; Based on determining that the starting portion of the first user input points to a fourth position within the first user interface object and not within the second area in the first user interface object, abandoning the indication that the corresponding resizing operation was initiated by the starting portion of the first user input.

257. The method of any one of claims 241 to 254, comprising: prior to resizing the first user interface object based on the first user interface object, displaying a resizable affordance having a corresponding spatial relationship to a portion of the first user interface object; as well as During resizing of the first user interface object in accordance with the first user input, moving the resizing affordance in accordance with the first user input, wherein: for a first value of the characteristic refresh rate of the first content, the resizing affordance having a first distance from the corresponding spatial relationship to the portion of the first user interface object, For a second value of the characteristic refresh rate of the first content, the resizing indicator can be represented by a second distance from the corresponding spatial relationship to the portion of the first user interface object, wherein the first distance is greater than the second distance when the first value of the characteristic refresh rate is higher than the second value of the characteristic refresh rate and / or when the first zoom amount is less than the second zoom amount.

258. A method according to claim 257, wherein the resizable indicator represents a shape and / or size based on spatial characteristics of the movable portion of the first user interface object.

259. A method according to claim 258, wherein for a first value of the spatial characteristic of the movable portion of the first user interface object, the resizing indicator is represented by a first length and / or size, and for a second value of the spatial characteristic of the movable portion of the first user interface object, the resizing indicator is represented by a second length and / or size, wherein when the first value of the spatial characteristic is different from the second value of the spatial characteristic, the first length and / or size is different from the second length and / or size.

260. A method according to any one of claims 257 to 259, wherein the resizing indication is displayed with a first curvature when the movable portion of the first user interface object has the first value of the spatial characteristic, and wherein the resizing indication is displayed with a second curvature when the movable portion of the first user interface object has the second value of the spatial characteristic, wherein the first curvature and the second curvature are different when the first value of the spatial characteristic is different from the second value of the spatial characteristic.

261. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generating component and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 241 to 260.

262. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 241 to 260.

263. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: A component for performing the method according to any one of claims 241 to 260.

264. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, the one or more programs comprising instructions for: displaying, via the display generation component, a first user interface object, wherein the first user interface object includes first content; while displaying, via the display generating component, the first user interface object including the first content, detecting, via the one or more input devices, a first user input directed to the first user interface object; In response to detecting the first user input directed to the first user interface object: based on determining that the first user input corresponds to a request to resize the first user interface object, resizing the first user interface object based on the first user input, Adjusting the size of the first user interface object according to the first user input includes: One or more temporary resize operations, including: based on determining that a characteristic refresh rate of the first content within the first user interface object is a first refresh rate, scaling the first user interface object having the first content by a first scaling amount before updating the first content within the first user interface object according to a first updated size of the first user interface object specified by the first user input; based on determining that the characteristic refresh rate of the first content within the first user interface object is a second refresh rate different from the first refresh rate, scaling the first user interface object having the first content by a second scaling amount different from the first scaling amount before updating the first content within the first user interface object based on the first updated size of the first user interface object specified by the first user input; After the one or more temporary resizing operations, the first user interface object is displayed at the first updated size specified by the first user input, and the first content within the first user interface object is updated according to the first updated size specified by the first user input.

265. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, a first user interface object, wherein the first user interface object includes first content; while displaying, via the display generating component, the first user interface object including the first content, detecting, via the one or more input devices, a first user input directed to the first user interface object; In response to detecting the first user input directed to the first user interface object: based on determining that the first user input corresponds to a request to resize the first user interface object, resizing the first user interface object based on the first user input, Adjusting the size of the first user interface object according to the first user input includes: One or more temporary resize operations, including: based on determining that a characteristic refresh rate of the first content within the first user interface object is a first refresh rate, scaling the first user interface object having the first content by a first scaling amount before updating the first content within the first user interface object according to a first updated size of the first user interface object specified by the first user input; based on determining that the characteristic refresh rate of the first content within the first user interface object is a second refresh rate different from the first refresh rate, scaling the first user interface object having the first content by a second scaling amount different from the first scaling amount before updating the first content within the first user interface object based on the first updated size of the first user interface object specified by the first user input; After the one or more temporary resizing operations, the first user interface object is displayed at the first updated size specified by the first user input, and the first content within the first user interface object is updated according to the first updated size specified by the first user input.

266. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: means for displaying, via the display generation component, a first user interface object, wherein the first user interface object comprises first content; means for detecting, via the one or more input devices, a first user input directed to the first user interface object, enabled when the first user interface object including the first content is displayed via the display generating component; In response to detecting the first user input directed to the first user interface object, means for: based on determining that the first user input corresponds to a request to resize the first user interface object, resizing the first user interface object based on the first user input, Adjusting the size of the first user interface object according to the first user input includes: One or more temporary resize operations, including: based on determining that a characteristic refresh rate of the first content within the first user interface object is a first refresh rate, scaling the first user interface object having the first content by a first scaling amount before updating the first content within the first user interface object according to a first updated size of the first user interface object specified by the first user input; based on determining that the characteristic refresh rate of the first content within the first user interface object is a second refresh rate different from the first refresh rate, scaling the first user interface object having the first content by a second scaling amount different from the first scaling amount before updating the first content within the first user interface object based on the first updated size of the first user interface object specified by the first user input; After the one or more temporary resizing operations, the first user interface object is displayed at the first updated size specified by the first user input, and the first content within the first user interface object is updated according to the first updated size specified by the first user input.

267. A method comprising: At a first computer system in communication with one or more display generating components and one or more input devices: displaying, via the one or more display generation components, a first application window at a first scale; detecting, via the one or more input devices, a first gesture directed toward the first application window while the first application window is displayed at the first scale; as well as In response to detecting the first gesture, based on determining that the first gesture points to a corresponding portion of the first application window that is not associated with an application-specific response to the first gesture, changing a corresponding scale of the first application window from the first scale to a second scale that is different from the first scale.

268. The method of claim 267, comprising, in response to detecting the first gesture: Based on determining that the first gesture points to a corresponding portion of the first application window associated with an application-specific response to the first gesture, a first application-specific response to the first gesture is performed by changing the appearance of corresponding content of the corresponding portion of the first application window associated with the application-specific response to the first gesture relative to the appearance of the first application window.

269. The method of claim 268, wherein performing the first application-specific response to the first gesture comprises resizing the corresponding content relative to other portions of the first application window in accordance with the first gesture.

270. The method of any one of claims 267 to 269, wherein the first gesture is an air gesture.

271. The method of claim 270, wherein the first gesture is a two-handed air gesture.

272. The method of any one of claims 267 to 271, wherein: The first application window includes a first content and a second content separate from the first content; and Changing the scale of the first application window from the first scale to the second scale includes simultaneously changing a corresponding size of the first content and a corresponding size of the second content.

273. The method of claim 267, wherein: The first application window includes first content and one or more window management controls separate from the first content; and Changing the corresponding scale of the first application window from the first scale to the second scale includes simultaneously changing a corresponding size of the first content and a corresponding size of the one or more window management controls.

274. The method of any one of claims 267 to 273, wherein: Changing the corresponding proportion of the first application window includes changing the corresponding size of the first application window without changing the content displayed in the first application window; and The method comprises: detecting a first window resize input directed to the first application window; as well as In response to detecting the first window resize input directed to the first application window, changing the corresponding size of the first application window, including maintaining the corresponding proportions of one or more elements of the first application window at the first proportion when changing at least some content of the first application window based on the corresponding size of the first application window that has been changed.

275. A method according to any one of claims 267 to 274, wherein the corresponding portion of the first application window that is not associated with the application-specific response to the first gesture includes one or more window management controls of the first application window (e.g., one or more accessory objects of the first application window).

276. The method of any one of claims 267 to 275, wherein: The first application window has a corresponding ratio limit; Changing the corresponding scale of the first application window from the first scale to a second scale different from the first scale comprises: When the first gesture is detected, changing the corresponding scale of the first application window to a third scale exceeding the corresponding scale limit; and After changing the corresponding scale of the first application window to the third scale exceeding the corresponding scale limit: detecting that the end of the corresponding input has been detected; and In response to detecting the end of the first gesture, changing the respective scale of the first application window to a fourth scale that is at or within the respective scale limit.

277. The method of claim 276, wherein: The first application window has an upper limit and a lower limit of the ratio; and Changing the corresponding scale of the first application window to the fourth scale at or within the corresponding scale limit comprises: In response to determining that the corresponding scale limit is the scale upper limit, reducing the corresponding scale of the first application window to a corresponding upper scale that is equal to or lower than the scale upper limit; and Based on determining that the corresponding scale limit is the scale lower limit, the corresponding scale of the first application window is reduced to a corresponding lower scale that is equal to or higher than the scale lower limit.

278. The method of any one of claims 267 to 277, comprising: detecting an end of the first gesture after changing the corresponding scale of the first application window from the first scale to a second scale different from the first scale; as well as In response to detecting the end of the first gesture, maintaining the corresponding scale of the first application window at the second scale.

279. The method of claim 278, comprising: When the first application window is displayed at the second scale, detecting a request to close the first application window; In response to detecting the request to close the first application window, closing the first application window; After closing the first application window, detecting a user input corresponding to a request to open the first application window; as well as In response to detecting the user input corresponding to the request to open the first application window, the first application window is displayed at the first scale.

280. The method of claim 278, comprising: When the first application window is displayed at the second scale, detecting a request to close the first application window; In response to detecting the request to close the first application window, closing the first application window; After closing the first application window, detecting a user input corresponding to a request to open the first application window; as well as In response to detecting the user input corresponding to the request to open the first application window, displaying the first application window at the second scale includes: displaying the first application window at the first adjusted scale in response to the user input corresponding to the request to open the first application window based on determining that the first application window was displayed at the first adjusted scale when the first application window was last closed; as well as The first application window is displayed at the second adjusted scale in response to the user input corresponding to the request to open the first application window based on determining that the first application window was displayed at the second adjusted scale when the first application window was last closed.

281. The method of any one of claims 267 to 280, wherein: The first application window is displayed in a three-dimensional environment visible via a viewport provided by the one or more display generation components; and The method comprises: When the first application window is displayed at the corresponding scale, detecting a change in the positioning of the first application window relative to the viewport, the change causing the first application window to cease being displayed in the viewport; After causing the first application window to cease display in the viewport, detecting a user input corresponding to a request to display the first application window in the viewport; and In response to detecting the user input corresponding to the request to display the first application window in the viewport, displaying the first application window in the viewport at the corresponding scale includes: displaying the first application window at the third adjusted scale in response to the user input corresponding to the request to display the first application window in the viewport based on determining to display the first application window at the third adjusted scale when the first application window ceased to be displayed in the viewport; and Based on determining to display the window at a fourth adjusted scale when the window ceased to be displayed in the viewport, the first application window is displayed at the fourth adjusted scale in response to the user input corresponding to the request to display the application window in the viewport.

282. The method of any one of claims 267 to 281, comprising, when displaying the first application window at the second scale: detecting one or more inputs directed to the first application window other than the first gesture; and In response to detecting one or more inputs directed to the first application window that are different from the first gesture, while maintaining the first application window at the second scale, perform one or more operations associated with the first application providing the first application window based on the one or more inputs directed to the first application window.

283. The method of any one of claims 267 to 282, comprising: while displaying the first application window at the second scale via the one or more display generating components, detecting, via the one or more input devices, a second gesture directed toward the first application window; In response to detecting the second gesture, based on determining that the second gesture points to a corresponding portion of the first application window that is not associated with an application-specific response to the second gesture, changing the corresponding scale of the first application window from the second scale to a fifth scale that is different from the second scale.

284. The method of claim 283, comprising, in response to detecting the second gesture: Based on determining that the second gesture points to a corresponding portion of the first application window associated with the application-specific response to the second gesture, a second application-specific response to the second gesture is performed by changing the appearance of corresponding content of the corresponding portion of the first application window associated with the application-specific response to the second gesture relative to the appearance of the first application window.

285. The method of any one of claims 267 to 284, wherein: evaluating, by one or more gesture recognizers associated with the first application window, one or more input events corresponding to respective gestures directed toward the first application window; and authorizing a system gesture recognizer to use the one or more input events for one or more system operations based on determining that the one or more gesture recognizers associated with the first application window determine that the one or more input events corresponding to the corresponding gesture directed to the application window do not correspond to input to be consumed by the one or more gesture recognizers associated with the first application window; as well as Based on determining that the one or more gesture recognizers associated with the first application window determine that the one or more input events corresponding to the corresponding gesture directed to the application window correspond to input to be consumed by the one or more gesture recognizers associated with the first application window, relinquishing authorization for the one or more input events to be used by the system gesture recognizers for one or more system operations.

286. The method of claim 285, wherein: The one or more gesture recognizers associated with the first application window include a set of gesture recognizers having different priorities for processing the one or more input events, and A first gesture recognizer in the set of gesture recognizers that is responsible for performing one or more application-specific responses to the corresponding gesture by changing the appearance of corresponding content relative to the appearance of the first application window has a higher priority for processing the one or more input events than a second gesture recognizer in the set of gesture recognizers that is responsible for authorizing the system gesture recognizer to use the one or more input events for one or more system operations.

287. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a first computer system that communicates with one or more display generating components and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 267 to 286.

288. A first computer system in communication with one or more display generation components and one or more input devices, the first computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 267 to 286.

289. A first computer system, the first computer system in communication with one or more display generation components and one or more input devices, the first computer system comprising: A component for performing a method according to any one of claims 267 to 286.

290. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a first computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for: displaying, via the one or more display generation components, a first application window at a first scale; While the first application window is displayed at the first scale, detecting, via the one or more input devices, a first gesture directed toward the first application window; and In response to detecting the first gesture, based on determining that the first gesture points to a corresponding portion of the first application window that is not associated with an application-specific response to the first gesture, changing a corresponding scale of the first application window from the first scale to a second scale that is different from the first scale.

291. A first computer system, the first computer system in communication with one or more display generation components and one or more input devices, the first computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the one or more display generation components, a first application window at a first scale; detecting, via the one or more input devices, a first gesture directed toward the first application window while the first application window is displayed at the first scale; as well as In response to detecting the first gesture, based on determining that the first gesture points to a corresponding portion of the first application window that is not associated with an application-specific response to the first gesture, changing a corresponding scale of the first application window from the first scale to a second scale that is different from the first scale.

292. A first computer system in communication with one or more display generation components and one or more input devices, the first computer system comprising: means for displaying, via the one or more display generating components, a first application window at a first scale; means enabled when the first application window is displayed at the first scale for detecting, via the one or more input devices, a first gesture directed toward the first application window; and A component enabled in response to detecting the first gesture and used for the following operation based on determining that the first gesture is directed to a corresponding portion of the first application window that is not associated with an application-specific response to the first gesture: changing the corresponding scale of the first application window from the first scale to a second scale different from the first scale.