Method for controlling and interacting with three-dimensional environment

By optimizing the user interface of virtual reality and augmented reality environments, using immersion, volume and focus mode control elements, the problem of inefficiency in the existing technology is solved, and more efficient user interaction and resource conservation is achieved.

CN120266077APending Publication Date: 2025-07-04APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380081359.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-02
Filing Date
2023-09-22
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing methods of interacting with virtual reality and augmented reality environments are inefficient, user input is cumbersome and error-prone, resulting in wasted computer system resources and reduced user experience.

Method used

By providing immersion control, volume control and focus mode control elements, the user interface is optimized, the number and nature of user input is reduced, the interaction efficiency is improved, and user input is detected and processed through computer systems to optimize the display and interaction of virtual content.

Benefits of technology

Improve user interaction efficiency, reduce computer system resource consumption, extend battery life, and improve user experience and device operability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120266077A_ABST
    Figure CN120266077A_ABST
Patent Text Reader

Abstract

A computer system displays an immersion control, a volume control, an element configured to allow or limit penetration in different operating modes of the computer system, and / or an option selectable to cause display of a representation of content from a second computer system via a display generation component of the computer system. While the second computer system is displaying the content, the first computer system detects an input corresponding to a request to display a representation of the content from the second computer system via a display generating component of the first computer system, and in response, displays the representation of the content from the second computer system. A process is initiated that displays a representation of the content and does not emphasize the content displayed by the second computer system. The first computer system facilitates disambiguating a second computer system from the plurality of computer systems to display a representation of content from the second computer system.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 377,028, filed on September 24, 2022, U.S. Provisional Application No. 63 / 505,690, filed on June 1, 2023, and U.S. Provisional Application No. 63 / 506,042, filed on June 2, 2023, the entire contents of which are incorporated herein by reference for all purposes. Technical Field

[0003] The present invention generally relates to computer systems that provide computer-generated experiences, including but not limited to electronic devices that provide virtual reality and mixed reality experiences via a display. Background Art

[0004] In recent years, the development of computer systems for augmented reality has grown significantly. Example augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays) for computer systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Example virtual elements include virtual objects such as digital images, videos, text, icons, and control elements (such as buttons and other graphics). Summary of the Invention

[0005] Some methods and interfaces for interacting with an environment that includes at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems where virtual object manipulation is complex, cumbersome, and error-prone can impose a significant cognitive burden on the user and detract from the experience of the virtual / augmented reality environment. In addition, these methods take longer than necessary, thereby wasting the energy of the computer system. This latter consideration is particularly important in battery-powered devices.

[0006] Accordingly, there is a need for computer systems with improved methods and interfaces for providing computer-generated experiences to users, such that the interaction between the user and the computer system is more effective and intuitive for the user. Such methods and interfaces optionally supplement or replace conventional methods for providing extended reality experiences. Such methods and interfaces form a more effective human-machine interface by helping the user understand the connection between the inputs provided and the device's response to those inputs, thereby reducing the quantity, degree, and / or nature of the inputs from the user.

[0007] The above-mentioned deficiencies and other problems associated with the user interface of a computer system are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a laptop computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touch screen" or "touch screen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, in addition to a display generation component, the computer system also has one or more output devices, which include one or more haptic output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, a memory, and one or more modules, a program or set of instructions stored in the memory for performing multiple functions. In some embodiments, the user interacts with the GUI through contact and gestures of a stylus and / or finger on a touch-sensitive surface, movement of the user's eyes and hands in space relative to the GUI (and / or the computer system) or in the space of the user's body (as captured by cameras and other motion sensors), and / or voice input (as captured by one or more audio input devices). In some embodiments, the functions performed through the interaction optionally include image editing, drawing, presentation, word processing, spreadsheet creation, playing games, making and receiving phone calls, video conferencing, sending and receiving emails, instant messaging, fitness support, digital photography, digital video recording, web browsing, digital music playing, note-taking, and / or digital video playing. The executable instructions for performing these functions are optionally included in a transient and / or non-transient computer-readable storage medium or other computer program product configured to be executed by one or more processors.

[0008] There is a need for electronic devices having improved methods and interfaces for interacting with content in a three-dimensional environment. Such methods and interfaces can supplement or replace conventional methods for interacting with content in a three-dimensional environment. Such methods and interfaces reduce the amount, degree, and / or nature of input from the user and result in a more efficient human-machine interface. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charges.

[0009] In some embodiments, a computer system displays immersion control elements for controlling the computer system to display virtual content at an immersion level. In some embodiments, the computer system displays volume control elements for controlling the volume level of a virtual environment and / or for controlling the volume level of a user interface of an application. In some embodiments, the computer system displays focus mode control elements that can be selected to allow or restrict reducing the salience of at least a portion of the virtual content relative to at least a portion of the physical environment. In some embodiments, the computer system displays options that can be selected to cause a representation of content from a second computer system to be displayed via a display generation component of the computer system. In some embodiments, when the second computer system is displaying content, the first computer system detects an input corresponding to a request to display a representation of content from the second computer system via a display generation component of the first computer system, and in response, initiates a process of displaying a representation of content from the second computer system and de-emphasizing the content displayed by the second computer system. In some embodiments, the first computer system facilitates disambiguating the second computer system from among multiple computer systems to display a representation of content from the second computer system.

[0010] Note that the various embodiments described above can be combined with any other embodiments described herein. The features and advantages described in this specification are not exhaustive, and in particular, many additional features and advantages will be apparent to those of ordinary skill in the art from the drawings, the specification, and the claims. Additionally, it should be noted that the language used in this specification has been selected for readability and guidance purposes in principle, and may not have been selected to depict or define the subject matter of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To better understand the various described embodiments, reference should be made to the following detailed description in conjunction with the following drawings, in which like reference numerals indicate corresponding parts in all the drawings.

[0012] Figure 1A is a block diagram illustrating an operating environment of a computer system for providing an XR experience according to some embodiments.

[0013] Figures 1B to 1P is for use in Figure 1A an example of a computer system for providing an XR experience in an operating environment.

[0014] Figure 2 is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's XR experience according to some embodiments.

[0015] Figure 3A block diagram of a display generation component configured to provide a visual component of an XR experience to a user in a computer system according to some embodiments.

[0016] Figure 4 A block diagram of a hand tracking unit configured to capture a user's gesture input in a computer system according to some embodiments.

[0017] Figure 5 A block diagram of an eye tracking unit configured to capture a user's gaze input in a computer system according to some embodiments.

[0018] Figure 6 A flowchart illustrating a flash-assisted gaze tracking pipeline according to some embodiments.

[0019] Figures 7A to 7H An example of a computer system that promotes control of the immersion level in a virtual environment according to some embodiments is illustrated.

[0020] Figures 8A to 8I A flowchart illustrating an exemplary method for promoting control of the immersion level in a virtual environment according to some embodiments.

[0021] Figures 9A to 9E An example of controlling the audio settings of a virtual environment according to some embodiments is illustrated.

[0022] Figures 10A to 10G A flowchart illustrating a method for controlling the audio settings of a virtual environment according to some embodiments.

[0023] Figures 11A to 11F An example of controlling the penetration settings of a computer system for displaying a three-dimensional environment via a display generation component according to some embodiments is illustrated.

[0024] Figures 12A to 12J A flowchart illustrating a method for controlling the penetration settings of a computer system for displaying a three-dimensional environment via a display generation component according to some embodiments.

[0025] Figures 13A to 13D An example of a first computer system that promotes the display of a representation of content from a second computer system in a three-dimensional environment according to some embodiments is illustrated.

[0026] Figures 14A to 14H A flowchart illustrating a method for promoting the display of a representation of content from a second computer system in a three-dimensional environment according to some embodiments.

[0027] Figures 15A to 15E An example of promoting the initiation of a virtual computer experience in a three-dimensional environment according to some embodiments is illustrated.

[0028] Figures 16A to 16J is a flowchart illustrating a method for facilitating the initiation of a virtual computer experience in a three - dimensional environment according to some embodiments.

[0029] Figures 17A to 17H An example of a first computer system that illustrates disambiguating a second computer system from multiple computer systems to display a representation of content from the second computer system in a three - dimensional environment according to some embodiments.

[0030] Figure 18 is a flowchart illustrating a method for facilitating the disambiguation of a second computer system from multiple computer systems to display a representation of content from the second computer system in a three - dimensional environment according to some embodiments.

[0031] Figure 19 is a flowchart illustrating a method for facilitating the disambiguation of a second computer system from multiple computer systems to display a representation of content from the second computer system in a three - dimensional environment according to some embodiments. Detailed Description

[0032] According to some embodiments, the present disclosure relates to a user interface for providing a computer - generated (CGR) experience to a user.

[0033] The systems, methods, and GUIs described herein provide improved ways for an electronic device to facilitate interaction with and manipulation of objects in a three - dimensional environment.

[0034] In some embodiments, a computer system displays an immersion control element for controlling the level of immersion at which the computer system displays virtual content. The level of immersion at which the computer system displays virtual content can be increased and / or decreased based on an input pointing to the immersion control element.

[0035] In some embodiments, a computer system displays a volume control element for controlling the volume level of a virtual environment and / or for controlling the volume level of a user interface of an application. The volume level of the virtual environment and / or the volume level of the user interface of the application can be increased and / or decreased based on an input pointing to the volume control element.

[0036] In some embodiments, a computer system displays a focus mode control element that can be selected to allow or restrict reducing the salience of at least a portion of virtual content relative to at least a portion of a physical environment. In response to selection of the focus mode control element, the operation mode of the computer system can be changed. For example, in a first operation mode of the computer system, optionally, reducing the salience of at least a portion of virtual content relative to at least a portion of the physical environment is performed in response to a first event satisfying one or more first criteria. In a second operation mode of the computer system, optionally, reducing the salience of at least a portion of virtual content relative to at least a portion of the physical environment is not performed in response to a second event satisfying one or more second criteria.

[0037] In some embodiments, the computer system displays an option that can be selected to cause a representation of content from a second computer system to be displayed via a display generation component of the computer system. In response to selection of the option, optionally, a representation of content from the second computer system is caused to be displayed via the display generation component of the computer system, and operations can be performed on the representation of content from the second computer system.

[0038] In some embodiments, when the second computer system is displaying content, the first computer system detects an input corresponding to a request to display a representation of content from the second computer system via a display generation component of the first computer system, and in response, initiates a process of displaying a representation of content from the second computer system and de-emphasizing the content displayed by the second computer system. Optionally, de-emphasizing the content displayed by the second computer system includes: displaying different content or ceasing to display any content from the display by the second computer system.

[0039] In some embodiments, the first computer system visually detects a second computer system in a physical environment corresponding to a three-dimensional environment visible via a display generation component via one or more cameras. In some embodiments, in response to visually detecting the second computer system, based on determining that the second computer system satisfies one or more connection criteria, the first computer system displays a first selectable option in the three-dimensional environment that can be selected to initiate a process of establishing a connection between the first computer system and the second computer system. In some embodiments, based on determining that the second computer system does not satisfy one or more connection criteria, the first computer system refrains from displaying the first selectable option in the three-dimensional environment.

[0040] In some embodiments, a first computer system detects, via one or more input devices, a request to establish a connection with a corresponding computer system that is different from the first computer system and that is within a corresponding region of the physical environment of the first computer system. In some embodiments, in response to detecting the request, the first computer system establishes a connection between the first computer system and a second computer system among the plurality of computer systems when the plurality of computer systems are within the corresponding region, where the second computer system meets one or more criteria, and does not establish a connection between the first computer system and other computer systems among the plurality of computer systems. In some embodiments, in response to determining that a third computer system that is different from the second computer system among the plurality of computer systems meets one or more criteria when the plurality of computer systems are within the corresponding region, the first computer system establishes a connection between the first computer system and the third computer system and does not establish a connection between the first computer system and other computer systems (including the second computer system) among the plurality of computer systems.

[0041] Figures 1A to 6 A description of an example computer system for providing an XR experience to a user is provided (such as described below with respect to methods 800, 1000, 1200, 1400, 1600, 1800, and / or 1900). Figures 7A to 7H An example of a computer system that promotes control of immersion in a virtual environment according to some embodiments is illustrated. Figures 8A to 8I FIG. is a flowchart illustrating an exemplary method that promotes control of immersion in a virtual environment according to some embodiments. Figures 7A to 7H The user interface in is used to illustrate Figures 8A to 8I the process in. Figures 9A to 9E An example of a computer system that controls audio settings of a virtual environment according to some embodiments is illustrated. Figures 10A to 10G FIG. is a flowchart of a method that controls audio settings of a virtual environment according to some embodiments. Figures 9A to 9E The user interface in is used to illustrate Figures 10A to 10G the process in. Figures 11A to 11F An example technique for controlling a penetration setting of a computer system that displays a three-dimensional environment via a display generation component according to some embodiments is illustrated. Figures 12A to 12J FIG. is a flowchart of a method for controlling a penetration setting of a computer system that displays a three-dimensional environment via a display generation component according to various embodiments. Figures 11A to 11F The user interface in is used to illustrate Figures 12A to 12J the process in. Figures 13A to 13D An example technique for promoting display of a representation of content from a second computer system in a three-dimensional environment according to some embodiments is illustrated. Figures 14A to 14HIt is a flowchart of a method for facilitating the display of a representation of content from a second computer system in a three-dimensional environment according to various embodiments. Figures 13A to 13D The user interface in Figures 14A to 14H illustrates the Figures 15A to 15E process in. Figures 16A to 16J It is a flowchart of a method for facilitating the initiation of a virtual computer experience in a three-dimensional environment according to various embodiments. Figures 15A to 15E The user interface in Figures 16A to 16J illustrates the Figures 17A to 17H process in. Figure 18 It is a flowchart of a method for facilitating the disambiguation of a second computer system from multiple computer systems in order to display a representation of content from the second computer system in a three-dimensional environment according to some embodiments. Figures 17A to 17H The user interface in Figure 18 illustrates the Figure 19 process in. Figures 17A to 17H It is a flowchart of a method for facilitating the disambiguation of a second computer system from multiple computer systems in order to display a representation of content from the second computer system in a three-dimensional environment according to some embodiments. Figure 19 The user interface in

[0042] The processes described below enhance the operability of a device and make the user-device interface more effective through various techniques (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), including providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional display controls, performing an operation when a set of conditions has been met without further user input, improving privacy and / or security, providing a more diverse, detailed, and / or realistic user experience while saving storage space, and / or additional techniques. These techniques also reduce power usage and extend the battery life of the device by enabling the user to use the device faster and more effectively. Saving battery power, and thus weight, improves the ergonomics of the device. These techniques also enable real-time communication, allow the use of fewer and / or less precise sensors, resulting in a more compact, lighter, and cheaper device, and enable the device to be used under various lighting conditions. These techniques reduce energy usage, thereby reducing the heat emitted by the device, which is particularly important for wearable devices, where it can become uncomfortable for the user to wear the device if the device generates too much heat while operating entirely within the operating parameters of the device components.

[0043] In addition, in methods described herein where one or more steps depend on one or more conditions being met, it should be understood that the method can be repeated in multiple iterations such that, during the course of the repetition, all conditions that determine the steps in the method are met in different iterations of the method. For example, if a method requires performing a first step (if a condition is met) and performing a second step (if the condition is not met), a person of ordinary skill in the art will know to repeat the stated steps until both the condition is met and the condition is not met (in no particular order). Thus, a method described as having one or more steps that depend on one or more conditions being met can be rewritten as a method that repeats until each condition described in the method has been met. However, this does not require a system or computer-readable medium to state that the system or computer-readable medium includes instructions for performing conditional operations based on the satisfaction of corresponding one or more conditions and is thus capable of determining whether a possible scenario has been met without explicitly repeating the steps of the method until all conditions that determine the steps in the method have been met. A person of ordinary skill in the art will also understand that, similar to methods with conditional steps, a system or computer-readable storage medium can repeat the steps of the method as many times as needed to ensure that all conditional steps have been performed.

[0044] In some embodiments, as Figure 1AAs shown, an XR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a haptic sensor, an orientation sensor, a proximity sensor, a temperature sensor, a position sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., household appliances, wearable devices, etc.). In some embodiments, one or more of the input devices 125, output devices 155, sensors 190, and peripheral devices 195 are integrated with the display generation component 120 (e.g., in a head-mounted device or a handheld device).

[0045] When describing an XR experience, various terms are used to distinctively refer to several related but different environments that a user can sense and / or interact with (e.g., interact using inputs detected by the computer system 101 that generates the XR experience, where the input causes the computer system generating the XR experience to generate audio, visual, and / or haptic feedback corresponding to the various inputs provided to the computer system 101). The following is a subset of these terms:

[0046] Physical environment: The physical environment refers to the physical world that people can sense and / or interact with without the help of an electronic system. Physical environments such as a physical park include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment, such as through vision, touch, hearing, taste, and smell.

[0047] Extended Reality: In contrast, an extended reality (XR) environment is a fully or partially simulated environment in which a person senses and / or interacts with an electronic system. In XR, a subset of a person's physical movements or representations thereof are tracked, and in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner consistent with at least one physical law. For example, an XR system can detect a person's head rotation and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. In some cases (e.g., for accessibility reasons), the adjustment of the characteristics of virtual objects in the XR environment can be made in response to a representation of a physical movement (e.g., a voice command). A person can use any of their senses to sense and / or interact with XR objects, including vision, hearing, touch, taste, and smell. For example, a person can sense and / or interact with an audio object that creates a 3D or spatial audio environment that provides the perception of point audio sources in 3D space. In another example, an audio object can implement audio transparency that selectively introduces ambient sounds from the physical environment with or without computer-generated audio. In some XR environments, a person can sense and / or interact only with audio objects.

[0048] Examples of XR include virtual reality and mixed reality.

[0049] Virtual Reality: A virtual reality (VR) environment is a simulated environment that is designed to be based entirely on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects with which a person can sense and / or interact. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment by way of a simulation of the person's presence within the computer-generated environment and / or by way of a simulation of a subset of the person's physical movements within the computer-generated environment.

[0050] Mixed Reality: Compared with a VR environment that is designed to be completely based on computer-generated sensory input, a mixed reality (MR) environment is an analog environment that is designed to include, in addition to computer-generated sensory input (e.g., virtual objects), sensory input from the physical environment or a representation thereof. On the virtual continuum, an MR environment is any condition between a completely physical environment at one end and a virtual reality environment at the other end, but excluding these two ends. In some MR environments, the computer-generated sensory input can respond to changes in the sensory input from the physical environment. Additionally, some electronic systems for presenting an MR environment can track the position and / or orientation relative to the physical environment so that virtual objects can interact with real objects (i.e., physical items from the physical environment or their representations). For example, the system can take motion into account so that virtual trees appear stationary relative to the physical ground.

[0051] Examples of mixed reality include augmented reality and augmented virtuality.

[0052] Augmented Reality: An augmented reality (AR) environment is a simulated environment in which one or more virtual objects are superimposed on a physical environment or a representation of a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system can be configured to present virtual objects on the transparent or translucent display such that a person using the system perceives the virtual objects superimposed on the physical environment. Alternatively, the system can have an opaque display and one or more imaging sensors that capture an image or video of the physical environment, which is a representation of the physical environment. The system combines the image or video with the virtual objects and presents the combination on the opaque display. A person uses the system to indirectly view the physical environment via the image or video of the physical environment and perceives the virtual objects superimposed on the physical environment. As used herein, a video of a physical environment displayed on an opaque display is referred to as a "passthrough video," meaning that the system uses one or more image sensors to capture an image of the physical environment and uses those images when presenting the AR environment on the opaque display. Further alternatively, the system can have a projection system that projects virtual objects into the physical environment, such as as a hologram or on a physical surface, such that a person using the system perceives the virtual objects superimposed on the physical environment. An augmented reality environment is also a simulated environment in which a representation of a physical environment is transformed by computer-generated sensory information. For example, in providing a passthrough video, the system can transform one or more sensor images to impose a selected perspective (e.g., a viewpoint) different from the perspective captured by the imaging sensors. As another example, a representation of a physical environment can be transformed by graphically modifying (e.g., magnifying) portions thereof such that the modified portions can be a representative but not photo-realistic version of the originally captured image. As yet another example, a representation of a physical environment can be transformed by graphically removing portions thereof or blurring portions thereof.

[0053] Augmented Virtuality: An augmented virtuality (AV) environment is a simulated environment in which a virtual environment or a computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory inputs can be representations of one or more characteristics of the physical environment. For example, an AV park can have virtual trees and virtual buildings, but a person's face is realistically reproduced from a photograph of the physical person. As another example, a virtual object can adopt the shape or color of a physical item imaged by one or more imaging sensors. As yet another example, a virtual object can adopt a shadow that conforms to the positioning of the sun in the physical environment.

[0054] In an augmented reality, mixed reality, or virtual reality environment, a view of a three-dimensional environment is visible to a user. The view of the three-dimensional environment is typically visible to the user through a virtual viewport via one or more display generation components (e.g., a display or a pair of display modules that provide stereoscopic content to different eyes of the same user), the virtual viewport having a viewport boundary that defines the extent of the three-dimensional environment visible to the user via the one or more display generation components. In some embodiments, the region defined by the viewport boundary is smaller than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size, optical properties, or other physical characteristics of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user's eyes). In some embodiments, the region defined by the viewport boundary is larger than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size, optical properties, or other physical characteristics of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user's eyes). The viewport and the viewport boundary typically move as the one or more display generation components move (e.g., for a head-mounted device as the user's head moves, or for a handheld device such as a tablet or smartphone as the user's hand moves). The user's viewpoint determines the content visible in the viewport, the viewpoint typically specifying a position and orientation relative to the three-dimensional environment, and as the viewpoint shifts, the view of the three-dimensional environment will also shift in the viewport. For a head-mounted device, the viewpoint is typically based on the position and orientation of the user's head, face, and / or eyes to provide a perceptually accurate view of the three-dimensional environment that provides an immersive experience while the user is using the head-mounted device. For a handheld or stationary device, the viewpoint shifts as the handheld or stationary device moves and / or as the user's positioning relative to the handheld or stationary device changes (e.g., the user moves toward, away from, up, down, right, and / or left). For a device that includes a display generation component with virtual passthrough, the portions of the physical environment visible (e.g., displayed and / or projected) via the one or more display generation components are based on the field of view of one or more cameras that communicate with the display generation component, the one or more cameras typically moving as the display generation component moves (e.g., for a head-mounted device as the user's head moves, or for a handheld device such as a tablet or smartphone as the user's hand moves), because the user's viewpoint moves as the field of view of the one or more cameras moves (and the appearance of one or more virtual objects displayed via the one or more display generation components is updated based on the user's viewpoint (e.g., the display position and pose of the virtual object are updated based on the movement of the user's viewpoint)).For a display generation component with optical passthrough, portions of the physical environment that are visible via one or more display generation components (e.g., optically visible through one or more partial or fully transparent portions of the display generation component) are based on the user's field of view through the partial or fully transparent portion of the display generation component (e.g., for a head-mounted device, moves as the user's head moves, or for a handheld device such as a tablet or smartphone, moves as the user's hand moves), because the user's viewpoint moves as the user's field of view through the partial or fully transparent portion of the display generation component moves (and the appearance of one or more virtual objects is updated based on the user's viewpoint).

[0055] In some embodiments, the representation of the physical environment (e.g., via virtual passthrough or optical passthrough display) may be partially or fully occluded by the virtual environment. In some embodiments, the amount of the virtual environment displayed (e.g., the amount of the physical environment not displayed) is based on the immersion level of the virtual environment (e.g., relative to the representation of the physical environment). For example, increasing the immersion level optionally causes more of the virtual environment to be displayed, thereby replacing and / or occluding more of the physical environment, and decreasing the immersion level optionally causes less of the virtual environment to be displayed, thereby revealing previously non-displayed and / or previously occluded portions of the physical environment. In some embodiments, at a particular immersion level, one or more first background objects (e.g., in the representation of the physical environment) are more visually de-emphasized (e.g., dimmed, blurred, displayed with increased transparency) compared to one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, the immersion level includes the associated degree to which virtual content (e.g., virtual environment and / or virtual content) displayed by a computer system occludes background content (e.g., content other than the virtual environment and / or virtual content) around / behind the virtual content, optionally including the number of items of the displayed background content and / or the displayed visual characteristics of the background content (e.g., color, contrast, and / or opacity), the angular range of the virtual content displayed by a display generation component (e.g., 60 degrees of content displayed at low immersion, 120 degrees of content displayed at medium immersion, or 180 degrees of content displayed at high immersion), and / or the proportion of the field of view displayed by the display generation component occupied by the virtual content (e.g., 33% of the field of view occupied by the virtual content at low immersion, 66% of the field of view occupied by the virtual content at medium immersion, or 100% of the field of view occupied by the virtual content at high immersion). In some embodiments, the background content includes the background on which the virtual content is displayed (e.g., background content in the representation of the physical environment). In some embodiments, the background content includes a user interface (e.g., a user interface generated by a computer system corresponding to an application), virtual objects not associated with and / or not included in the virtual environment and / or virtual content (e.g., files or representations of other users generated by a computer system, etc.), and / or real objects (e.g., passthrough objects representing real objects in the physical environment around the user, which are visible such that they are displayed by the display generation component and / or visible via a transparent or translucent component of the display generation component because the computer system does not occlude / hinder their visibility through the display generation component). In some embodiments, at a low immersion level (e.g., a first immersion level), the background, virtual, and / or real objects are displayed in an unoccluded manner.For example, a virtual environment with a low level of immersion is optionally displayed concurrently with background content, which is optionally displayed at full brightness, color, and / or semi-transparency. In some embodiments, at a higher level of immersion (e.g., a second level of immersion higher than a first level of immersion), background, virtual, and / or real objects are displayed in an occluded manner (e.g., dimmed, blurred, or removed from the display). For example, a corresponding virtual environment with a high level of immersion is displayed without concurrently displaying background content (e.g., in full-screen or fully immersive mode). As another example, a virtual environment displayed at a medium level of immersion is displayed concurrently with background content that is dimmed, blurred, or otherwise de-emphasized. In some embodiments, the visual characteristics of background objects vary among the background objects. For example, at a particular level of immersion, one or more first background objects are more visually de-emphasized (e.g., dimmed, blurred, and / or displayed with increased transparency) compared to one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, a non-immersion or zero-immersion level corresponds to a virtual environment that ceases to be displayed, and instead a representation of the physical environment is displayed (optionally with one or more virtual objects, such as applications, windows, or virtual three-dimensional objects), and the representation of the physical environment is not occluded by the virtual environment. Adjusting the level of immersion using physical input elements provides a fast and effective way to adjust immersion, which enhances the operability of the computer system and makes the user-device interface more effective.

[0056] Viewpoint-Locked Virtual Object: When a computer system displays a virtual object at the same location and / or orientation in the user's viewpoint, the virtual object is viewpoint-locked even if the user's viewpoint shifts (e.g., changes). In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the forward direction of the user's head (e.g., when the user looks straight ahead, the user's viewpoint is at least a part of the user's field of view); thus, without moving the user's head, the user's viewpoint remains fixed even when the user's gaze shifts. In embodiments where the computer system has a display generation component (e.g., a display screen) that can be repositioned relative to the user's head, the user's viewpoint is an augmented reality view presented to the user on the display generation component of the computer system. For example, a viewpoint-locked virtual object that is displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's viewpoint even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the location and / or orientation at which the viewpoint-locked virtual object is displayed in the user's viewpoint is independent of the user's location and / or orientation in the physical environment. In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head such that the virtual object is also referred to as a "head-locked virtual object".

[0057] Environment-Locked Visual Objects: When a computer system displays a virtual object at a location and / or orientation in the user's field of view, the virtual object is environment-locked (alternatively, "world-locked"), where the location and / or orientation is based on a location and / or object in a three-dimensional environment (e.g., a physical environment or a virtual environment) (e.g., selected and / or anchored to the location and / or object with reference to the location and / or object). As the user's field of view shifts, the location and / or object in the environment relative to the user's field of view changes, which causes the environment-locked virtual object to be displayed at a different location and / or orientation in the user's field of view. For example, an environment-locked virtual object locked to a tree directly in front of the user is displayed at the center of the user's field of view. When the user's field of view shifts to the right (e.g., the user's head turns to the right) such that the tree is now to the left of center in the user's field of view (e.g., the location of the tree in the user's field of view has shifted), the environment-locked virtual object locked to the tree is displayed to the left of center in the user's field of view. In other words, the location and / or orientation at which the environment-locked virtual object is displayed in the user's field of view depends on the location and / or orientation of the location and / or object in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system anchored to a fixed location and / or object in a physical environment) to determine the orientation at which the environment-locked virtual object is displayed in the user's field of view. The environment-locked virtual object can be locked to a stationary part of the environment (e.g., a floor, a wall, a table, or other stationary object), or can be locked to a movable part of the environment (e.g., a vehicle, an animal, a person, or even a representation of a part of the user's body such as a hand, a wrist, an arm, or a foot that moves independently of the user's field of view) such that the virtual object moves as the field of view or that part of the environment moves to maintain a fixed relationship between the virtual object and that part of the environment.

[0058] In some embodiments, an environment-locked or view-locked virtual object exhibits a lazy follow behavior that reduces or delays the movement of the environment-locked or view-locked virtual object relative to the movement of a reference point that the virtual object follows. In some embodiments, when exhibiting the lazy follow behavior, when a movement of a reference point (e.g., a portion of the environment, a view point, or a point fixed relative to the view point, such as a point between 5 cm and 300 cm from the view point) that the virtual object is following is detected, the computer system intentionally delays the movement of the virtual object. For example, when the reference point (e.g., that portion of the environment or the view point) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point, but moves at a second speed that is slower compared to the first speed (e.g., until the reference point stops moving or slows down, at which time the virtual object starts to catch up with the reference point). In some embodiments, when the virtual object exhibits the lazy follow behavior, the device ignores small movements of the reference point (e.g., ignores movements of the reference point below a threshold movement amount, such as a movement of 0 degrees to 5 degrees or 0 cm to 50 cm). For example, when the reference point (e.g., the portion of the environment or the view point to which the virtual object is locked) moves a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed to maintain a fixed or substantially fixed position relative to a view point or a portion of the environment different from the reference point to which the virtual object is locked), and when the reference point (e.g., the portion of the environment or the view point to which the virtual object is locked) moves a second amount that is greater than the first amount, the distance between the reference point and the virtual object initially increases (e.g., because the virtual object is being displayed to maintain a fixed or substantially fixed position relative to a view point or a portion of the environment different from the reference point to which the virtual object is locked), and then decreases when the movement amount of the reference point increases above a threshold (e.g., the "lazy follow" threshold), because the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point. In some embodiments, the virtual object maintaining a substantially fixed position relative to the reference point includes the virtual object being displayed within a threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 15 cm, 20 cm, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward of the position relative to the reference point).

[0059] Hardware: There are many different types of electronic systems that enable a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields integrated with display capabilities, windows integrated with display capabilities, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system can have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system can be configured to receive an external opaque display (e.g., a smartphone). A head-mounted system can incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system can have a transparent or translucent display instead of an opaque display. The transparent or translucent display can have a medium through which light representing an image is directed to a person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light sources, or any combination of these technologies. The medium can be an optical waveguide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display can be configured to selectively become opaque. A projection-based system can employ retinal projection techniques that project graphic images onto a person's retina. The projection system can also be configured to project virtual objects into the physical environment, such as as a hologram or on a physical surface. In some embodiments, the controller 110 is configured to manage and coordinate a user's XR experience. In some embodiments, the controller 110 includes a suitable combination of software, firmware, and / or hardware. Below with respect to Figure 2Controller 110 is described in more detail. In some embodiments, controller 110 is a computing device that is local or remote relative to scene 105 (e.g., a physical environment). For example, controller 110 is a local server located within scene 105. In another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside of scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., an HMD, a display, a projector, a touch screen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, controller 110 is included within the housing (e.g., a physical enclosure) of one or more of display generation component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors, etc.), input device 125, output device 155, one or more sensors in sensor 190, and / or one or more peripheral devices in peripheral device 195, or shares the same physical housing or support structure with one or more of the above devices.

[0060] In some embodiments, display generation component 120 is configured to provide an XR experience to a user (e.g., at least the visual component of the XR experience). In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Display generation component 120 is described in more detail below with respect to Figure 3 In some embodiments, the functionality of controller 110 is provided by and / or in combination with display generation component 120.

[0061] According to some embodiments, display generation component 120 provides an XR experience to a user when the user is virtually and / or physically present within scene 105.

[0062] In some embodiments, the display generation component is worn on a part of the user's body (e.g., on his / her head, on his / her hand, etc.). Accordingly, the display generation component 120 includes one or more XR displays provided to display XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smart phone or a tablet) configured to present XR content, and the user holds the device having a display facing the user's field of view and a camera facing the scene 105. In some embodiments, the handheld device is optionally placed in a housing worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is an XR room, housing, or chamber configured to present XR content, where the user does not wear or hold the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) can be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing an interaction with XR content triggered based on an interaction occurring in the space in front of a handheld device or a tripod-mounted device can be similarly implemented with an HMD, where the interaction occurs in the space in front of the HMD and the response to the XR content is displayed via the HMD. Similarly, a user interface showing an interaction with XR content triggered based on the movement of a handheld device or a tripod-mounted device relative to the physical environment (e.g., the scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)) can be similarly implemented with an HMD, where the movement is caused by the movement of the HMD relative to the physical environment (e.g., the scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)).

[0063] Although relevant features of the operating environment 100 are illustrated in Figure 1A for the sake of brevity and to not obscure more relevant aspects of the example embodiments disclosed herein, various other features are not illustrated.

[0064] Figures 1A to 1PShows various examples of computer systems for performing methods and providing audio, visual, and / or tactile feedback as part of the user interfaces described herein. In some embodiments, the computer system includes one or more display generation components (e.g., first display assembly 1-120a and second display assembly 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b) for displaying virtual elements and / or representations of the physical environment to a user of the computer system, the virtual elements and / or representations of the physical environment being optionally generated based on detected events and / or user input detected by the computer system. The user interface generated by the computer system is optionally corrected by one or more corrective lenses 11.3.2-216, the one or more corrective lenses being optionally removably attached to one or more of the optical modules so that the user interface is more easily viewed by users who would otherwise use glasses or contact lenses to correct their vision. Although many of the user interfaces shown herein show a single view of the user interface, the user interface in the HMD is optionally displayed using two optical modules (e.g., first display assembly 1-120a and second display assembly 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b), one optical module for the user's right eye and a different optical module for the user's left eye, and presenting slightly different images to the two different eyes to create an illusion of stereoscopic depth, the single view of the user interface is typically the right-eye view or the left-eye view, and the depth effect is explained in the text or using other schematic diagrams or views. In some embodiments, the computer system includes one or more external displays (e.g., display assembly 1-108) for displaying status information of the computer system to a user of the computer system (when the computer system is not being worn) and / or to others near the computer system, the status information being optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more audio output components (e.g., electronics 1-112) for generating audio feedback, the audio feedback being optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting information about the physical environment of the device (e.g., one or more sensors in sensor assembly 1-356, and / or Figure 1I ), the one or more sensors being capable of (optionally in combination with one or more illuminators, such as Figure 1IThe illuminator) is used to generate digital pass-through images, capture visual media corresponding to the physical environment (e.g., photos and / or videos), or determine the pose (e.g., position and / or orientation) of physical objects and / or surfaces in the physical environment, such that virtual objects can be placed based on the detected pose of the physical objects and / or surfaces. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting hand positioning and / or movement (e.g., one or more of the sensors in sensor assembly 1-356, and / or Figure 1I ), which can be used (optionally in combination with one or more illuminators, such as Figure 1I the illuminator 6-124 described) to determine when one or more air gestures have been performed. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting eye movement (e.g., Figure 1I the eye tracking and gaze tracking sensors in), which can be used (optionally in combination with one or more lights, such as Figure 1OThe lights in (11.3.2 - 110) determine attention or gaze localization and / or gaze movement, which can optionally be used to detect gaze-only input based on gaze movement and / or dwell. Combinations of the various sensors described above can be used to determine a user's facial expression and / or hand movement for generating a user avatar or representation, such as an anthropomorphic avatar or representation for a real-time communication session, where the avatar has facial expressions, hand movements, and / or body movements based on or similar to the detected facial expressions, hand movements, and / or body movements of the user of the device. Gaze and / or attention information is optionally combined with hand tracking information to determine interaction between the user and one or more user interfaces based on direct and / or indirect input, such as an air gesture or input using one or more hardware input devices, such as one or more buttons (e.g., first button 1 - 128, button 11.1.1 - 114, second button 1 - 132, and / or dial or button 1 - 328), knobs (e.g., first button 1 - 128, button 11.1.1 - 114, and / or dial or button 1 - 328), digital crowns (e.g., a first button 1 - 128, button 11.1.1 - 114, and / or dial or button 1 - 328 that is pressable and twistable or rotatable), touchpads, touchscreens, keyboards, mice, and / or other input devices. One or more buttons (e.g., first button 1 - 128, button 11.1.1 - 114, second button 1 - 132, and / or dial or button 1 - 328) are optionally used to perform system operations, such as re-centering content in a three-dimensional environment visible to the user of the device, displaying a main user interface for launching an application, starting a real-time communication session, or initiating the display of a virtual three-dimensional background. A knob or digital crown (e.g., a first button 1 - 128, button 11.1.1 - 114, and / or dial or button 1 - 328 that is pressable and twistable or rotatable) is optionally rotatable to adjust parameters of visual content, such as the level of immersion of a virtual three-dimensional environment (e.g., the extent to which virtual content occupies the user's viewport in the three-dimensional environment) or other parameters associated with the three-dimensional environment and virtual content displayed via an optical module (e.g., first display component 1 - 120a and second display component 1 - 120b and / or first optical module 11.1.1 - 104a and second optical module 11.1.1 - 104b).

[0065] Figure 1BA front top perspective view of an example of a head - mountable display (HMD) device 1 - 100 configured to be worn by a user and provide virtual and augmented reality (VR / AR) experiences is illustrated. The HMD 1 - 100 may include a display unit 1 - 102 or component, an electronic strip assembly 1 - 104 connected to and extending from the display unit 1 - 102, and a strap assembly 1 - 106 fixed to the electronic strip assembly 1 - 104 at either end. The electronic strip assembly 1 - 104 and the strap 1 - 106 may be part of a retention assembly configured to wrap around the user's head to hold the display unit 1 - 102 against the user's face.

[0066] In at least one example, the strap assembly 1 - 106 may include a first strap 1 - 116 configured to wrap around the back side of the user's head and a second strap 1 - 117 configured to extend over the top of the user's head. As shown, the second strap may extend between a first electronic strip 1 - 105a and a second electronic strip 1 - 105b of the electronic strip assembly 1 - 104. The strip assembly 1 - 104 and the strap assembly 1 - 106 may be part of a fixation mechanism that extends rearward from the display unit 1 - 102 and is configured to hold the display unit 1 - 102 against the user's face.

[0067] In at least one example, the fixation mechanism includes a first electronic strip 1 - 105a that includes a first proximal end 1 - 134 coupled to the display unit 1 - 102 (e.g., the housing 1 - 150 of the display unit 1 - 102) and a first distal end 1 - 136 opposite the first proximal end 1 - 134. The fixation mechanism may also include a second electronic strip 1 - 105b that includes a second proximal end 1 - 138 coupled to the housing 1 - 150 of the display unit 1 - 102 and a second distal end 1 - 140 opposite the second proximal end 1 - 138. The fixation mechanism may also include a first strap 1 - 116 and a second strap 1 - 117, the first strap including a first end 1 - 142 coupled to the first distal end 1 - 136 and a second end 1 - 144 coupled to the second distal end 1 - 140, and the second strap extending between the first electronic strip 1 - 105a and the second electronic strip 1 - 105b. The strips 1 - 105a - b and the strap 1 - 116 may be coupled via a connection mechanism or component 1 - 114. In at least one example, the second strap 1 - 117 includes a first end 1 - 146 coupled to the first electronic strip 1 - 105a between the first proximal end 1 - 134 and the first distal end 1 - 136 and a second end 1 - 148 coupled to the second electronic strip 1 - 105b between the second proximal end 1 - 138 and the second distal end 1 - 140.

[0068] In at least one example, the first electronic strip 1-105a and the second electronic strip 1-105b include plastic, metal, or other structural materials that form the shape of substantially rigid strips 1-105a-1-105b. In at least one example, the first band 1-116 and the second band 1-117 are formed of a resilient flexible material including a woven textile, rubber, etc. The first band 1-116 and the second band 1-117 can be flexible to conform to the shape of the user's head when the HMD 1-100 is worn.

[0069] In at least one example, one or more of the first electronic strip 1-105a and the second electronic strip 1-105b can define an inner strip volume and include one or more electronic components disposed in the inner strip volume. Figure 1B As shown, the first electronic strip 1-105a may include an electronic component 1-112. In one example, the electronic component 1-112 may include a speaker. In one example, the electronic component 1-112 may include a computing component, such as a processor.

[0070] In at least one example, the housing 1-150 defines a first front opening 1-152. Figure 1B 1-152 in dashed lines because the display assembly 1-108 is configured to shield the first opening 1-152 from view when the HMD 1-100 is assembled. The housing 1-150 may also define a rear-mounted second opening 1-154. The housing 1-150 also defines an internal volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display assembly 1-108, which may include a front cover and a display screen (shown in other figures) disposed in or across the front opening to shield the front opening 1-152. In at least one example, the display screen of the display assembly 1-108, and the display assembly 1-108 in general, has a curvature configured to follow the curvature of the user's face. The display screen of the display assembly 1-108 may be curved as shown to complement the user's facial features and the overall curvature from one side of the face to the other, such as from left to right and / or from top to bottom, where the display unit 1-102 is pressed.

[0071] In at least one example, the housing 1-150 may define a first aperture 1-126 between a first opening 1-152 and a second opening 1-154 and a second aperture 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 may also include a first button 1-126 disposed in the first aperture 1-128 and a second button 1-132 disposed in the second aperture 1-130. The first button 1-128 and the second button 1-132 are capable of being pressed through the respective apertures 1-126, 1-130. In at least one example, the first button 1-126 and / or the second button 1-132 may be a twist dial as well as a push button. In at least one example, the first button 1-128 is a pushable and twistable dial button and the second button 1-132 is a push button.

[0072] Figure 1C A rear perspective view of the HMD 1-100 is illustrated. The HMD 1-100 may include a light seal 1-110 that extends rearwardly from the housing 1-150 of the display assembly 1-108 around a perimeter of the housing 1-150 as shown. The light seal 1-110 may be configured to extend from the housing 1-150 to a user's face and around the user's eyes to block external light from being visible. In one example, the HMD 1-100 may include a first display assembly 1-120a and a second display assembly 1-120b that are disposed at or within a second opening 1-154 defined by the housing 1-150 that faces rearward and / or are disposed within an interior volume of the housing 1-150 and are configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a-1-120ab may include a respective display screen 1-122a, 1-122b that is configured to project light through the second opening 1-154 in a rearward direction toward the user's eyes.

[0073] In at least one example, referring Figure 1B and Figure 1C both, the display assembly 1-108 may be a front forward display assembly that includes a display screen configured to project light in a first forward direction, and the rear display screens 1-122a-1-122ab may be configured to project light in a second rearward direction that is opposite the first direction. As noted above, the light seal 1-110 may be configured to block light external to the HMD 1-100 (including light emitted by Figure 1BThe light projected by the forward display screen of the display assembly 1-108 shown in the front perspective view reaches the user's eyes. In at least one example, the HMD 1-100 may also include a curtain 1-124 for the second opening 1-154 between the shielding housing 1-150 and the rear display assembly 1-120a-1-120b. In at least one example, the curtain 1-124 may be elastic or at least partially elastic.

[0074] Figure 1B and Figure 1C Any of the features, components, and / or parts shown (including their arrangements and configurations) may be included individually or in any combination in Figures 1D to 1F any other examples of the devices, features, components, and parts shown and described herein. Similarly, reference Figures 1D to 1F to any of the features, components, and / or parts shown or described (including their arrangements and configurations) may be included individually or in any combination in Figure 1B and Figure 1C the examples of the devices, features, components, and parts shown.

[0075] Figure 1D An exploded view illustrating an example of an HMD 1-200 including various parts or components, which are separated according to the modularity and selective coupling of these components. For example, the HMD 1-200 may include a strap 1-216, which may be selectively coupled to a first electronic strip 1-205a and a second electronic strip 1-205b. The first fixed strip 1-205a may include a first electronic component 1-212a, and the second fixed strip 1-205b may include a second electronic component 1-212b. In at least one example, the first strip 1-205a and the second strip 1-205b can be removably coupled to the display unit 1-202.

[0076] In addition, the HMD 1-200 may include a light seal 1-210 configured to be removably coupled to the display unit 1-202. The HMD 1-200 may also include a lens 1-218, which may be removably coupled to the display unit 1-202, for example, on a first component including a display screen and a second display component. The lens 1-218 may include a custom prescription lens configured for vision correction. As noted, in Figure 1DIn the exploded view shown and described above, each of the parts can be removably coupled, attached, reattached, and replaced to update the parts or swap out parts for different users. For example, straps such as straps 1-216, light seals such as light seals 1-210, lenses such as lens 1-218, and electronic strips such as electronic strips 1-205a-1-205b can be swapped out according to the user so that these parts are customized to fit and correspond to a single user of the HMD 1-200.

[0077] Figure 1D Any of the features, components, and / or parts shown (including their arrangements and configurations) can be included individually or in any combination in Figure 1B , Figure 1C and Figures 1E to 1F any other examples of the devices, features, components, and parts shown and described herein. Similarly, reference Figure 1B , Figure 1C and Figures 1E to 1F any of the features, components, and / or parts shown or described (including their arrangements and configurations) can be included individually or in any combination in Figure 1D the examples of the devices, features, components, and parts shown.

[0078] Figure 1E An exploded view illustrating an example of the display unit 1-306 of the HMD is shown. The display unit 1-306 can include a front display assembly 1-308, a frame / housing assembly 1-350, and a curtain assembly 1-324. The display unit 1-306 can also include a sensor assembly 1-350, a logic board assembly 1-358, and a cooling assembly 1-360 disposed between the frame assembly 1-356 and the front display assembly 1-308. In at least one example, the display unit 1-306 can also include a rear display assembly 1-320 that includes a first rear display screen 1-322a and a second rear display screen 1-322b disposed between the frame 1-350 and the curtain assembly 1-324.

[0079] In at least one example, the display unit 1-306 can also include a motor assembly 1-362 configured as an adjustment mechanism for adjusting the positioning of the display screens 1-322a-b of the display assembly 1-320 relative to the frame 1-350. In at least one example, the display assembly 1-320 is mechanically coupled to the motor assembly 1-362, and each display screen 1-322a-b has at least one motor such that the motors can translate the display screens 1-322a-b to match the pupil spacing of the user's eyes.

[0080] In at least one example, the display unit 1-306 may include a dial or button 1-328 that is pressable relative to the frame 1-350 and accessible by a user external to the frame 1-350. The button 1-328 may be electronically connected to the motor assembly 1-362 via a controller such that the button 1-328 can be manipulated by the user to cause the motor of the motor assembly 1-362 to adjust the positioning of the display screens 1-322a-b.

[0081] Figure 1E Any of the illustrated features, components, and / or parts (including their arrangements and configurations) may be included, either individually or in any combination, in Figures 1B to 1D and Figure 1F any other examples of the devices, features, components, and parts shown and described herein. Similarly, reference Figures 1B to 1D and Figure 1F to any of the illustrated and described features, components, and / or parts (including their arrangements and configurations) may be included, either individually or in any combination, in Figure 1E the examples of the devices, features, components, and parts shown.

[0082] Figure 1F An exploded view of another example of a display unit 1-406 of an HMD device similar to other HMD devices described herein is illustrated. The display unit 1-406 may include a front display assembly 1-402, a sensor assembly 1-456, a logic board assembly 1-458, a cooling assembly 1-460, a frame assembly 1-450, a rear display assembly 1-421, and a curtain assembly 1-424. The display unit 1-406 may also include a motor assembly 1-462 for adjusting the positioning of a first display sub-assembly 1-420a and a second display sub-assembly 1-420b of the rear display assembly 1-421, including a first corresponding display screen and a second corresponding display screen for inter-pupillary adjustment, as described above.

[0083] Figure 1F The various parts, systems, and components shown in the exploded view are described in more detail herein with reference to Figures 1B to 1E and the subsequent figures referenced in this disclosure. Figure 1F The illustrated display unit 1-406 may be assembled and integrated with Figures 1B to 1E the illustrated fixing mechanism that includes an electronic strip, a band, and other components including a light seal, a connection assembly, etc.

[0084] Figure 1F Any of the illustrated features, components, and / or parts (including their arrangements and configurations) may be included, either individually or in any combination, in Figures 1B to 1E any other examples of the devices, features, components, and parts shown and described herein. Similarly, reference Figures 1B to 1EAny of the features, components, and / or parts shown and described (including their arrangements and configurations) may be included individually or in any combination in Figure 1F the examples of the devices, features, components, and parts shown

[0085] Figure 1G FIG. 6 illustrates an exploded perspective view of a front cover assembly 3-100 of the HMD device described herein, such as Figure 1G the front cover assembly 3-1 of the HMD 3-100 shown or any other HMD device shown and described herein. Figure 1G The front cover assembly 3-100 shown may include a transparent or translucent cover 3-102, a shield 3-104 (or “hood”), an adhesive layer 3-106, a display assembly 3-108 including a bi-convex lens panel or array 3-110, and a structural trim 3-112. The adhesive layer 3-106 may secure the shield 3-104 and / or the transparent cover 3-102 to the display assembly 3-108 and / or the trim 3-112. The trim 3-112 may secure the various components of the front cover assembly 3-100 to the frame or base of the HMD device.

[0086] In at least one example, as Figure 1G shown, the transparent cover 3-102, the shield 3-104, and the display assembly 3-108 including the bi-convex lens array 3-110 may be bendable to conform to the curvature of the user's face. The transparent cover 3-102 and the shield 3-104 may be bent in two or three dimensions, e.g., bent vertically in the Z direction inside and outside the Z-X plane, and bent horizontally in the X direction inside and outside the Z-X plane. In at least one example, the display assembly 3-108 may include the bi-convex lens array 3-110 and a display panel having pixels configured to project light through the shield 3-104 and the transparent cover 3-102. The display assembly 3-108 may be bent in at least one direction (e.g., the horizontal direction) to conform to the curvature of the user's face from one side (e.g., the left side) to the other side (e.g., the right side) of the face. In at least one example, each layer or component of the display assembly 3-108 (which will be shown and described in more detail in subsequent figures, but which may include the bi-convex lens array 3-110 and the display layer) may be similarly or concentrically bent in the horizontal direction to conform to the curvature of the user's face.

[0087] In at least one example, the shroud 3-104 may include a transparent or translucent material through which the display component 3-108 projects light. In one example, the shroud 3-104 may include one or more opaque portions, such as an opaque ink printed portion or other opaque film portion on the back surface of the shroud 3-104. When the HMD device is worn, the rear surface may be the surface of the shroud 3-104 facing the user's eyes. In at least one example, the opaque portion may be on the front surface of the shroud 3-104 opposite the rear surface. In at least one example, one or more opaque portions of the shroud 3-104 may include a peripheral portion that visually hides any components around the outer perimeter of the display screen of the display component 3-108. In this way, the opaque portions of the shroud hide any other components of the HMD device that would otherwise be visible through the transparent or translucent cover 3-102 and / or the shroud 3-104, including electronic components, structural components, and the like.

[0088] In at least one example, the shroud 3-104 may define one or more apertured transparent portions 3-120 through which sensors may transmit and receive signals. In one example, portion 3-120 is an aperture through which the sensor may extend or through which the sensor may transmit and receive signals. In one example, portion 3-120 is a transparent portion, or a portion that is more transparent than the surrounding translucent or opaque portions of the shroud, through which the sensor may transmit and receive signals through the shroud and through the transparent cover 3-102. In one example, the sensors may include a camera, an IR sensor, a LUX sensor, or any other visual or non-visual environmental sensor of the HMD device.

[0089] Figure 1G Any of the illustrated features, components, and / or parts (including their arrangements and configurations) may be included, either alone or in any combination, in any other examples of the devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts (including their arrangements and configurations) illustrated and described herein may be included, either alone or in any combination, in Figure 1G the examples of the devices, features, components, and parts illustrated.

[0090] Figure 1H An exploded view of an example of an HMD device 6-100 is illustrated. The HMD device 6-100 may include a sensor array or system 6-102 that includes one or more sensors, cameras, projectors, etc. mounted to one or more components of the HMD 6-100. In at least one example, the sensor system 6-102 may include a bracket 1-338 to which one or more sensors of the sensor system 6-102 may be fixed / fastened.

[0091] Figure 1I Illustrates a portion of an HMD device 6-100 including a front transparent cover 6-104 and a sensor system 6-102. The sensor system 6-102 may include a plurality of different sensors, transmitters, receivers, including cameras, IR sensors, projectors, etc. The transparent cover 6-104 is shown in front of the sensor system 6-102 to show the relative positioning of the various sensors and transmitters and the orientation of each sensor / transmitter of the system 6-102. As referred to herein, "lateral", "side", "transverse", "horizontal", and other similar terms refer to the orientation or direction indicated by the X-axis as Figure 1J shown. Terms such as "vertical", "upward", "downward", and similar terms refer to the orientation or direction indicated by the Z-axis as Figure 1J shown. Terms such as "forward", "backward", "frontward", "backward", and similar terms refer to the orientation or direction indicated by the Y-axis as Figure 1J shown.

[0092] In at least one example, the transparent cover 6-104 may define the front outer surface of the HMD device 6-100, and the sensor system 6-102 including various sensors and their components may be disposed behind the cover 6-104 in the Y-axis / direction. The cover 6-104 may be transparent or translucent to allow light to pass through the cover 6-104, including both light detected by the sensor system 6-102 and light emitted therefrom.

[0093] As described elsewhere herein, the HMD device 6-100 may include one or more controllers that include processors for electrically coupling the various sensors and transmitters of the sensor system 6-102 to one or more motherboards, processing units, and other electronic devices such as display screens, etc. In addition, as will be shown in more detail with reference to other figures below, the various sensors, transmitters, and other components of the sensor system 6-102 may be coupled to Figure 1I various structural frame members, brackets, etc. of the HMD device 6-100 not shown. For clarity, Figure 1I illustrates that the components of the sensor system 6-102 are not attached and not electrically coupled to other components.

[0094] In at least one example, the device may include one or more controllers having a processor configured to execute instructions stored on a memory component electrically coupled to the processor. These instructions may include or cause the processor to execute one or more algorithms for self-correcting the angles and positions of the various cameras described herein over time as the initial positioning, angle, or orientation of the camera undergoes collisions or deformations due to accidental drop events or other events.

[0095] In at least one example, the sensor system 6-102 can include one or more scene cameras 6-106. The system 6-102 can include two scene cameras 6-102, respectively disposed on both sides of the bridge of the nose or the arch structure of the HMD device 6-100, such that each of the two cameras 6-106 is positioned generally corresponding to the left and right eyes of the user behind the cover 6-103. In at least one example, the scene cameras 6-106 are generally oriented forward in the Y direction to capture images in front of the user during the use of the HMD 6-100. In at least one example, the scene cameras are color cameras and, when the HMD device 6-100 is used, provide images and content for MR video passthrough to the display screen facing the user's eyes. The scene cameras 6-106 can also be used for environment and object reconstruction.

[0096] In at least one example, the sensor system 6-102 can include a first depth sensor 6-108 that is generally pointed forward in the Y direction. In at least one example, the first depth sensor 6-108 can be used for environment and object reconstruction and for tracking the user's hands and body. In at least one example, the sensor system 6-102 can include a second depth sensor 6-110 that is centered along the width of the HMD device 6-100 (e.g., along the X axis). For example, the second depth sensor 6-110 can be disposed above the central bridge of the nose or on an adapter structure above the nose when the user wears the HMD 6-100. In at least one example, the second depth sensor 6-110 can be used for environment and object reconstruction and for hand and body tracking. In at least one example, the second depth sensor can include a LIDAR sensor.

[0097] In at least one example, the sensor system 6-102 can include a depth projector 6-112 that is generally oriented forward to project electromagnetic waves (e.g., in the form of a pre-determined pattern of light points) into the field of view or within the field of view of the user and / or the scene cameras 6-106, or into a field of view that includes and extends beyond the field of view of the user and / or the scene cameras 6-106. In at least one example, the depth projector is capable of projecting electromagnetic waves of light in the form of a pattern of light points that are reflected from objects and return to the above-mentioned depth sensors, including depth sensors 6-108, 6-110. In at least one example, the depth projector 6-112 can be used for environment and object reconstruction and for hand and body tracking.

[0098] In at least one example, the sensor system 6-102 can include a downward-facing camera 6-114, whose field of view generally points downward relative to the HDM device 6-100 on the Z-axis. In at least one example, the downward camera 6-114 can be disposed on the left and right sides of the HMD device 6-100 as shown and is used for hand and body tracking, headset tracking, and face avatar detection and creation for displaying a user avatar on the forward display screen of the HMD device 6-100 described elsewhere herein. For example, the downward camera 6-114 can be used to capture the facial expressions and movements of the user's face below the HMD device 6-100, including the cheeks, mouth, and chin.

[0099] In at least one example, the sensor system 6-102 can include a jaw camera 6-116. In at least one example, the jaw camera 6-116 can be disposed on the left and right sides of the HMD device 6-100 as shown and is used for hand and body tracking, headset tracking, and face avatar detection and creation for displaying a user avatar on the forward display screen of the HMD device 6-100 described elsewhere herein. For example, the jaw camera 6-116 can be used to capture the facial expressions and movements of the user's face below the HMD device 6-100, including the user's jaw, cheeks, mouth, and chin. For hand and body tracking, headset tracking, and face avatar

[0100] In at least one example, the sensor system 6-102 can include a side camera 6-118. The side camera 6-118 can be oriented to capture left and right views in the X-axis or in the direction relative to the HMD device 6-100. In at least one example, the side camera 6-118 can be used for hand and body tracking, headset tracking, and face avatar detection and recreation.

[0101] In at least one example, the sensor system 6-102 can include a plurality of eye tracking and gaze tracking sensors for determining the identity, status, and gaze direction of the user's eyes during and / or before use. In at least one example, the eye / gaze tracking sensors can include a nose-eye camera 6-120, which is disposed on either side of the user's nose and near the user's nose when wearing the HMD device 6-100. The eye / gaze sensors can also include a bottom eye camera 6-122 disposed below the corresponding user's eye for capturing an image of the eye for face avatar detection and creation, gaze tracking, and iris identification functions.

[0102] In at least one example, the sensor system 6-102 can include an infrared illuminator 6-124 that points outward from the HMD device 6-100 to illuminate the external environment and any objects therein with IR light for IR detection using one or more IR sensors of the sensor system 6-102. In at least one example, the sensor system 6-102 can include a flicker sensor 6-126 and an ambient light sensor 6-128. In at least one example, the flicker sensor 6-126 can detect the top light refresh rate to avoid display flicker. In one example, the infrared illuminator 6-124 can include a light emitting diode and can be particularly used in low light environments to illuminate the user's hand and other objects in low light for detection by the infrared sensors of the sensor system 6-102.

[0103] In at least one example, multiple sensors (including the scene camera 6-106, the downward camera 6-114, the jaw camera 6-116, the side camera 6-118, the depth projector 6-112, and the depth sensors 6-108, 6-110) can be used in combination with an electrically coupled controller to combine depth data with camera data for hand tracking and for size determination to better perform the hand tracking and object recognition and tracking functions of the HMD device 6-100. In at least one example, the downward camera 6-114, the jaw camera 6-116, and the side camera 6-118 described above and shown in Figure 1I can be wide-angle cameras that can operate in the visible and infrared spectra. In at least one example, these cameras 6-114, 6-116, 6-118 can operate only in black and white light detection to simplify image processing and obtain sensitivity.

[0104] Figure 1I Any of the features, components, and / or parts shown (including their arrangement and configuration) can be included alone or in any combination in Figures 1J to 1L any other example of the devices, features, components, and parts shown and described herein. Similarly, any of the features, components, and / or parts shown and described with reference to Figures 1J to 1L can be included alone or in any combination in Figure 1I the example of the devices, features, components, and parts shown.

[0105] Figure 1JA bottom perspective view of an example of an HMD 6-200 is illustrated that includes a cover or shield 6-204 fixed to a frame 6-230. In at least one example, sensors 6-202 of a sensor system 6-203 can be disposed around the perimeter of the HDM 6-200 such that the sensors 6-203 are disposed outwardly around the perimeter of a display area or zone 6-232 so as not to obstruct the viewing of the displayed light. In at least one example, the sensors can be disposed behind the shield 6-204 and aligned with a transparent portion of the shield, thereby allowing the sensors and projectors to allow light to pass back and forth through the shield 6-204. In at least one example, an opaque ink or other opaque material or film / layer can be disposed on the shield 6-204 around the display zone 6-232 to hide components of the HMD 6-200 outside the display zone 6-232 other than the transparent portion defined by the opaque portion, through which the sensors and projectors transmit and receive light and electromagnetic signals during operation. In at least one example, the shield 6-204 allows light to pass through from a display (e.g., within the display area 6-232), but does not allow light to pass radially outward from the display area around the perimeter of the display and the shield 6-204.

[0106] In some examples, the shield 6-204 includes a transparent portion 6-205 and an opaque portion 6-207, as described above and elsewhere herein. In at least one example, the opaque portion 6-204 of the shield 6-207 can define one or more transparent regions 6-209 through which sensors 6-203 of the sensor system 6-202 can transmit and receive signals. In the illustrated example, sensors 6-203 of the sensor system 6-202 transmit and receive signals through the shield 6-204, or more specifically through the transparent regions 6-209 (or defined thereby) of the opaque portion 6-207 of the shield 6-204, and the sensors can include the same or similar sensors as those shown in the examples of Figure 1I , such as depth sensors 6-108 and 6-110, depth projectors 6-112, a first scene camera and a second scene camera 6-106, a first downward camera and a second downward camera 6-114, a first side camera and a second side camera 6-118, and a first infrared illuminator and a second infrared illuminator 6-124. These sensors are also shown in the examples of Figure 1K and Figure 1L . Other sensors, sensor types, sensor quantities, and their relative positioning can be included in one or more other examples of the HMD.

[0107] Figure 1J Any of the features, components, and / or parts shown (including their arrangement and configuration) can be included individually or in any combination in Figure 1I and Figures 1K to 1LIn any other examples of the devices, features, components, and parts shown and described herein. Similarly, reference Figure 1I and Figures 1K to 1L Any one of the features, components, and / or parts shown or described (including their arrangements and configurations) may be included individually or in any combination in Figure 1J the examples of the devices, features, components, and parts shown herein.

[0108] Figure 1K A front view illustrating a part of an example of the HMD device 6-300, including a display 6-334, brackets 6-336, 6-338, and a frame or housing 6-330. Figure 1K The example shown does not include a front cover or shield in order to illustrate the brackets 6-336, 6-338. For example, Figure 1J the shield 6-204 shown includes an opaque portion 6-207 that will visually cover / block the viewing of anything external to the display / display area 6-334 (e.g., radially / peripherally external), including the sensor 6-303 and the bracket 6-338.

[0109] In at least one example, the various sensors of the sensor system 6-302 are coupled to the brackets 6-336, 6-338. In at least one example, the scene cameras 6-306 include tight tolerances on the angles relative to each other. For example, the tolerance on the mounting angle between two scene cameras 6-306 can be 0.5 degrees or less, such as 0.3 degrees or less. To achieve and maintain such tight tolerances, in one example, the scene cameras 6-306 may be mounted to the bracket 6-338 rather than the shield. The bracket may include a cantilever on which the scene cameras 6-306 and other sensors of the sensor system 6-302 may be mounted to maintain their position and orientation unchanged in the event of a drop where other brackets 6-226, the housing 6-330, and / or the shield are deformed by the user.

[0110] Figure 1K Any one of the features, components, and / or parts shown or described (including their arrangements and configurations) may be included individually or in any combination in Figures 1I to 1J and Figure 1L any other examples of the devices, features, components, and parts shown and described herein. Similarly, reference Figures 1I to 1J and Figure 1L Any one of the features, components, and / or parts shown or described (including their arrangements and configurations) may be included individually or in any combination in Figure 1K the examples of the devices, features, components, and parts shown in

[0111] Figure 1LA bottom view of an example of an HMD 6-400 is illustrated that includes a front display / cover assembly 6-404 and a sensor system 6-402. The sensor system 6-402 can be similar to other sensor systems described above and elsewhere herein, including reference Figures 1I to 1K as described. In at least one example, the chin camera 6-416 can face downward to capture images of the user's lower facial features. In one example, the chin camera 6-416 can be directly coupled to the frame or housing 6-430 or one or more internal brackets that are directly coupled to the illustrated frame or housing 6-430. The frame or housing 6-430 can include one or more holes / openings 6-415 through which the chin camera 6-416 can transmit and receive signals.

[0112] Figure 1L Any of the illustrated features, components, and / or parts (including their arrangements and configurations) can be included individually or in any combination in Figures 1I to 1K any other example of the devices, features, components, and parts shown and described herein. Similarly, reference Figures 1I to 1K Any of the illustrated features, components, and / or parts (including their arrangements and configurations) can be included individually or in any combination in Figure 1L the example of the devices, features, components, and parts shown.

[0113] Figure 1M A rear perspective view of an interpupillary distance (IPD) adjustment system 11.1.1-102 is illustrated that includes first and second optical modules 11.1.1-104a and 11.1.1-104b that are slidably engaged / coupled to respective guide rods 11.1.1-108a - 11.1.1-108b and motors 11.1.1-110a - 11.1.1-110b of a left adjustment subsystem 11.1.1-106a and a right adjustment subsystem 11.1.1-106b. The IPD adjustment system 11.1.1-102 can be coupled to a bracket 11.1.1-112 and includes buttons 11.1.1-114 that are in electrical communication with the motors 11.1.1-110a - 11.1.1-110b. In at least one example, the buttons 11.1.1-114 can be in electrical communication with the first motor 11.1.1-110a and the second motor 11.1.1-110b via a processor or other circuit components such that the first motor 11.1.1-110a and the second motor 11.1.1-110b are activated and cause the first optical module 11.1.1-104a and the second optical module 11.1.1-104b to change their positions relative to each other, respectively.

[0114] In at least one example, the first optical module 11.1.1-104a and the second optical module 11.1.1-104b may include respective display screens configured to project light toward a user's eyes when the HMD 11.1.1-100 is worn. In at least one example, a user may manipulate (e.g., press and / or rotate) button 11.1.1-114 to activate positioning adjustment of the optical modules 11.1.1-104a-11.1.1-104b to match the user's interpupillary distance (IPD). The optical modules 11.1.1-104a-11.1.1-104b may also include one or more cameras or other sensors / sensor systems for imaging and measuring the user's IPD such that the optical modules 11.1.1-104a-11.1.1-104b may be adjusted to match the IPD.

[0115] In one example, a user may manipulate button 11.1.1-114 to cause automatic positioning adjustment of the first optical module 11.1.1-104a and the second optical module 11.1.1-104b. In one example, a user may manipulate button 11.1.1-114 to cause manual adjustment such that the optical modules 11.1.1-104a-11.1.1-104b move farther or closer (e.g., when the user rotates button 11.1.1-114 in one way or another) until the user visually matches her / his own IPD. In one example, the manual adjustment is communicated electronically via one or more circuits and power for moving the optical modules 11.1.1-104a-11.1.1-104b via motors 11.1.1-110a-11.1.1-110b is provided by a power source. In one example, adjustment and movement of the optical modules 11.1.1-104a-11.1.1-104b via manipulation of button 11.1.1-114 is actuated mechanically via movement of button 11.1.1-114.

[0116] Figure 1M Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included, either alone or in any combination, in any other example of a device, feature, component, and part shown in any of the other figures shown and described herein. Similarly, any of the features, components, and / or parts shown or described (including their arrangement and configuration) with reference to any of the other figures shown and described herein may be included, either alone or in any combination, in Figure 1M the example of a device, feature, component, and part shown.

[0117] Figure 1NFront perspective view of a portion of HMD 11.1.2-100, including an outer structural frame 11.1.2-102 and an inner or intermediate structural frame 11.1.2-104 that define a first aperture 11.1.2-106a and a second aperture 11.1.2-106b. The apertures 11.1.2-106a - 11.1.2-106b are shown in Figure 1N dashed lines in, because the view of the apertures 11.1.2-106a - 11.1.2-106b may be blocked by one or more other components of the HMD 11.1.2-100 that are coupled to the inner frame 11.1.2-104 and / or the outer frame 11.1.2-102, as shown. In at least one example, the HMD 11.1.2-100 may include a first mounting bracket 11.1.2-108 that is coupled to the inner frame 11.1.2-104. In at least one example, the mounting bracket 11.1.2-108 is coupled to the inner frame 11.1.2-104 between the first aperture 11.1.2-106a and the second aperture 11.1.2-106b.

[0118] The mounting bracket 11.1.2-108 may include an intermediate or central portion 11.1.2-109 that is coupled to the inner frame 11.1.2-104. In some examples, the intermediate or central portion 11.1.2-109 may not be the geometric middle or center of the bracket 11.1.2-108. Instead, the intermediate / central portion 11.1.2-109 may be disposed between a first cantilevered extension arm and a second cantilevered extension arm that extend away from the intermediate portion 11.1.2-109. In at least one example, the mounting bracket 108 includes a first cantilever 11.1.2-112 and a second cantilever 11.1.2-114 that extend away from the intermediate portion 11.1.2-109 of the mounting bracket 11.1.2-108 that is coupled to the inner frame 11.1.2-104.

[0119] As Figure 1NAs shown, the outer frame 11.1.2-102 can define a curved geometry on its lower side to accommodate the user's nose when the user wears the HMD 11.1.2-100. The curved geometry can be referred to as the nose bridge 11.1.2-111 and is centered on the lower side of the HMD 11.1.2-100 as shown. In at least one example, the mounting bracket 11.1.2-108 can be connected to the inner frame 11.1.2-104 between the holes 11.1.2-106a - 11.1.2-106b such that the cantilever arms 11.1.2-112, 11.1.2-114 extend downward and laterally outward away from the middle portion 11.1.2-109 to be complementary to the geometry of the nose bridge 11.1.2-111 of the outer frame 11.1.2-102. In this way, the mounting bracket 11.1.2-108 is configured to accommodate the user's nose, as described above. The geometry of the nose bridge 11.1.2-111 accommodates the nose because the nose bridge 11.1.2-111 provides a curvature that conforms to the shape of the user's nose, providing a comfortable fit from above, over, and around.

[0120] The first cantilever arm 11.1.2-112 can extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a first direction, and the second cantilever arm 11.1.2-114 can extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-10 in a second direction opposite to the first direction. The first cantilever arm 11.1.2-112 and the second cantilever arm 11.1.2-114 are referred to as "cantilevered" or "cantilever" arms because each arm 11.1.2-112, 11.1.2-114 includes a free distal end 11.1.2-116, 11.1.2-118 respectively that is not attached to the inner frame 11.1.2-102 and the outer frame 11.1.2-104. In this way, the arms 11.1.2-112, 11.1.2-114 overhang from the middle portion 11.1.2-109 which can be connected to the inner frame 11.1.2-104 while the distal ends 11.1.2-102, 11.1.2-104 are not attached.

[0121] In at least one example, the HMD 11.1.2-100 may include one or more components coupled to the mounting bracket 11.1.2-108. In one example, the components include a plurality of sensors 11.1.2-110a-11.1.2-110f. Each of the plurality of sensors 11.1.2-110a-11.1.2-110f may include various types of sensors, including cameras, IR sensors, etc. In some examples, one or more of the sensors 11.1.2-110a-11.1.2-110f may be used for object recognition in three-dimensional space, such that it is important to maintain the precise relative positioning of two or more of the plurality of sensors 11.1.2-110a-11.1.2-110f. The cantilever nature of the mounting bracket 11.1.2-108 may protect the sensors 11.1.2-110a-11.1.2-110f from damage and misalignment in the event of an accidental drop by the user. Since the sensors 11.1.2-110a-11.1.2-110f are cantilevered on the arms 11.1.2-112, 11.1.2-114 of the mounting bracket 11.1.2-108, the stress and deformation of the inner frame and / or the outer frames 11.1.2-104, 11.1.2-102 are not transferred to the cantilevers 11.1.2-112, 11.1.2-114 and thus do not affect the relative positioning of the sensors 11.1.2-110a-11.1.2-110f coupled / mounted to the mounting bracket 11.1.2-108.

[0122] Figure 1N Any of the illustrated features, components, and / or parts (including their arrangement and configuration) may be included, either individually or in any combination, in any other example of the devices, features, components described herein. Similarly, any of the features, components, and / or parts (including their arrangement and configuration) shown and described herein may be included, either individually or in any combination, in Figure 1N the examples of the devices, features, components, and parts shown.

[0123] Figure 1O An example of an optical module 11.3.2-100 for use in an electronic device (such as an HMD, including the HDM devices described herein) is illustrated. As shown in one or more other examples described herein, the optical module 11.3.2-100 may be one of two optical modules within the HMD, where each optical module is aligned to project light toward the user's eyes. In this manner, the first optical module may project light to the user's first eye via a display screen, and the second optical module of the same device may project light to the user's second eye via another display screen.

[0124] In at least one example, the optical module 11.3.2-100 may include an optical frame or housing 11.3.2-102, which may also be referred to as a barrel or an optical module barrel. The optical module 11.3.2-100 may further include a display 11.3.2-104 coupled to the housing 11.3.2-102, the display including one or more display screens. The display 11.3.2-104 may be coupled to the housing 11.3.2-102 such that the display 11.3.2-104 is configured to project light toward a user's eyes during use when wearing the HMD to which the display module 11.3.2-100 belongs. In at least one example, the housing 11.3.2-102 may surround the display 11.3.2-104 and provide connection features for coupling other components of the optical module described herein.

[0125] In one example, the optical module 11.3.2-100 may include one or more cameras 11.3.2-106 coupled to the housing 11.3.2-102. The cameras 11.3.2-106 may be positioned relative to the display 11.3.2-104 and the housing 11.3.2-102 such that the cameras 11.3.2-106 are configured to capture one or more images of the user's eyes during use. In at least one example, the optical module 11.3.2-100 may further include a light strip 11.3.2-108 surrounding the display 11.3.2-104. In one example, the light strip 11.3.2-108 is disposed between the display 11.3.2-104 and the cameras 11.3.2-106. The light strip 11.3.2-108 may include a plurality of lights 11.3.2-110. The plurality of lights may include one or more light-emitting diodes (LEDs) or other lights configured to project light toward a user's eyes when wearing the HMD. Each light 11.3.2-110 in the light strip 11.3.2-108 may be spaced apart around the light strip 11.3.2-108 and thus may be spaced apart evenly or unevenly around the display 11.3.2-104 at various positions on the light strip 11.3.2-108 and around the display 11.3.2-104.

[0126] In at least one example, the housing 11.3.2-102 defines a viewing opening 11.3.2-101 through which a user may view the display 11.3.2-104 when wearing the HMD device. In at least one example, the LEDs are configured and arranged to emit light through the viewing opening 11.3.2-101 onto the user's eyes. In one example, the cameras 11.3.2-106 are configured to capture one or more images of the user's eyes through the viewing opening 11.3.2-101.

[0127] As noted above, Figure 1OEach of the components and features of the optical module 11.3.2-100 shown can be replicated in another (e.g., second) optical module provided with the HMD to interact with the user's other eye (e.g., project light and capture images).

[0128] Figure 1O Any one of the features, components, and / or parts shown (including their arrangement and configuration) can be included, either individually or in any combination, in Figure 1P any other example of the devices, features, components, and parts shown or otherwise described herein. Similarly, reference Figure 1P to or any one of the features, components, and / or parts shown or otherwise described herein (including their arrangement and configuration) can be included, either individually or in any combination, in Figure 1O the examples of the devices, features, components, and parts shown.

[0129] Figure 1P A cross-sectional view illustrating an example of the optical module 11.3.2-200 is shown, including a housing 11.3.2-202, a display assembly 11.3.2-204 coupled to the housing 11.3.2-202, and a lens 11.3.2-216 coupled to the housing 11.3.2-202. In at least one example, the housing 11.3.2-202 defines a first aperture or passage 11.3.2-212 and a second aperture or passage 11.3.2-214. The passages 11.3.2-212, 11.3.2-214 can be configured to slidably engage corresponding tracks or guide rods of the HMD device to allow the optical module 11.3.2-200 to be positioned relative to the user's eyes to match the user's interpupillary distance (IPD). The housing 11.3.2-202 is capable of slidably engaging the guide rods to secure the optical module 11.3.2-200 in place within the HMD.

[0130] In at least one example, the optical module 11.3.2-200 may further include a lens 11.3.2-216 coupled to the housing 11.3.2-202 and disposed between the display component 11.3.2-204 and the user's eyes when the HMD is worn. The lens 11.3.2-216 may be configured to direct light from the display component 11.3.2-204 to the user's eyes. In at least one example, the lens 11.3.2-216 may be part of a lens assembly including a corrective lens removably attached to the optical module 11.3.2-200. In at least one example, the lens 11.3.2-216 is disposed above the light bar 11.3.2-208 and one or more eye tracking cameras 11.3.2-206 such that the cameras 11.3.2-206 are configured to capture images of the user's eyes through the lens 11.3.2-216, and the light bar 11.3.2-208 includes lights configured to project light through the lens 11.3.2-216 onto the user's eyes during use.

[0131] Figure 1P Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included, either alone or in any combination, in any other example of the devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included, either alone or in any combination, in Figure 1P the examples of the devices, features, components, and parts shown.

[0132] Figure 2FIG. 0 is a block diagram of an example of controller 110 in accordance with some embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features have not been shown for the sake of brevity and to not obscure more relevant aspects of the embodiments disclosed herein. To that end, as a non-limiting example, in some embodiments, controller 110 includes one or more processing units 202 (e.g., microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., universal serial bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, global system for mobile communications (GSM), code division multiple access (CDMA), time division multiple access (TDMA), global positioning system (GPS), infrared (IR), Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these components and various other components.

[0133] In some embodiments, one or more communication buses 204 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0134] Memory 220 includes high-speed random access memory such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes non-transitory computer-readable storage medium. In some embodiments, memory 220 or the non-transitory computer-readable storage medium of memory 220 stores the following programs, modules, and data structures, or subsets thereof, including optionally operating system 230 and XR experience module 240.

[0135] The operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate single or multiple XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for respective groups of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.

[0136] In some embodiments, the data acquisition unit 241 is configured to obtain data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 of Figure 1A and optionally from one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.

[0137] In some embodiments, the tracking unit 242 is configured to map the scene 105 and track the positioning / location of at least the display generation component 120 relative to Figure 1A the scene 105, and optionally track the location of one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the positioning / location of one or more parts of the user's hand, and / or the movement of one or more parts of the user's hand relative to Figure 1A the scene 105, relative to the display generation component 120, and / or relative to a coordinate system (which is defined relative to the user's hand). The hand tracking unit 244 is described in more detail below with respect to Figure 4 . In some embodiments, the eye tracking unit 243 is configured to track the positioning or movement of the user's gaze (or more generally, the user's eyes, face, or head) relative to the scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hand)) or relative to the XR content displayed via the display generation component 120. The eye tracking unit 243 is described in more detail below with respect to Figure 5 .

[0138] In some embodiments, the coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by the display generation component 120 and, optionally, by one or more of the output device 155 and / or the peripheral device 195. To this end, in various embodiments, the coordination unit 246 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0139] In some embodiments, the data sending unit 248 is configured to send data (e.g., presentation data, location data, etc.) to at least the display generation component 120 and, optionally, to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data sending unit 248 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0140] Although the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 are shown as residing on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 may be located in separate computing devices.

[0141] In addition, Figure 2 Rather, it is more of a functional description of the various features that may be present in a particular implementation, as opposed to the structural schematic of the embodiments described herein. As will be recognized by those of ordinary skill in the art, items shown separately may be combined, and some items may be separated. For example, Figure 2 some of the functional modules shown separately in may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the specific partitioning of functions, as well as how the features are allocated therein, will vary depending on the particular implementation and, in some embodiments, will depend in part on the specific combination of hardware, software, and / or firmware selected for the particular implementation.

[0142] Figure 3FIG. is a block diagram of an example of a display generation component 120 according to some embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and to not obscure more relevant aspects of the embodiments disclosed herein. To that end, as a non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., a microprocessor, an ASIC, an FPGA, a GPU, a CPU, a processing core, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional internal and / or external image sensors 314, a memory 320, and one or more communication buses 304 for interconnecting these components and various other components.

[0143] In some embodiments, one or more communication buses 304 include circuitry for interconnecting and controlling communication between the various system components. In some embodiments, one or more I / O devices and sensors 306 include an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, and / or one or more depth sensors (e.g., structured light, time of flight, etc.).

[0144] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more XR displays 312 correspond to diffractive, reflective, polarized, holographic, and other waveguide displays. For example, the display generation component 120 (e.g., an HMD) includes a single XR display. In another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, one or more XR displays 312 are capable of presenting MR and VR content. In some embodiments, one or more XR displays 312 are capable of presenting MR or VR content.

[0145] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to the user's hand and optionally at least a portion of the user's arm (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene that the user would see in the absence of the display generation component 120 (e.g., an HMD) (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.

[0146] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes non-transitory computer-readable storage medium. In some embodiments, memory 320 or the non-transitory computer-readable storage medium of memory 320 stores the following programs, modules, and data structures or subsets thereof, including optionally operating system 330 and XR rendering module 340.

[0147] Operating system 330 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, XR rendering module 340 is configured to present XR content to a user via one or more XR displays 312. To this end, in various embodiments, XR rendering module 340 includes a data acquisition unit 342, an XR presentation unit 344, an XR graph generation unit 346, and a data transmission unit 348.

[0148] In some embodiments, data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) at least from Figure 1A controller 110. To this end, in various embodiments, data acquisition unit 342 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0149] In some embodiments, XR presentation unit 344 is configured to present XR content via one or more XR displays 312. To this end, in various embodiments, XR presentation unit 344 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0150] In some embodiments, XR graph generation unit 346 is configured to generate an XR graph (e.g., a 3D graph of a mixed reality scene or a graph of a physical environment in which computer-generated objects can be placed to generate extended reality) based on media content data. To this end, in various embodiments, XR graph generation unit 346 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0151] In some embodiments, the data sending unit 348 is configured to send data (e.g., presentation data, location data, etc.) to at least the controller 110, and optionally to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data sending unit 348 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0152] Although the data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data sending unit 348 are shown as residing on a single device (e.g., Figure 1A the display generation component 120 of), it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data sending unit 348 may be located in separate computing devices.

[0153] In addition, Figure 3 It is more used as a functional description of the various features that may exist in a particular implementation, different from the schematic diagram of the structure of the embodiments described herein. As those of ordinary skill in the art will recognize, the items shown separately can be combined, and some items can be separated. For example, Figure 3 Some of the functional modules shown separately in can be implemented in a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules and the specific division of functions, as well as how the features are allocated therein, will vary according to the specific implementation, and in some embodiments, partly depend on the specific combination of hardware, software, and / or firmware selected for the particular implementation.

[0154] Figure 4 is a schematic diagram of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 ( Figure 1A ) is controlled by the hand tracking unit 244 ( Figure 2 ) to track the positioning / position of one or more parts of the user's hand, and / or the movement of one or more parts of the user's hand relative to the scene 105 of FIG. 1 (e.g., relative to a part of the physical environment around the user, relative to the display generation component 120, or relative to a part of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system that is defined relative to the user's hand). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0155] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures an image of the hand with sufficient resolution such that the fingers and their respective positions can be distinguished. The image sensor 404 typically captures images of other parts of the user's body, and may also or possibly capture images of all parts of the body, and may have a zooming capability or a dedicated sensor with increased magnification to capture an image of the hand with a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in combination with other image sensors to capture the physical environment of the scene 105, or serves as the image sensor for capturing the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment in such a way that the field of view of the image sensor 404 or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are treated as inputs to the controller 110.

[0156] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and additionally, possibly, color image data) to the controller 110, which extracts high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API), and the application drives the display generation component 120 accordingly. For example, a user can interact with software running on the controller 110 by moving his hand 406 and changing his hand pose.

[0157] In some embodiments, the image sensor 404 projects a speckle pattern onto the scene containing the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral displacement of the speckles in the pattern. This method is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. This method gives the depth coordinates of points in the scene at a specific distance from the image sensor 404 relative to a pre-determined reference plane. In the present disclosure, it is assumed that the image sensor 404 defines an orthogonal set of x-axis, y-axis, and z-axis such that the depth coordinates of points in the scene correspond to the z-component measured by the image sensor. Alternatively, the image sensor 404 (e.g., the hand tracking device) may use other 3D mapping methods, such as stereoscopy or time-of-flight measurement, based on a single or multiple cameras or other types of sensors.

[0158] In some embodiments, the hand tracking device 140 captures and processes a time series of depth maps containing the user's hand as the user moves his hand (e.g., the entire hand or one or more fingers). Software running on a processor in the image sensor 404 and / or the controller 110 processes the 3D map data to extract patch descriptors of the hand in these depth maps. The software can match these descriptors to patch descriptors stored in the database 408 based on a previous learning process to estimate the pose of the hand in each frame. The pose generally includes the 3D positions of the user's hand joints and finger tips.

[0159] The software can also analyze the trajectories of the hand and / or fingers over multiple frames in the sequence to identify gestures. The pose estimation function described herein can alternate with the motion tracking function such that the patch-based pose estimation is only performed once every two (or more) frames, while tracking is used to find changes in the pose that occur in the remaining frames. Pose, motion, and gesture information is provided to an application running on the controller 110 via the aforementioned API. The program can, for example, move and modify the image presented on the display generation component 120 in response to the pose and / or gesture information, or perform other functions.

[0160] In some embodiments, gestures include air gestures. An air gesture is detected without the user touching an input element that is part of a device (e.g., the computer system 101, one or more input devices 125, and / or the hand tracking device 140) (or independent of an input element that is part of a device) and is based on the detected movement of a part of the user's body (e.g., the head, one or more arms, one or more hands, one or more fingers, and / or one or more legs) through the air (including the movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one of the user's hands relative to the other of the user's hands, and / or the movement of the user's finger relative to another finger or part of the hand), and / or the absolute movement of a part of the user's body (e.g., a tap gesture that includes the hand moving in a predetermined pose by a predetermined amount and / or speed, or a shake gesture that includes a predetermined rotational speed or amount of rotation of a part of the user's body)).

[0161] In some embodiments, according to some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures performed by the movement of a user's finger relative to other fingers or parts of the user's hand for interacting with an XR environment (e.g., a virtual or mixed reality environment). In some embodiments, the air gesture is detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device)) and is based on the detected movement of a part of the user's body through the air (including the movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), the movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one hand of the user relative to the other hand of the user, and / or the movement of the user's finger relative to another finger or part of the hand), and / or the absolute movement of a part of the user's body (e.g., a tap gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shake gesture that includes a predetermined speed or amount of rotation of a part of the user's body)).

[0162] In some embodiments where the input gesture is an air gesture (e.g., in the absence of physical contact with an input device that provides information to the computer system about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touch screen, or contact with a mouse or touchpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Thus, in a particular implementation involving an air gesture, the input gesture is, for example, attention (e.g., gaze) towards a user interface element detected in combination with (e.g., concurrently with) the movement of the user's finger and / or hand performing a pinch and / or tap input, as described in more detail below.

[0163] In some embodiments, an input gesture directed to a user interface object is performed with direct or indirect reference to the user interface object. For example, user input is performed directly on the user interface object by performing the input at a location corresponding to the positioning of the user's hand relative to the positioning of the user interface object in a three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, when attention (e.g., gaze) of the user to the user interface object is detected, the input gesture is performed indirectly on the user interface object based on the positioning of the user's hand not being at the location corresponding to the positioning of the user interface object in the three-dimensional environment while the user performs the input gesture. For example, for a direct input gesture, the user can direct their input to the user interface object by initiating a gesture at or near a location corresponding to the display positioning of the user interface object (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or between 0 and 5 cm measured from the outer edge of the option or the central portion of the option). For an indirect input gesture, the user can direct their input to the user interface object by focusing on the user interface object (e.g., by gazing at the user interface object), and while focusing on the option, the user initiates an input gesture (e.g., at any location detectable by the computer system) (e.g., at a location not corresponding to the display positioning of the user interface object).

[0164] In some embodiments, according to some embodiments, input gestures (e.g., air gestures) used in the various examples and embodiments described herein include pinch inputs and tap inputs for interacting with a virtual or mixed reality environment. For example, the pinch inputs and tap inputs described below are performed as air gestures.

[0165] In some embodiments, a pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double pinch gesture. For example, a pinch gesture as an air gesture includes the movement of two or more fingers of a hand to contact each other, i.e., optionally, followed by an immediate interruption of contact with each other (e.g., within 0 to 1 second). A long pinch gesture as an air gesture includes the movement of two or more fingers of a hand contacting each other for at least a threshold amount of time (e.g., at least 1 second) before detecting an interruption of contact with each other. For example, a long pinch gesture includes a user holding a pinch gesture (e.g., where two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between the two or more fingers is detected. In some embodiments, a double pinch gesture as an air gesture includes two (e.g., or more) pinch inputs (e.g., performed by the same hand) detected consecutively and immediately with each other (e.g., within a predefined time period). For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., interrupts contact between two or more fingers), and performs a second pinch input within a predefined time period after releasing the first pinch input (e.g., within 1 second or within 2 seconds).

[0166] In some embodiments, a pinch-and-drag gesture, as an air gesture (e.g., an air drag gesture or an air swipe gesture), includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in combination with (e.g., following) a drag input that changes the positioning of the user's hand from a first positioning (e.g., the starting positioning of the drag) to a second positioning (e.g., the ending positioning of the drag). In some embodiments, the user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., opens their two or more fingers) to end the drag gesture (e.g., at the second positioning). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., the user pinches two or more fingers together and moves the same hand into a second positioning in the air using the drag gesture). In some embodiments, the pinch input is performed by the user's first hand, and the drag input is performed by the user's second hand (e.g., while the user continues the pinch input with the user's first hand, the user's second hand moves in the air from the first positioning to the second positioning). In some embodiments, an input gesture as an air gesture includes an input performed using the user's two hands (e.g., a pinch and / or tap input). For example, the input gesture includes two (e.g., or more) pinch inputs performed in combination with each other (e.g., concurrently or within a predefined time period). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch-and-drag input) is performed using the user's first hand, and in combination with performing the pinch input using the first hand, a second pinch input is performed using the other hand (e.g., the second hand of the user's two hands).

[0167] In some embodiments, a tap input performed as an air gesture (e.g., pointing to a user interface element) includes the movement of the user's finger towards the user interface element, the movement of the user's hand towards the user interface element (optionally, the user's finger extends towards the user interface element), the downward movement of the user's finger (e.g., mimicking a mouse click movement or a tap on a touch screen), or other predefined movements of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on the movement characteristics of the finger or hand performing the tap gesture movement, which is the movement of the finger or hand away from the user's viewing point and / or towards an object that is the target of the tap input, followed by the end of the movement. In some embodiments, the end of the movement is detected based on a change in the movement characteristics of the finger or hand performing the tap gesture (e.g., the end of the movement away from the user's viewing point and / or towards an object that is the target of the tap input, the reversal of the movement direction of the finger or hand, and / or the reversal of the acceleration direction of the movement of the finger or hand).

[0168] In some embodiments, the determination of the user's attention being directed to a portion of the three-dimensional environment is made based on the detection of a gaze directed to the portion of the three-dimensional environment (optionally, without the need for other conditions). In some embodiments, the determination of the user's attention being directed to a portion of the three-dimensional environment is made based on the detection of a gaze directed to the portion of the three-dimensional environment under one or more additional conditions, such as requiring the gaze to be directed to the portion of the three-dimensional environment for at least a threshold duration (e.g., dwell duration) and / or requiring the gaze to be directed to the portion of the three-dimensional environment when the user's viewpoint is within a distance threshold from the portion of the three-dimensional environment, so that the device determines that the user's attention is directed to the portion of the three-dimensional environment, wherein if one of these additional conditions is not met, the device determines that the attention is not directed to the portion of the three-dimensional environment to which the gaze is directed (e.g., until the one or more additional conditions are met).

[0169] In some embodiments, the detection of the readiness state configuration of the user or a part of the user is detected by the computer system. The detection of the readiness state configuration of the hand is used by the computer system as an indication that the user may be about to use one or more air gesture inputs (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein) performed by the hand to interact with the computer system. For example, based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape where the thumb and one or more fingers are extended and spaced apart to prepare for a pinch or grab gesture, or one or more fingers are extended and the back of the hand faces the user for a pre-tap), based on whether the hand is in a predetermined orientation relative to the user's viewpoint (e.g., below the user's head and above the user's waist and extending at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or based on whether the hand has moved in a particular manner (e.g., moving towards an area in front of the user that is above the user's waist and below the user's head or moving away from the user's body or legs) to determine the readiness state of the hand. In some embodiments, the readiness state is used to determine whether interactive elements of the user interface respond to attention (e.g., gaze) inputs.

[0170] In a scenario where input is described with reference to air gestures, it should be understood that a hardware input device attached to one or more of the user's hands or held by one or more of the user's hands can be used to detect similar gestures, where optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units can be used to track the positioning of the hardware input device in space, and the positioning and / or movement of the hardware input device is used in place of the positioning and / or movement of one or more hands in the corresponding air gesture. In a scenario where input is described with reference to air gestures, it should be understood that a hardware input device attached to one or more of the user's hands or held by one or more of the user's hands can be used to detect similar gestures. User input can be detected using controls contained in the hardware input device, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, one or more hand or finger overlays that can detect the positioning or change in positioning of parts of the hand and / or fingers relative to each other, relative to the user's body, and / or relative to the user's physical environment, and / or other hardware input device controls, where user input using the controls contained in the hardware input device is used in place of hand and / or finger gestures such as an air tap or an air pinch in the corresponding air gesture. For example, a selection input described as being performed using an air tap or an air pinch input can alternatively be detected using a button press, a tap on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input. As another example, a movement input described as being performed using an air pinch and drag (e.g., an air drag gesture or an air swipe gesture) can alternatively be detected based on an interaction with a hardware input control (such as a button press and hold, a touch on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input after the movement of the hardware input device (e.g., along with the hand associated with the hardware input device) through space). Similarly, a two-handed input that includes movement of the hands relative to each other can be performed using one air gesture and a hardware input device in a hand that is not performing the air gesture, two hardware input devices held in different hands, or two air gestures performed by different hands and / or various combinations of inputs detected by the one or more hardware input devices described above.

[0171] In some embodiments, the software can be downloaded electronically to the controller 110, for example, via a network, or alternatively can be provided on a tangible non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in the memory associated with the controller 110. Alternatively or in addition, some or all of the described functions of the computer can be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). Although the controller 110 is shown in Figure 4 , some or all of the processing functions of the controller, for example, as a unit separate from the image sensor 404, can be performed by a suitable microprocessor and software or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand tracking device) or other devices associated with the image sensor 404. In some embodiments, at least some of these processing functions can be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television receiver, a handheld device, or a head-mounted device) or integrated with any other suitable computerized device, such as a game console or a media player. The sensing function of the image sensor 404 can similarly be integrated into a computer or other computerized device controlled by the sensor output.

[0172] Figure 4 Also included is a schematic representation of a depth map 410 captured by the image sensor 404 according to some embodiments. As described above, the depth map includes a matrix of pixels with corresponding depth values. The pixels 412 corresponding to the hand 406 have been segmented from the background and the hand wrist in the figure. The brightness of each pixel within the depth map 410 is inversely proportional to its depth value (i.e., the measured z-distance from the image sensor 404), where the gray shading becomes darker as the depth increases. The controller 110 processes these depth values to identify and segment the components (i.e., a group of adjacent pixels) of the image that have the characteristics of a human hand. These characteristics can include, for example, the overall size, shape, and motion from frame to frame in a sequence of depth maps.

[0173] Figure 4 Also schematically illustrated is a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406. In Figure 4In [the figure], the hand bones 414 are superimposed on the hand background 416 that has been segmented from the original depth map. In some embodiments, key feature points on the hand and optionally on the hand wrist or arm connected to the hand (e.g., points corresponding to knuckles, finger tips, the center of the hand palm, the end of the hand connected to the hand wrist, etc.) are identified and located on the hand bones 414. In some embodiments, the controller 110 uses the positions and movements of these key feature points across multiple image frames to determine, according to some embodiments, the gesture performed by the hand or the current state of the hand.

[0174] Figure 5 An example embodiment of an eye tracking device 130 is shown ( Figure 1A ). In some embodiments, the eye tracking device 130 is controlled by an eye tracking unit 243 ( Figure 2 ) to track the positioning and movement of the user's gaze relative to the scene 105 or relative to the XR content displayed via the display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as, a head-mounted headset, helmet, goggles, or glasses) or a hand-held device placed in a wearable frame, the head-mounted device includes both components for generating XR content for the user to view and components for tracking the user's gaze relative to the XR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a hand-held device or an XR room, the eye tracking device 130 is optionally a device separate from the hand-held device or the XR room. In some embodiments, the eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used in combination with a display generation component that is also head-mounted or a display generation component that is not head-mounted. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally part of a non-head-mounted display generation component.

[0175] In some embodiments, the display generation component 120 uses a display mechanism (e.g., a left near-eye display panel and a right near-eye display panel) to display a frame including a left image and a right image in front of the user's eyes, thereby providing the user with a 3D virtual view. For example, the head-mounted display generation component may include a left optical lens and a right optical lens (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display, and virtual objects are displayed on the transparent or translucent display through which the user can directly view the physical environment. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects may be projected, for example, onto a physical surface or as a hologram, such that an individual using the system observes the virtual objects superimposed over the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.

[0176] As Figure 5 shown, in some embodiments, the eye tracking device 130 (e.g., a gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) or near-infrared (NIR) camera), and an illumination source (e.g., an IR or NIR light source, such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera may be directed at the user's eyes to receive the IR or NIR light from the light source reflected directly from the eyes, or alternatively may be directed at a "hot" mirror located between the user's eyes and the display panel that reflects the IR or NIR light from the eyes to the eye tracking camera while allowing visible light to pass through. The eye tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60 frames per second - 120 frames per second (fps)), analyzes the images to generate gaze tracking information, and transmits the gaze tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by corresponding eye tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by corresponding eye tracking cameras and illumination sources.

[0177] In some embodiments, a device-specific calibration process is used to calibrate the eye tracking device 130 to determine the parameters of the eye tracking device for a particular operating environment 100, such as the 3D geometric relationships and parameters of the LED, camera, hot mirror (if present), eye lens, and display screen. The device-specific calibration process can be performed at the factory or another facility before delivering the AR / VR equipment to the end user. The device-specific calibration process can be an automatic calibration process or a manual calibration process. According to some embodiments, the user-specific calibration process can include the estimation of the eye parameters of a particular user, such as pupil position, fovea position, optical axis, visual axis, eye spacing, etc. According to some embodiments, once the device-specific parameters and user-specific parameters are determined for the eye tracking device 130, a flash-assisted method can be used to process the images captured by the eye tracking camera to determine the user's current visual axis and fixation point relative to the display.

[0178] As Figure 5 shown, the eye tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a fixation tracking system that includes at least one eye tracking camera 540 (e.g., an infrared (IR) or near-infrared (NIR) camera) positioned on the side of the user's face where eye tracking is to be performed, and an illumination source 530 (e.g., an IR or NIR light source, such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eyes 592. The eye tracking camera 540 can be pointed at a mirror 550 (which reflects IR or NIR light from the eyes 592 while allowing visible light to pass through) (e.g., as shown at the top of Figure 5 ), or alternatively can be pointed at the user's eyes 592 to receive the reflected IR or NIR light from the eyes 592 (e.g., as shown at the bottom of Figure 5 ).

[0179] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames for the left and right display panels) and provides the frames 562 to the display 510. The controller 110 uses the fixation tracking input 542 from the eye tracking camera 540 for various purposes, such as for processing the frames 562 for display. The controller 110 optionally estimates the user's fixation point on the display 510 based on the fixation tracking input 542 obtained using a flash-assisted method or other suitable method. The fixation point estimated based on the fixation tracking input 542 is optionally used to determine the direction the user is currently looking.

[0180] The following describes several possible use cases of the current gaze direction of a user and is not intended to be limiting. As an example use case, the controller 110 may render virtual content differently based on the determined direction of the user's gaze. For example, the controller 110 may generate virtual content in the foveal region determined according to the user's current gaze direction at a higher resolution compared to the resolution in the peripheral region. As another example, the controller may position or move virtual content in the view at least in part based on the user's current gaze direction. As another example, the controller may display specific virtual content in the view at least in part based on the user's current gaze direction. As another example use case in an AR application, the controller 110 may direct an external camera for capturing the physical environment of the XR experience to focus in the determined direction. Then, the autofocus mechanism of the external camera may focus on an object or surface in the environment that the user is currently looking at on the display 510. As another example use case, the eye lens 520 may be a focusable lens, and the controller uses the gaze tracking information to adjust the focus of the eye lens 520 such that the virtual object that the user is currently looking at has an appropriate vergence to match the convergence of the user's eyes 592. The controller 110 may utilize the gaze tracking information to direct the eye lens 520 to adjust the focus such that a nearby object that the user is looking at appears at the correct distance.

[0181] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510) mounted in a wearable housing, two eye lenses (e.g., eye lenses 520), an eye tracking camera (e.g., eye tracking camera 540), and a light source (e.g., illumination source 530 (e.g., IR or NIR LED)). The light source emits light (e.g., IR or NIR light) towards the user's eyes 592. In some embodiments, the light source may be arranged in a ring or circle around each of the lenses, as Figure 5 shown. In some embodiments, for example, eight illumination sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer illumination sources 530 may be used, and other arrangements and positions of the illumination sources 530 may be used.

[0182] In some embodiments, the display 510 emits light in the visible light range and does not emit light in the IR or NIR range, and thus does not introduce noise in the gaze tracking system. Note that the position and angle of the eye tracking camera 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.

[0183] As Figure 5 The illustrated embodiments of the gaze tracking system may be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experience to the user.

[0184] Figure 6 An example of a flash-assisted gaze tracking pipeline according to some embodiments is illustrated. In some embodiments, the gaze tracking pipeline is implemented by a flash-assisted gaze tracking system (e.g., the eye tracking device 130 as illustrated in Figure 1A and Figure 5 ). The flash-assisted gaze tracking system may maintain a tracking state. Initially, the tracking state is off or "no". When in the tracking state, when analyzing the current frame to track the pupil contour and flash in the current frame, the flash-assisted gaze tracking system uses the previous information from the previous frame. When not in the tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and flash in the current frame, and if successful, initializes the tracking state to "yes" and continues with the next frame in the tracking state.

[0185] As Figure 6 shown, the gaze tracking camera may capture left and right images of the user's left and right eyes. The captured images are then input into the gaze tracking pipeline for processing starting at 610. As indicated by the arrow returning to element 600, the gaze tracking system may continue to capture images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images may be input into the pipeline for processing. However, in some embodiments or under some conditions, not all of the captured frames are processed by the pipeline.

[0186] At 610, for the currently captured image, if the tracking status is yes, the method proceeds to element 640. At 610, if the tracking status is no, then as indicated at 620, the image is analyzed to detect the user's pupil and flash in the image. At 630, if the pupil and flash are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eye.

[0187] At 640, if proceeding from element 610, the current frame is analyzed to track the pupil and flash based in part on previous information from a previous frame. At 640, if proceeding from element 630, the tracking status is initialized based on the pupil and flash detected in the current frame. The processing result at element 640 is checked to verify that the result of the tracking or detection can be trusted. For example, the result can be checked to determine whether the pupil and a sufficient number of flashes for performing gaze estimation are successfully tracked or detected in the current frame. At 650, if the result cannot be trusted, then at element 660, the tracking status is set to no, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is trusted, the method proceeds to element 670. At 670, the tracking status is set to yes (if it is not already yes), and the pupil and flash information is passed to element 680 to estimate the user's point of gaze.

[0188] Figure 6 It is intended to be used as an example of an eye tracking technique that can be used for a particular specific implementation. As will be recognized by those of ordinary skill in the art, according to various embodiments, in the computer system 101 for providing an XR experience to a user, other eye tracking techniques that currently exist or are developed in the future can be used to replace the flash-assisted eye tracking technique described herein or used in combination with the flash-assisted eye tracking technique.

[0189] In some embodiments, a captured portion of the real-world environment 602 is used to provide an XR experience to the user, such as a mixed reality environment in which one or more virtual objects are superimposed over a representation of the real-world environment 602.

[0190] Accordingly, the description herein describes some implementations of a three-dimensional environment (e.g., an XR environment) that includes a representation of real-world objects and a representation of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table that exists in a physical environment, which is captured and displayed in the three-dimensional environment (e.g., actively displayed via a camera and a display of a computer system or passively displayed via a transparent or semi-transparent display of the computer system). As previously described, the three-dimensional environment is optionally a mixed reality system, where the three-dimensional environment is based on a physical environment captured by one or more sensors of a computer system and displayed via a display generation component. As a mixed reality system, the computer system is optionally capable of selectively displaying portions and / or objects of the physical environment such that the corresponding portions and / or objects of the physical environment appear as if they exist in the three-dimensional environment displayed by the computer system. Similarly, the computer system is optionally capable of displaying virtual objects in the three-dimensional environment at corresponding locations that have corresponding positions in the real world such that the virtual objects appear as if they exist in the real world (e.g., the physical environment). For example, the computer system optionally displays a vase such that the vase appears as if a real vase is placed on top of a table in the physical environment. In some implementations, the corresponding locations in the three-dimensional environment have corresponding positions in the physical environment. Thus, when the computer system is described as displaying a virtual object at a corresponding location relative to a physical object (e.g., a location at or near the user's hand or a location at or near a physical table), the computer system displays the virtual object at a specific location in the three-dimensional environment such that it appears as if the virtual object is at or near the physical object in the physical environment (e.g., the virtual object is displayed at a location in the three-dimensional environment that corresponds to the location in the physical environment where the virtual object would be displayed if the virtual object were a real object at that specific location).

[0191] In some implementations, real-world objects that exist in a physical environment and are displayed in the three-dimensional environment (e.g., and / or visible via a display generation component) can interact with virtual objects that exist only in the three-dimensional environment. For example, the three-dimensional environment can include a table and a vase placed on top of the table, where the table is a view (or representation) of a physical table in the physical environment and the vase is a virtual object.

[0192] In a three-dimensional environment (e.g., a real environment, a virtual environment, or an environment that includes a mixture of real and virtual objects), an object is sometimes said to have depth or simulated depth, or an object is said to be visible, displayed, or placed at different depths. In this context, depth refers to a dimension that is different from height or width. In some embodiments, depth is defined relative to a fixed set of coordinates (e.g., where a room or object has a height, depth, and width defined relative to a fixed set of coordinates). In some embodiments, depth is defined relative to the position or viewpoint of a user, in which case the depth dimension varies based on the position of the user and / or the position and angle of the user's viewpoint. In some embodiments where depth is defined relative to the position of the user relative to the surface of the environment (e.g., the surface of the floor or ground of the environment), an object that is further away from the user along a line extending parallel to the surface is considered to have a greater depth in the environment, and / or the depth of the object is measured along an axis that extends outward from the user's position and is parallel to the surface of the environment (e.g., depth is defined in a cylindrical or substantially cylindrical coordinate system, where the user's position is at the center of a cylinder that extends from the user's head towards the user's feet). In some embodiments where depth is defined relative to the user's viewpoint (e.g., the direction relative to a point in space that determines which part of the environment is visible via a head-mounted device or other display), an object that is further away from the user's viewpoint along a line extending parallel to the user's viewpoint is considered to have a greater depth in the environment, and / or the depth of the object is measured along an axis that extends outward from the user's viewpoint and is parallel to the direction of the user's viewpoint (e.g., depth is defined in a spherical or substantially spherical coordinate system, where the origin of the viewpoint is at the center of a sphere that extends outward from the user's head). In some embodiments, depth is defined relative to a user interface container (e.g., a window or application in which an application and / or system content is displayed), where the user interface container has a height and / or width, and depth is a dimension that is orthogonal to the height and / or width of the user interface container. In some embodiments, in the case where depth is defined relative to a user interface container, when the container is placed in a three-dimensional environment or is initially displayed (e.g., such that the depth dimension of the container extends outward away from the user or the user's viewpoint), the height and / or width of the container is typically orthogonal or substantially orthogonal to a straight line that extends from the user's position (e.g., the user's viewpoint or the user's position) to the user interface container (e.g., the center of the user interface container or another feature point of the user interface container). In some embodiments, in the case where depth is defined relative to a user interface container, the depth of an object relative to the user interface container refers to the positioning of the object along the depth dimension of the user interface container. In some embodiments, multiple different containers can have different depth dimensions (e.g., different depth dimensions that extend away from the user or the user's viewpoint in different directions and / or from different starting points).In some embodiments, when defining depth relative to a user interface container, the direction of the depth dimension remains constant for the user interface container as the position of the user interface container, the user, and / or the user's viewing point changes (e.g., or when multiple different viewers are viewing the same container in a three-dimensional environment, such as during an in-person collaboration session and / or when multiple participants are in a real-time communication session with shared virtual content that includes the container). In some embodiments, for a curved container (e.g., including a container having a curved surface or a curved content area), the depth dimension optionally extends into the surface of the curved container. In some cases, the z-spacing (e.g., the spacing between two objects in the depth dimension), the z-height (e.g., the distance of one object from another in the depth dimension), the z-position (e.g., the position of one object in the depth dimension), the z-depth (e.g., the position of one object in the depth dimension), or the simulated z-dimension (e.g., the depth used as a dimension of an object, the dimension of an environment, the orientation in space, and / or the simulated orientation in space) are used to refer to the concept of depth as described above.

[0193] In some embodiments, the user can optionally interact with virtual objects in a three-dimensional environment using one or both hands as if the virtual objects were real objects in a physical environment. For example, as described above, one or more sensors of the computer system optionally capture one or more hands of the user and display a representation of the user's hands in the three-dimensional environment (e.g., in a manner similar to that described above for displaying real-world objects in the three-dimensional environment), or in some embodiments, the user's hands can be seen via the display generation component, due to the transparency / translucency of the portion of the user interface being displayed by the display generation component, or due to the projection of the user interface onto a transparent / translucent surface or onto the user's eyes or into the user's field of view. Thus, in some embodiments, the user's hands are displayed at their corresponding positions in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment that can interact with virtual objects in the three-dimensional environment as if those virtual objects were physical objects in a physical environment. In some embodiments, the computer system can update the display of the representation of the user's hands in the three-dimensional environment in conjunction with the movement of the user's hands in the physical environment.

[0194] In some of the embodiments described below, the computer system is optionally capable of determining an "effective" distance between a physical object in the physical world and a virtual object in a three-dimensional environment, e.g., for determining whether the physical object is directly interacting with the virtual object (e.g., whether a hand is touching, grasping, holding, etc., the virtual object or is within a threshold distance of the virtual object). For example, a hand directly interacting with a virtual object optionally includes one or more of the following: a finger of the hand pressing a virtual button, a hand of the user grasping a virtual vase, the hands of the user brought together and pinching / holding the user interface of an application, and two fingers performing any other type of interaction described herein. For example, when determining whether a user is interacting with a virtual object and / or how the user is interacting with the virtual object, the computer system optionally determines the distance between the user's hand and the virtual object. In some embodiments, the computer system determines the distance between the user's hand and the virtual object by determining the distance between the position of the hand in the three-dimensional environment and the position of the virtual object of interest in the three-dimensional environment. For example, the one or more hands of the user are located at a particular location in the physical world, and the computer system optionally captures the one or more hands and displays the one or more hands at a particular corresponding location in the three-dimensional environment (e.g., the location where the hand would be displayed in the three-dimensional environment if the hand were a virtual hand rather than a physical hand). Optionally, the location of the hand in the three-dimensional environment is compared with the location of the virtual object of interest in the three-dimensional environment to determine the distance between the one or more hands of the user and the virtual object. In some embodiments, the computer system optionally determines the distance between the physical object and the virtual object by comparing locations in the physical world (e.g., rather than comparing locations in the three-dimensional environment). For example, when determining the distance between one or more hands of the user and a virtual object, the computer system optionally determines the corresponding position of the virtual object in the physical world (e.g., the location where the virtual object would be located in the physical world if the virtual object were a physical object rather than a virtual object), and then determines the distance between the corresponding physical location and the one or more hands of the user. In some embodiments, the same techniques are optionally used to determine the distance between any physical object and any virtual object. Thus, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the computer system optionally performs any of the techniques described above to map the position of the physical object to the three-dimensional environment and / or map the position of the virtual object to the physical environment.

[0195] In some embodiments, the same or similar techniques are used to determine where and what the user's gaze is directed at, and / or where and what a physical stylus held by the user is directed at. For example, if the user's gaze is directed at a specific location in the physical environment, the computer system optionally determines the corresponding location in the three-dimensional environment (e.g., the virtual location of the gaze), and if a virtual object is located at the corresponding virtual location, the computer system optionally determines that the user's gaze is directed at the virtual object. Similarly, the computer system is optionally able to determine the direction in the physical environment that the stylus is directed based on the orientation of the physical stylus. In some embodiments, based on this determination, the computer system determines the corresponding virtual location in the three-dimensional environment that corresponds to the location in the physical environment where the stylus is directed, and optionally determines that the stylus is directed at the corresponding virtual location in the three-dimensional environment.

[0196] Similarly, the embodiments described herein may refer to the location of the user (e.g., the user of the computer system) in the three-dimensional environment and / or the location of the computer system in the three-dimensional environment. In some embodiments, the user of the computer system is holding, wearing, or otherwise located at or near the computer system. Thus, in some embodiments, the location of the computer system serves as a proxy for the location of the user. In some embodiments, the location of the computer system and / or the user in the physical environment corresponds to the corresponding location in the three-dimensional environment. For example, the location of the computer system will be the location in the physical environment (and its corresponding location in the three-dimensional environment) where, if the user were standing at that location and facing the corresponding portion of the physical environment visible via the display generation component, the user would see in the physical environment those objects that are in the same location, orientation, and / or size (e.g., in an absolute sense and / or relative to each other) as the objects that are displayed or visible in the three-dimensional environment by the display generation component of the computer system. Similarly, if the virtual objects displayed in the three-dimensional environment are physical objects in the physical environment (e.g., physical objects placed in the physical environment at the same location as the location of these virtual objects in the three-dimensional environment, and physical objects having the same size and orientation in the physical environment as when in the three-dimensional environment), the location of the computer system and / or the user is the location from which the user would see in the physical environment those virtual objects that are in the same location, orientation, and / or size (e.g., in an absolute sense and / or relative to each other and real-world objects) as the virtual objects displayed in the three-dimensional environment by the display generation component of the computer system.

[0197] In this disclosure, various input methods are described relative to interaction with a computer system. When one input device or input method is used to provide an example and another input device or input method is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the input device or input method described relative to the other example. Similarly, various output methods are described relative to interaction with a computer system. When one output device or output method is used to provide an example and another output device or output method is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the output device or output method described relative to the other example. Similarly, various methods are described relative to interaction with a virtual environment or a mixed reality environment via a computer system. When interaction with a virtual environment is used to provide an example and a mixed reality environment is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the methods described relative to the other example. Accordingly, this disclosure discloses embodiments that are combinations of features of multiple examples without exhaustively listing all features of the embodiments in the description of each example embodiment.

[0198] User Interface and Associated Processes

[0199] Attention is now directed to embodiments of a user interface (“UI”) and associated processes that can be implemented on a computer system, such as a portable multifunctional device or a head-mounted device, having a display generation component, one or more input devices, and optionally one or more cameras.

[0200] Figures 7A to 7H An example of a computer system that facilitates control of the degree of immersion in a virtual environment according to some embodiments is illustrated.

[0201] Figure 7A An illustration shows computer system 101 in a real-world environment 702 and displaying a three-dimensional environment 704 on a user interface via a display generation component (e.g., display generation component 120 of FIG. 1). As described above with reference to FIGS. 1 to Figure 6 As described, computer system 101 optionally includes a display generation component (e.g., a touch screen) and a plurality of image sensors (e.g., Figure 3image sensor 314). The image sensor optionally includes one or more of the following: a visible light camera; an infrared camera; a depth sensor; or any other sensor that the computer system 101 will be able to use to capture one or more images of the user or a part of the user when the user interacts with the computer system 101. In some embodiments, the user interface described below is implemented on a head-mounted display that includes: a display generation component that displays the user interface to the user, and sensors that detect the movement of the physical environment and / or the user's hand (such as movement that is interpreted by the computer system as a gesture such as an air gesture) (e.g., an external sensor facing away from the user), and / or sensors that detect the user's gaze (e.g., an internal sensor facing inward toward the user's face). The figures herein illustrate a three-dimensional environment presented to the user by the computer system 101 (e.g., and displayed by the display generation component of the computer system 101) and a top view 718 of the physical environment and / or the three-dimensional environment 704 associated with the computer system 101, which is used to illustrate the relative position of objects in the real-world environment and the position of virtual objects in the three-dimensional environment.

[0202] As Figure 7A shown, the computer system 101 captures one or more images of the real-world environment 702 (e.g., the operating environment 100) around the computer system 101, including one or more objects in the real-world environment 702 around the computer system 101. In some embodiments, the computer system 101 displays a representation of the real-world environment 702 in the three-dimensional environment 702, or a portion of the real-world environment 704 is visible in the three-dimensional environment via the display generation component 120. For example, the three-dimensional environment 704 includes a room that includes a representation of a corner table 708a (corner table 708b in the top view 718), a representation of a desk 710a (e.g., the real object desk 710b in the top view 718), a representation of a coffee table 714a (e.g., the real object coffee table 714b in the top view 718), and a representation of an end table 712a (e.g., the real object end table 712b in the top view 718), each of these representations being optionally a photo-realistic representation, a simplified representation, a cartoon, a comic, and / or a digital or passive see-through representation, as described in reference method 800.

[0203] As shown in the top view 718, the user 720 of the computer system 101 is sitting on a couch 719 and holds the computer system 101 (e.g., or wears the computer system 101 if the computer system 101 is a head-mounted device, for example) in such a way that one or more sensors face the other end of the room, thus capturing the corner table 708b, the desk 710b, the end table 712b, and the coffee table 714b, and displaying a representation of the objects in the three-dimensional environment 704.

[0204] In Figure 7A , computer system 101 is displaying an immersion level indicator 716. The immersion level indicator 716 indicates that computer system 101 is at the current immersion level at which it is displaying the three-dimensional environment 704 (e.g., outside the maximum immersion level). In some embodiments, the immersion level corresponds to the amount by which the view of the physical environment (e.g., the view of objects in the real-world environment 702) is occluded by the virtual environment (e.g., an analog environment optionally different from the real-world environment 702 surrounding the user), or to the amount by which the objects of the physical environment are modified to achieve a particular spatial effect (e.g., as will be described in further detail below with respect to method 800). For example, the maximum immersion level (e.g., full immersion) optionally refers to a state in which no physical environment is visible in the three-dimensional environment 704 via the display generation component 120 and the entire three-dimensional environment 704 is surrounded by the virtual environment. In some embodiments, an intermediate immersion level (e.g., an immersion level less than the maximum immersion level and greater than the no-immersion level) refers to a state in which a portion of the real-world environment 702 is visible in the three-dimensional environment 704 via the display generation component 120 and the portion of the real-world environment 702 that would otherwise be visible (e.g., if not for the immersion) has been replaced by the virtual environment. In some embodiments, the immersion level indicator 716 optionally includes a plurality of elements associated with a plurality of immersion levels. In some embodiments, as the immersion level increases, computer system 101 presents more elements of the virtual environment.

[0205] In Figure 7A , the immersion level indicator 716 indicates that the current immersion level is a first level that corresponds to the amount of the virtual environments 722-1a (e.g., 722-1b in top view 718), 722-2a (e.g., 722-2b in top view 718) being displayed by computer system 101 as illustrated by Figure 7A . The virtual environments 722-1a, 722-2a include a display corresponding to background 1 (BKGD 1) that optionally includes features simulating a location or atmosphere. For example, the display corresponding to BKGD 1 optionally includes a sunny day at a hill. Further details of the virtual environment and immersion level are described with reference to method 800.

[0206] In an illustrative embodiment, computer system 101 displays a three-dimensional environment 704 that includes a control center user interface 724a (e.g., a system user interface, and / or a first user interface of the control center user interface) and a video application user interface 726a (e.g., a user interface of an application). As shown in the top view 718, the control center user interface 724b and the video application user interface 726b are located at different positions in the three-dimensional environment 704. The control center user interface includes an immersion slider user interface element 728a, a system environment setting user interface element 728b, an auto-dim user interface element 728c, a volume control user interface element 728d, and a focus mode control user interface element 728e. The immersion slider user interface element 728a is displayed at a first fill level corresponding to the current immersion level (e.g., the current immersion level of the immersion indicator 716). Similarly, the volume control user interface element 728d includes the display of a slider element at a position corresponding to the current volume level of the computer system 101. The auto-dim user interface element 728c is active. Further details regarding the control center user interface 724a are described with reference to methods 800, 1000, and / or 1200.

[0207] In Figure 7A the illustrative embodiment, the user attention 730a - 730c (e.g., the gaze of user 720) and the input from the hand 732 of user 720 alternatively point to the immersion slider user interface element 728a, the system environment setting user interface element 728b, and the auto-dim user interface element 728c. In some embodiments, the user interface elements may be selected via user attention or input from the hand 732 (e.g., an air gesture) or via a combination of both user attention 730 and input from the hand 732, and such characteristics of the input and the processes for detecting such input are described in more detail with reference to method 800.

[0208] Figure 7A1 Illustrates concepts similar and / or identical to Figure 7A those shown (with many of the same reference numerals). It should be understood that, unless otherwise indicated below, Figures 7A to 7H the elements with the same reference numerals as Figure 7A1 those shown have one or more or all of the same characteristics. Figure 7A1 Includes computer system 101, which includes a display generation component 120 (or the same as the display generation component). In some embodiments, computer system 101 and display generation component 120 respectively have Figures 7A to 7H the computer system 101 shown as well as in FIG. 1 and Figure 3One or more of the characteristics of the display generation component 120 shown, and in some embodiments, Figures 7A to 7H the computer system 101 and the display generation component 120 shown have Figure 7A1 one or more of the characteristics of the computer system 101 and the display generation component 120 shown.

[0209] In Figure 7A1 it, the display generation component 120 includes one or more internal image sensors 314a oriented towards the user's face (e.g., referring to Figure 5 the eye tracking camera 540 shown). In some embodiments, the internal image sensor 314a is used for eye tracking (e.g., detecting the user's gaze). The internal image sensor 314a is optionally arranged on the left and right portions of the display generation component 120 so as to enable eye tracking of the user's left and right eyes. The display generation component 120 further includes external image sensors 314b and 314c facing outwards from the user to detect and / or capture the physical environment and / or the movement of the user's hand. In some embodiments, the image sensors 314a, 314b, and 314c have one or more of the characteristics of the image sensor 314 referred to Figures 7A to 7H to.

[0210] In Figure 7A1 it, the display generation component 120 is illustrated as displaying content optionally corresponding to the content described as being displayed and / or visible via the display generation component 120. In some embodiments, the content is displayed by a single display included in the display generation component 120 (e.g., Figures 7A to 7H the display 510). In some embodiments, the display generation component 120 includes two or more displays (e.g., a left display panel and a right display panel respectively for the user's left and right eyes, as referred to Figure 5 to), and the two or more displays have displayed outputs that are combined (e.g., by the user's brain) to create Figure 5 the view of the content shown. Figure 7A1 shown.

[0211] The display generation component 120 has a field of view corresponding to Figure 7A1 the content shown (e.g., the field of view captured by the external image sensors 314b and 314c and / or visible to the user via the display generation component 120, which is indicated by a dashed line in the top view). Since the display generation component 120 is optionally a head-mounted device, the field of view of the display generation component 120 is optionally the same as or similar to the user's field of view.

[0212] In Figure 7A1, the user is depicted as performing an air pinch gesture (e.g., using hand 732) to provide input to computer system 101 to provide user input directed to content displayed by computer system 101. This depiction is intended to be exemplary and not limiting; the user optionally uses different air gestures and / or uses the same techniques as described in reference to FIG. Figures 7A to 7H Other forms of input are used to provide user input.

[0213] In some embodiments, the computer system 101 is configured to Figures 7A to 7H The user input is responded to.

[0214] exist Figure 7A1 In the example of , because the user's hand is within the field of view of the display generation component 120, the user's hand is visible in the three-dimensional environment. That is, the user can optionally see any part of his or her own body within the field of view of the display generation component 120 in the three-dimensional environment. It should be understood that Figures 7A to 7H One or more or all aspects of the present disclosure shown or described with reference to these figures and / or described with reference to corresponding methods are optionally provided with Figure 7A1 Similar or analogous methods are implemented on the computer system 101 and the display generation unit 120 .

[0215] Figure 7B 704 is illustrated with a three-dimensional environment responsive to an input pointing to an immersion slider user interface element 728a corresponding to a request to increase the immersion level according to some embodiments of the present disclosure. Figure 7A The illustrated first immersion level is greater than the second immersion level. Figures 7A to 7B , computer system 101 detects user attention 730a directed to element 728a while hand 732 performs an air pinch gesture, and then hand 732 moves upward while in a pinch hand shape, as described in more detail with reference to method 800. In response, Figure 7B As shown, the immersion slider user interface element 728a includes Figure 7A The illustrated representation of the current immersion level is displayed compared to a representation of a higher current immersion level because Figure 7B The current immersion level in the three-dimensional environment 704 has been increased in response to the input pointing to the immersion slider user interface element 728a corresponding to the request to increase the immersion level. In addition, the increase in the immersion level is illustrated in the immersion level indicator 716. The amount of the three-dimensional environment 704 that has been replaced by the virtual environment 722a is increased (compared to Fig. 7A The method 800 is used to increase the size of the virtual environment 722a in the three-dimensional environment 704 (e.g., increase the size of the visual "portal" leading to the virtual environment 722a). Further details on improving immersion are described with reference to method 800. Figure 7B In the illustrated embodiment of , in response to the increase in the immersion level, the control center user interface 724a and the video application user interface 726a remain in their respective positions.

[0216] Figure 7C A three-dimensional environment 704 is illustrated that includes a second user interface 724a for controlling a user interface 724a in response to a pointing device according to some embodiments of the present disclosure. Fig. 7A In the illustrated embodiment, in response to the input of the system environment setting user interface element 728b displayed in the display, the display mode of the virtual environment displayed by the display generation component is changed. Fig. 7A In response to an input from the system environment setting user interface element 728b in the control center user interface (e.g., input including a user's attention for a threshold duration, such as described in the present disclosure with reference to method 800), the second user interface of the control center user interface replaces and / or overlies the Fig. 7A The illustrated control center user interface 724a displays a first user interface, such as Figure 7C The second user interface of the control center user interface 724a includes selectable options for changing the display mode (eg, lighting settings) of the virtual environment 722 (eg, virtual environments 722-1a, 722-2a). Specifically, the illustrated second user interface includes a selectable display mode 1 option 736a for displaying the virtual environment 722 (e.g., simulated light (e.g., including simulated light sources at a first brightness level and / or from a simulated sun), daytime or daylight lighting settings), a selectable display mode 2 option 736b for displaying the virtual environment 722 (e.g., simulated darkness (e.g., including simulated light sources at a second brightness level lower than the first brightness level and / or from a simulated moon and stars), simulated nighttime or simulated nightlight lighting settings), a selectable display mode 3 option 736c (AUTO) for displaying the virtual environment 722 (e.g., a lighting setting that causes the computer system 101 to transition between different simulated lighting settings based on satisfying one or more criteria (such as the current time of day at the computer system being a particular time of day)), and a change background option 736d for displaying the virtual environment 722. Further details regarding lighting settings are described with reference to method 800.

[0217] Additionally, in Figure 7CIn the illustrated embodiments, the user attentions 730d and 730e (e.g., the gaze of user 720) and the input from the hand 732 of user 720 alternatively point to the selectable display mode 2 option 736b and the change background option 736d. In some embodiments, user interface elements may be selected via user attention or input from the hand 732 or via a combination of both user attention (e.g., gaze) and input from the hand 732, and such characteristics of the input and the processes for detecting such input are described in more detail with reference to method 800. In the illustrated embodiment, the selectable display mode 1 option 736a is currently selected, as illustrated illustratively in the shading below the selectable display mode 1 option 736a, which causes Figure 7C the virtual environment in

[0218] Fig.7D to be displayed with a visual appearance corresponding to the display mode 1 option 736a (e.g., light settings). Figure 7C Figure 6 illustrates a three-dimensional environment 704 including a second user interface that controls the user interface 724a according to some embodiments of the present disclosure, the second user interface for changing the display mode of the virtual environment to be displayed via a display generation component from display mode 1 to display mode 2 in response to an input (e.g., user attention 730d (e.g., gaze or another type of user attention) and / or input from Fig.7D Figure 7 illustrates the state of the three-dimensional environment 704 displayed in response to an input pointing to the selectable display mode 2 option 736b of Figure 7C Figure 7. Figure 7C In

[0219] Figure 8, the selectable display mode 2 option 736b is selected. In response, the computer system 101 optionally modifies the virtual environment 722 according to the selection. Thus, in Fig.7D Figure 9, the virtual environments 722-1a, 722-1b have been changed from BKGD 1 to BKGD 2 while maintaining the same level of immersion. The transition includes a visual transition of the lighting settings in which the virtual environment 722 is displayed from display mode 1 such as daytime lighting settings to display mode 2 such as nighttime lighting settings (e.g., dimming one or more simulated light sources, reducing brightness, turning off, or changing the light source for the virtual environment from a daytime light source (e.g., a simulated sun) to a nighttime light source (e.g., a simulated moon and stars)). Further details regarding the types of lighting settings and the transitions between them are discussed with reference to method 800. Fig.7D In

[0220] Fig. 7EIllustrated is a three-dimensional environment 704 including a third user interface controlling a user interface 724a, the third user interface being configured to change display parameters of a virtual environment displayed via a display generation component in response to an input that points to Fig.7D a selectable background option 736d that changes. In the illustrated embodiment, the current immersion level is non-immersion, but it should be noted that in some embodiments, the current immersion level is higher than non-immersion, and thus the three-dimensional environment 704 includes a virtual environment, such as Fig.7D the virtual environment 722 (e.g., 722-1a, 722-2a).

[0221] In Fig. 7E , the third user interface of the control user interface 724a for changing display parameters of the virtual environment includes a selectable option for changing a display parameter A related to the three-dimensional environment and a selectable option for changing a display parameter B related to the three-dimensional environment. Specifically, the third user interface includes selectable options 740a, 740b, 740c, 740d for changing the display parameter A, and includes selectable options 742a, 742b, 742c for changing the display parameter B. When selected, the display parameter A optionally corresponds to a simulation of a physical location, and when an option corresponding to the display parameter A is selected, the computer system 101 optionally displays a three-dimensional environment including the simulated physical location corresponding to the selected option. For example, when the selectable option 740a is selected, a lake or water body scene is optionally simulated; when the selectable option 740b is selected, a street scene is optionally simulated; when the selectable option 740c is selected, a ship or dock scene is optionally simulated; when the selectable option 740d is selected, a hill or mountain scene is optionally simulated. The display parameter B optionally corresponds to a simulated atmosphere effect displayed by the display generation component, and when an option corresponding to the display parameter B is selected, the computer system 101 optionally displays a three-dimensional environment including the simulated atmosphere effect corresponding to the selected option. For example, when the selectable option 742a is selected, a dew atmosphere is optionally simulated; when the selectable option 742b is selected, a clear atmosphere is optionally simulated; when the selectable option 742c is selected, a cloudy atmosphere is optionally simulated. The features corresponding to the display parameter A and the display parameter B are described in detail with reference to method 800.

[0222] In Fig. 7EIn the illustrated embodiments, the user attention 730f, 730g (e.g., the gaze of user 720) and the input from the hand 732 of user 720 alternatively point to the user interface element 740a for display parameter A and the user interface element 742b for display parameter B. In some embodiments, the user interface element may be selected via user attention or the input from the hand 732 or a combination of both user attention and the input from the hand 732, and such characteristics of the input and the process for detecting such input are described in more detail with reference to method 800.

[0223] Figure 7F Illustrates a three-dimensional environment 704 for input to a third user interface that controls the user interface 724a for changing display parameters of a virtual environment displayed via a display generation component in response to pointing Fig. 7E The three-dimensional environment includes a virtual environment 722 (e.g., 722a) that simulates BKGD 3 at a current immersion level indicated by the immersion level indicator 716. For example, in response to user attention 730f pointing to the user interface element 740a for display parameter A in Fig. 7E , the computer system 101 optionally causes BKGD 3 (e.g., a preview of the virtual environment 722a) to be displayed according to the selectable option 740a. Even when there is no immersion in Fig. 7E (e.g., the virtual environment is not displayed), the computer system automatically and temporarily increases the immersion to display BKGD 3 according to the selectable option 740a and then optionally reverts to the Fig. 7E immersion level. In another example, in response to user attention 730g pointing to the user interface element 742a for display parameter B in Fig. 7E , the computer system 101 optionally causes BKGD 3 (e.g., a preview of the virtual environment 722a) to be displayed according to the selectable option 742a. Additionally, even when there is no immersion in Fig. 7E , the computer system 101 automatically and temporarily increases the immersion to display BKGD 3 according to the selectable option 742a and then optionally reverts to the Fig. 7E immersion level. Further details regarding the virtual preview are described with reference to method 800.

[0224] Figure 7G Illustrates a three-dimensional environment 704 that includes a third user interface controlling the user interface 724a for changing display parameters of a virtual environment displayed via a display generation component Figure 7F after a preview of the virtual environment 722a in BKGD 3 being displayed (e.g., after a threshold time period such as 2s, 5s, 10s, 50s, 100s, or another threshold time period) has elapsed since the preview was displayed. Figure 7Gthe three-dimensional environment 704 in returns to when receiving Fig. 7E the input, the state of the three-dimensional environment. In the illustrated embodiment, when receiving Fig. 7E the input, the three-dimensional environment 704 does not include a virtual environment and / or the current immersion level is zero. Thus, in Figure 7G the illustrated embodiment in, the three-dimensional environment 704 returns to the state where the virtual environment is not displayed and / or the current immersion level is zero.

[0225] Figure 7H Illustrated is the three-dimensional environment 704 in response to an input (e.g., user attention 730b and / or input from Fig. 7A the hand 732) that automatically dims the user interface element 728c and the content item 750 of the video application user interface 726a is playing. In Fig. 7A the illustrated embodiment in, both the virtual environments 722-1a, 722-2a and parts of the physical environment are reduced in visual salience (e.g., reduced in brightness and / or dimmed), while the video application user interface 726a (e.g., the user interface of the application) is not reduced in visual salience (e.g., reduced in brightness and / or dimmed). In some embodiments, parts of the video application user interface 726a are dimmed while the content item 750 of the video application user interface 726a is not dimmed. Additionally, in Figure 7H it, the virtual environments 722-1a, 722-2a have changed from BKGD 1 (of Figure 7H the) to BKGD 2 while maintaining the same immersion level. This transition optionally includes the virtual environments 722-1a, 722-2a changing from Fig. 7A the display mode 1 option 736a (e.g., simulated light (e.g., including a simulated light source at a first brightness level), simulated daytime lighting, or simulated sunlight lighting settings) of Figure 7C to Figure 7C the display mode 2 option 736b (e.g., simulated darkness (e.g., including a simulated light source at a second brightness level lower than the first brightness level), simulated night, or simulated night light lighting settings) of

[0226] It should be noted that in some embodiments, the user interfaces discussed above, such as the control center user interface 724a and the video application user interface 726a, are displayed in the three-dimensional environment 704 in an orientation facing the viewpoint of the user 720 (e.g., the normal of the control center user interface 724a and / or the video application user interface 726a intersects the viewpoint of the user 720 and / or faces the viewpoint of the user).

[0227] Among other aspects of the disclosed embodiments, further details regarding aspects of the FIG. 11A to FIG. 11F exemplary embodiments are discussed with reference to method 800.

[0228] FIG. 8A to FIG. 8I FIG. 8 is a flowchart that illustrates an exemplary method for facilitating control of immersion in a virtual environment according to some embodiments. In some embodiments, method 800 is performed at a computer system (e.g., computer system 101 in FIG. 1, such as a tablet computer, a smart phone, a wearable computer, or a head-mounted device), the computer system including a display generation component (e.g., the display generation component 120 in FIGS. 1, Figure 3 and Figure 4 and Figure 1A )(e.g., a head-up display, a monitor, a touch screen, and / or a projector) and one or more cameras (e.g., a camera pointing downward at a user's hand (e.g., a color sensor, an infrared sensor, or other depth-sensing camera) or a camera pointing forward from a user's head). In some embodiments, method 800 is governed by instructions stored in a non-transitory computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., Figure 1A the control unit 110 in

[0229] In some embodiments, method 800 is performed at a computer system that communicates with a display generation component and one or more input devices. For example, a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer or other electronic device. In some embodiments, the display generation component is a display integrated with the electronic device (optionally a touchscreen display), an external display such as a monitor, projector, television, or a hardware component (optionally integrated or external) for projecting a user interface or making the user interface visible to one or more users. In some embodiments, the one or more input devices include electronic devices or components capable of receiving user input (e.g., capturing or detecting user input) and sending information associated with the user input to the computer system. Examples of input devices include touchscreens, mice (e.g., external), trackpads (optionally integrated or external), touchpads (optionally integrated or external), remote control devices (e.g., external), another mobile device (e.g., separate from the computer system), handheld devices (e.g., external), controllers (e.g., external), cameras, depth sensors, eye tracking devices, and / or motion sensors (e.g., hand tracking devices, hand motion sensors). In some embodiments, the computer system communicates with a hand tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touchscreens, touchpads)). In some embodiments, the hand tracking device is a wearable device, such as a smart hand glove. In some embodiments, the hand tracking device is a handheld input device, such as a remote control or a stylus.

[0230] In some embodiments, when displaying virtual content (e.g., a virtual environment or virtual elements that enhance a physical environment (e.g., an AR setting), such as a virtual setting and / or a virtual environment) at a first immersion level via the display generation component (e.g., corresponding to a level that immerses a user of the computer system into a virtual environment or other virtual content, as described below), the computer system displays (802a) the system user interface of the computer system via the display generation component, wherein displaying the system user interface includes: displaying an immersion control element (e.g., a slider, dial, toggle device, segmented control, or another type of control element) configured to control the computer system to the immersion level at which it displays the virtual content, such as Fig. 7A and Fig.7A1 the three-dimensional environment 704 in Fig. 7A and Fig.7A1The control center user interface 724a in it. In some embodiments, the computer system is displaying virtual content in a three-dimensional environment. In some embodiments, the three-dimensional environment is an extended reality (XR) environment, such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment. The virtual content is optionally any type of content that is not in the physical environment of the user and / or the computer system. For example, the virtual content is optionally a virtual representation of a location corresponding to a geographical location and / or an atmosphere corresponding to that location at a specific time (e.g., a hill with the "HOLLYWOOD" sign in Hollywood, California during a sunny day, or the shore of Lake Houston in Houston, Texas at night time, where the night time corresponds to the night at that shore), or a user interface of an application on the computer system (e.g., a messaging application, a content playback application, or a presentation application). The system user interface is optionally a virtual interface that displays control elements (e.g., a volume control element or a focus control element) for controlling one or more aspects or functions of the computer system. As used herein, the term "or" optionally corresponds to an inclusive "or". In some embodiments, although the system user interface is a virtual element, it is optionally separate and / or different from the virtual content. For example, when the system user interface is displayed, it is optionally displayed in a constant immersion state regardless of the immersion level at which the virtual content is displayed. Thus, the system user interface and / or one or more elements of the system user interface are optionally displayed with constant display characteristics regardless of changes in the display characteristics of other virtual elements displayed via the display generation component, as will be described later.

[0231] In some embodiments, the level of immersion includes the associated degree to which virtual content (e.g., a virtual environment and / or virtual content) displayed by a computer system obscures background content (e.g., content other than the virtual environment and / or virtual content) around / behind the virtual content, optionally including the number of items of the displayed background content and / or the displayed visual characteristics (e.g., color, contrast, and / or opacity) of the background content, the angular range of the virtual content displayed via a display generation component (e.g., 60 degrees for content displayed at a low level of immersion, 120 degrees for content displayed at a medium level of immersion, or 180 degrees for content displayed at a high level of immersion), and / or the proportion of the field of view displayed via the display generation component occupied by the virtual content (e.g., 33% of the field of view occupied by the virtual content at a low level of immersion, 66% of the field of view occupied by the virtual content at a medium level of immersion, or 100% of the field of view occupied by the virtual content at a high level of immersion). In some embodiments, the background content is included in the background on which the virtual content is displayed. In some embodiments, the background content includes a user interface (e.g., a user interface generated by the computer system corresponding to an application), virtual objects not associated with or not included in the virtual environment and / or virtual content (e.g., files or other user representations generated by the computer system, etc.), and / or real objects (e.g., passthrough objects representing real objects in the physical environment around the user, which are visible such that they are displayed via the display generation component and / or visible via a transparent or translucent component of the display generation component because the computer system does not obscure / hinder their visibility through the display generation component). In some embodiments, at a low level of immersion (e.g., a first level of immersion), the background, virtual, and / or real objects are displayed in an unobscured manner. For example, a virtual environment having a low level of immersion is optionally displayed concurrently with the background content, which is optionally displayed at full brightness, color, and / or semi-transparency. In some embodiments, at a higher level of immersion (e.g., a second level of immersion higher than the first level of immersion), the background, virtual, and / or real objects are displayed in an obscured manner (e.g., dimmed, blurred, or removed from the display). For example, the corresponding virtual environment having a high level of immersion is displayed without concurrently displaying the background content (e.g., in full-screen or fully immersive mode). As another example, a virtual environment displayed at a medium level of immersion is displayed concurrently with the background content that is dimmed, blurred, or otherwise de-emphasized. In some embodiments, the visual characteristics of the background objects differ among the background objects. For example, at a particular level of immersion, one or more first background objects are more visually de-emphasized (e.g., dimmed, blurred, and / or displayed with increased transparency) compared to one or more second background objects, and one or more third background objects cease to be displayed.

[0232] In some embodiments, when displaying virtual content at a first immersion level and displaying a system user interface that includes an immersion control element, the computer system receives (802b), via one or more input devices, an input that points to the immersion control element, such as Fig. 7A and Fig.7A1User attention 730a therein. In some embodiments, the input directed to the immersion control element includes or is an air gesture or gaze input from the user. In some embodiments, the input directed to the immersion control element includes the user's attention directed to the immersion control element (e.g., the line of sight or gaze directed to the immersion control element), the user's hand in a particular pose that is greater than a threshold hand distance from the immersion control element (e.g., 0.2 cm, 0.5 cm, 1 cm, 2 cm, 3 cm, 5 cm, 10 cm, 20 cm, 40 cm, 100 cm, 200 cm, or 500 cm) (e.g., lifted at a position in front of the user, in a pre-pinching hand shape, or the user's hand is in a pinching hand shape for a certain period of time), or any combination of the user's attention, the user's hand in a particular pose, and / or the user's hand at a greater than threshold hand distance. Additionally, in some embodiments, the input directed to the immersion control element includes vector data corresponding to the movement of the user's hand in a particular direction and / or from a first position to a second position and / or the movement of the user's attention in a particular direction, to indicate the user's request to modify the immersion level by interacting with the immersion control element. For example, the immersion control element is optionally a horizontal or vertical slider bar, and when the computer system displays virtual content at a first immersion level, the horizontal or vertical slider bar displays an indication of the slider or control element at a first position corresponding to the first immersion level on the horizontal or vertical slider bar. The immersion control element is optionally configured to be modified in response to the input directed to the immersion control element. In some embodiments, the input directed to the immersion control element includes vector data corresponding to the movement of the user's hand in a particular direction and / or from a first position to a second position and / or the movement of the user's attention in a particular direction from a first position to a second position, so as to correspond to the request to move the position of the slider to a position corresponding to the vector data and / or move the position of the slider in the direction corresponding to the vector data (e.g., move the slider control element to the right according to the rightward movement of the user's hand, optionally corresponding to increased immersion, and move the slider control element to the left according to the leftward movement of the user's hand, optionally corresponding to decreased immersion). As another example, in some embodiments, the input directed to the immersion control element includes the user's attention directed to the slider control element, the user's hand in a particular pose such as a pinching hand shape, and the movement of the user's hand in a direction when in a particular pose (e.g., corresponding to the pinching hand shape). In some embodiments, the direction and / or magnitude of the change in the immersion / immersion control element is based on the direction and / or magnitude of the hand movement. In some embodiments, the input directed to the immersion control element includes a touch input detected on a touch-sensitive surface (e.g., a touch screen).In some embodiments, the input directed to the immersion control element includes a user pressing a control element on a mouse (e.g., left click). In some embodiments, the input directed to the immersion control element is a gaze input and does not include other inputs such as air gestures.

[0233] In some embodiments, in response to receiving the input directed to the immersion control element, the computer system displays (802c) virtual content at a second immersion level different from the first immersion level via a display generation component according to the input, such as Figure 7BThe three-dimensional environment 704 therein. For example, an input pointing to an immersion control element optionally corresponds to a request to decrease the immersion level (e.g., a leftward or downward movement of the user's hand). When the input pointing to the immersion control element corresponds to a request to decrease the immersion level while the computer system is displaying virtual content at a first immersion level, the computer system optionally displays the immersion control element modified according to the decrease and / or displays the virtual content at a second immersion level lower than the first immersion level. Additionally, when the display generation component displays the virtual content at the second immersion level and the second immersion level is lower than the first immersion level, the display generation component optionally displays less virtual content and / or more portions of the user's physical environment. For example, compared to the first immersion level, portions of the physical environment (e.g., the real environment) become less occluded by the virtual content (e.g., the virtual environment), the virtual content becomes more transparent at the second immersion level than at the first immersion level, the angular range of the virtual content displayed via the display generation component decreases relative to the angular range at the first immersion level, and / or the proportion of the field of view displayed via the display generation component consumed by the virtual environment decreases. As another example, an input pointing to the immersion control element optionally corresponds to a request to increase the immersion level. When the input pointing to the immersion control element corresponds to a request to increase the immersion level while the computer system is displaying virtual content at a first immersion level, the computer system optionally displays the immersion control element modified according to the increase and / or displays the virtual content at a second immersion level higher than the first immersion level. Additionally, when the display generation component displays the virtual content at the second immersion level and the second immersion level is higher than the first immersion level, the display generation component optionally displays more virtual content and / or fewer portions of the user's physical environment. For example, compared to the first immersion level, portions of the physical environment (e.g., the real environment) become more occluded by the virtual content (e.g., the virtual environment), the virtual content becomes less transparent at the second immersion level than at the first immersion level, the angular range of the virtual content displayed via the display generation component increases relative to the angular range at the first immersion level, and / or the proportion of the field of view displayed via the display generation component consumed by the virtual environment increases. Regarding the immersion control element, in some embodiments, the immersion control element is optionally a slider bar, where the slider control element of the slider bar is at a first position corresponding to the first immersion level. Thus, in response to receiving an input pointing to the immersion control element, the slider control element of the slider bar optionally displays at a second position on the slider bar different from the first position, where the second position on the slider bar corresponds to a second immersion level different from the first immersion level.

[0234] Changing the immersion level of virtual content in response to receiving an input pointing to an immersion control element allows a user to easily control the immersion level when using a computer system and reduces errors in immersion control.

[0235] In some embodiments, the virtual content is a virtual reality experience (804a) in which the physical environment of the display generation component is not visible, such as a Fig. 7A and Fig.7A1 corner table 708b that is occluded and not visible in the three-dimensional environment 704 (e.g., the display of a virtual reality (VR) experience (e.g., virtual content) occludes the display of the physical environment via active or passive transparency by the display generation component, where the virtual content is displayed via the display generation component); when the virtual content is displayed at a first immersion level, the virtual content is displayed within an augmented reality experience (804b) in which the physical environment of the display generation component is visible, such as a Fig. 7A and Fig.7A1 visible desk 710 in (e.g., the display of an augmented reality (AR) experience does not (fully) occlude the display of the physical environment via active or passive transparency by the display generation component so as to suppress the visibility of one or more physical objects via the display generation component); when the virtual content is displayed at a first immersion level, the virtual content occupies a first proportion (804c) of the augmented reality experience, such as Fig. 7A and Fig.7A1 virtual environments 722-1a, 722-1b in (e.g., the display device simultaneously displays VR and AR experiences, and the VR experience consumes a first proportion of the AR experience, such as 10%, 20%, 30%, or 40% of the AR experience); and in response to receiving an input pointing to an immersion control element, changing the proportion of the augmented reality experience occupied by the virtual content to a second proportion (804d) different from the first proportion (e.g., less than or greater than the first proportion) (e.g., such as 20%, 30%, 50%, 70%, or 80% of the AR experience), such as the space occupied by the virtual environment 722a in Figure 7B relative to the space occupied by the virtual environments 722-1a, 722-2a in Fig. 7A and Fig.7A1 . In some embodiments, the amount of change in the AR experience is proportional (e.g., indirectly proportional) to the change in the immersion level caused by the input pointing to the immersion control element. For example, when the immersion level is increased, the AR experience optionally decreases proportionally to the change in the immersion level. Changing the proportion of the augmented reality experience occupied by the virtual content in response to receiving an input pointing to an immersion control element increases user control over the AR / VR experience by reducing the input involved in modifying the AR / VR experience and can reduce user fatigue or discomfort caused by using the computer system.

[0236] In some embodiments, in response to receiving an input (806a) directed to an immersion control element, virtual content is displayed (806b) at a third immersion level in accordance with determining that the input corresponds to a request to change the immersion level of the virtual content by a first amount (e.g., increasing the immersion from 0%, 3%, 5%, 10%, or 20% immersion to 50%, 60%, 70%, 80%, or 100% immersion, or similarly decreasing the immersion), such as Figure 7B the immersion of the virtual environment 722a in Fig. 7E ; in accordance with determining that the input corresponds to a request to change the immersion level of the virtual content by a second amount different from the first amount (e.g., increasing the immersion from 0%, 3%, 5%, 10%, or 20% immersion to 50%, 60%, 70%, 80, or 100% immersion, or similarly decreasing the immersion), virtual content is displayed (806c) at a fourth immersion level different from the third immersion level, such as

[0237] the immersion of the virtual environment of Fig. 7A and Fig.7A1 which is no immersion. Thus, the immersion level of the virtual content can be adjusted by a range of values. Changing the immersion level of the virtual content by an amount based on an input directed to the immersion control element increases user control of the virtual experience by reducing the input involved in modifying the immersion level, and can reduce user fatigue or discomfort caused by using the computer system.

[0238] In some embodiments, the virtual content includes a user interface (808) of an application (e.g., an email, internet, or content playback application). The display of the user interface is optionally changed in terms of the immersion level (e.g., transparency or another immersion aspect discussed above with reference to step 802) based on an input directed to an immersion control element, such as Fig. 7A and Fig.7A1One or more of the video application user interfaces 726a). The display of the first user interface and the second user interface is optionally changed at an immersion level (e.g., transparency or another immersion aspect discussed above) based on an input directed to the immersion control element. The immersion levels of the two user interfaces are optionally changed in the same manner and / or by the same amount in response to an input directed to the immersion control element. For example, if the input directed to the immersion control element is an input for increasing immersion, the user interface of the application optionally occupies a larger portion of the three-dimensional environment and / or the display area and / or the user's field of view, or if the input directed to the immersion control element is an input for decreasing immersion, the user interface of the application optionally occupies a smaller portion of the three-dimensional environment and / or the display area and / or the user's field of view. Changing the immersion levels of the user interfaces of multiple different applications by a certain amount based on an input directed to the immersion control element increases user control over the virtual experience of multiple applications without the need for separate inputs to do so and can reduce user fatigue or discomfort caused by using the computer system.

[0239] In some embodiments, the virtual content includes a system virtual environment (812) (e.g., a virtual location, setting, and / or atmosphere displayed via a display generation component of a computer system, such as referenced 7A to 7H and / or as described in reference step 802), such as Fig. 7A and Fig.7A1The background 1 (BKGD1) in the virtual environments 722-1a, 722-2a. For example, the system virtual environment optionally includes a virtual representation of a location corresponding to a geographical location and / or an atmosphere corresponding to the location at a specific time (e.g., a hill with the "HOLLYWOOD" sign during sunny days in Hollywood, California, or the shore of Lake Houston in Houston, Texas during night time, where the night time corresponds to the night at the shore), including simulations of objects in the location (e.g., rocks, wind, water, bugs, birds, etc. as characteristics of the simulated location). The user can interact in the location (e.g., walk, move, turn), and the computer system optionally changes the display based on the user's interaction with the location. For example, when the user is immersed (e.g., fully immersed) in the virtual content and bends down towards the ground while looking at the ground, the ground optionally occupies a larger view of the display of the three-dimensional environment compared to when the user is not looking at the ground, so as to increase the realism of the virtual experience. Similarly, when the user walks towards an object in the system virtual environment, the object optionally occupies more display area relative to the display generation component, so as to increase the realism of the virtual experience. The system virtual environment optionally changes at the immersion level based on an input pointing to an immersion control element. In addition, different applications can be placed within the same system virtual environment. For example, the user interfaces of content playback applications (such as the user interface of a movie application and the user interface of an Internet application) can be placed and / or positioned within the system virtual environment, either concurrently or at different times. Additionally, it should be noted that the system virtual environment is optionally similar or identical to the virtual environment discussed above with reference to step 802, but is the virtual environment displayed when the user does not indicate or select a specific virtual environment to display when displaying the virtual environment. In some embodiments, the system virtual environment is (optionally by default) displayed in response to an input corresponding to a request to display the virtual environment. Changing the immersion level of the system virtual environment by a certain amount based on an input pointing to an immersion control element, increasing user control over the virtual experience by reducing the input involved in modifying the immersion level, and can reduce fatigue or discomfort of the user caused by using the computer system.

[0240] In some embodiments, the virtual content includes a first virtual environment, such as Fig. 7A and Fig.7A1 the virtual environments 722-1a, 722-2a, and the system user interface includes a lighting control element (814a) that can be selected to change the lighting settings of the first virtual environment, such as Figure 7CThe control center user interface 724. In some embodiments, when displaying a first virtual environment where the lighting settings optionally correspond to different simulated times of day in the virtual environment (e.g., daytime (i.e., 1:00 p.m. at a California beach) versus sunset (i.e., 6:00 p.m. at the beach) versus night or before sunrise (i.e., 3:00 a.m. at the beach)) with a first value (e.g., a first brightness, a first color, and / or a first amount of virtual objects (e.g., dew or no dew on the ground of the first virtual environment, or sunlight or no sunlight displayed by the display generation component)), the computer system receives (814b) a second input pointing to a lighting control element via one or more input devices, such as Figure 7C the user's attention 730d. In some embodiments, in response to receiving the second input, the computer system displays (814c) the first virtual environment via the display generation component with a second value different from the first value that optionally corresponds to the simulated time of day in the virtual environment (e.g., daytime (i.e., 1:00 p.m. at a California beach) versus sunset (i.e., 6:00 p.m. at the beach) versus night or before sunrise (i.e., 3:00 a.m. at the beach)) in the lighting settings, such as Figure 7B the background 2 (BKGD2) in the virtual environments 722-1a, 722-2a. The first value and the second value are optionally associated with different lighting characteristics. For example, the first value is optionally associated with a brighter (e.g., greater intensity) and / or lighter display value compared to the second value. Additionally, the second value optionally includes more or fewer virtual objects compared to the first value. Thus, in addition to controlling other characteristics of the first virtual environment (such as the amount of virtual objects displayed in the first virtual environment), the lighting settings at the first value optionally relate to the lighting characteristics applied to the virtual objects in the first virtual environment. In some embodiments, in response to the selection of a lighting control element that can be selected to change the lighting settings of the first virtual environment, the system user interface displays a set of selectable options for setting the lighting settings of the first virtual environment. In some embodiments, the lighting control element that can be selected to change the lighting settings of the first virtual environment is displayed concurrently with the display of other selectable lighting control elements. Displaying the first virtual environment with the second value in the lighting settings after receiving a second input pointing to the lighting control element when the first virtual environment is displayed with the first value in the lighting settings increases user control over the virtual experience by reducing the inputs involved in changing the lighting settings, and can reduce adverse health effects on the user due to using the computer system.

[0241] In some embodiments, when displaying a first virtual environment with a corresponding value for the lighting setting, the computer system receives (816a), via one or more input devices, a third input corresponding to a request to display a second virtual environment different from the first virtual environment (e.g., a virtual location or setting displayed via a display generation component of the computer system, such as those referenced 7A to 7H and / or as described in reference method 800), such as Fig. 7E user attention 740d. For example, the third input is optionally a selection of a selectable element corresponding to the second virtual environment displayed in an environment selection user interface, such as those described in reference steps 832 - 836 below. In some embodiments, in response to receiving the third input (816b), and based on determining that the corresponding value is a first value, the computer system displays (816c) the second virtual environment with the lighting setting having the first value, such as according to Figure 7C display mode 2 option 736b, and based on determining that the corresponding value is a second value, the computer system displays (816d) the second virtual environment with the lighting setting having the second value, such as according to Figure 7C display mode 1 option 736a. Thus, the lighting setting is optionally persistent across different virtual environments being displayed. It is contemplated that although the corresponding value of the lighting setting is optionally persistent when transitioning from the first virtual environment to the second virtual environment in response to the third input, in some embodiments, the corresponding value of the lighting setting in the second virtual environment optionally corresponds to the time of day in the second virtual environment, which is different from the time of day corresponding to the corresponding value of the lighting setting in the first virtual environment. For example, when the corresponding value of the lighting setting in the first virtual environment is a first value (e.g., corresponding to 10:00 am in the first virtual environment), the second virtual environment displayed with the lighting setting having the first value optionally corresponds to the time of day in the second virtual environment (e.g., 1:00 pm in the second virtual environment), which is different from the time of day corresponding to the first value of the lighting setting in the first virtual environment. Similarly, when the corresponding value of the lighting setting in the first virtual environment is a second value (e.g., corresponding to 10:00 pm in the first virtual environment), the second virtual environment displayed with the lighting setting having the second value optionally corresponds to the time of day in the second virtual environment (e.g., 11:30 pm in the second virtual environment), which is different from the time of day corresponding to the second value of the lighting setting in the first virtual environment. Making the lighting setting persistent across virtual environments alleviates lighting disruptions when switching the display of virtual environments and reduces the input involved in switching the display of virtual environments.

[0242] In some embodiments, the system user interface includes an automatic lighting control element that can be selected to automatically set (818) the lighting settings of a first virtual environment (e.g., set to a first value or a second value) at least based on the current date and time of the computer system (and / or at the location of the computer system), such as the selectable display mode 3 option 736c of FIG. 7c. For example, the current date and time are optionally determined by the global positioning system (GPS) component of the computer system. For example, at noon on a summer day in California, the lighting settings are automatically set to a first value, and at 10 PM on the same summer day in California, the lighting settings are automatically set to a second value. Thus, when the automatic lighting control element is selected, both the date and time at the computer system and the location of the computer system are optionally used to determine the lighting settings of the first virtual environment. In some embodiments, the specific times during which the lighting settings of the first virtual environment switch are user-configurable. For example, the user can set the lighting settings to automatically switch to the first value when the current date and time is 1:32 PM, and / or to automatically switch to the second value when the current date and time is 8:03 PM. Additionally or alternatively, the times at which the lighting settings switch are optionally based on the sunrise / sunset times at the location of the computer system and thus optionally (automatically (e.g., without user input)) change over the course of a year as the sunrise and sunset times change for the location of the computer system. Additionally or alternatively, when the location of the computer system changes (e.g., moves towards or away from the equator or to a different time zone), the times at which the lighting settings (automatically (e.g., without user input)) switch are optionally changed so as to correspond, for example, to the sunrise and sunset times of the current location of the computer system. Including an automatic lighting control element in the system user interface that can be selected to automatically set the lighting settings of the first virtual environment based on the current date and time reduces unwanted lighting disruptions for the user between the physical environment and the virtual environment of the computer system, reduces the number of inputs involved in switching lighting settings, and reduces fatigue or discomfort experienced by the user due to using the computer system.

[0243] In some embodiments, the automatic lighting control element can be selected to further based on the type of corresponding virtual content concurrently displayed with the first virtual environment (such as based on Figure 7CThe video application user interface 726a) automatically sets the lighting settings (820) of the first virtual environment. For example, at noon on a summer day in California, when the type of the corresponding virtual content includes or is an email application, the lighting settings are automatically set to a first value (e.g., daytime settings), while at the same time, when the type of the corresponding virtual content includes or is a user interface of a movie, TV show, or video playback application, the lighting settings are automatically set to a second value (e.g., nighttime settings). Thus, based on the selection of the automatic lighting control element and the type of the corresponding virtual content being concurrently displayed with the first virtual environment, the lighting settings are optionally automatically set to a value that enhances the user's immersive experience of the content playback application. The inclusion of an automatic lighting control element in the system user interface that can be selected to automatically set the lighting settings of the first virtual environment based on the current day time and the type of the virtual content being displayed reduces the amount of input involved in setting the lighting settings.

[0244] In some embodiments, the system user interface includes (822a): a lighting control element that can be selected to change the lighting settings of the first virtual environment, where the lighting control element that can be selected to change the lighting settings of the first virtual environment can be selected to set the lighting settings of the first virtual environment to a second value (822b) (e.g., daytime lighting settings), such as according to Figure 7C the display mode 1 option 736a; and a second lighting control element that can be selected to set the lighting settings of the first virtual environment to a first value (822c) (e.g., nighttime lighting settings), such as according to Figure 7CDisplay mode 1 option 736a. The lighting control element and the second lighting control element are optionally located at a different level of the system user interface than the immersion control element (e.g., not displayed concurrently with the immersion control element). For example, to reach these elements, as described in reference step 822, the input optionally points to a system environment control element (e.g., a slider, a dial, a switching device, a segmented control, or another type of control element), which is optionally displayed concurrently with the immersion control element in the system user interface. In response to detecting an input pointing to the system environment control element, such as the user's attention and / or the user's hand pointing to the system environment control element for a certain period of time, the system user interface optionally displays the lighting control element and the second lighting control element. Additionally, in some embodiments, the automatic lighting control element discussed above is optionally displayed concurrently with the display of the lighting control element and the second lighting control element. In some embodiments, the method includes: detecting an input towards the lighting control element or the second lighting control element; and in response to detecting the input, the computer system optionally sets the lighting setting for the first visual environment to a first value or a second value based on which lighting control element the input points to. For example, in response to detecting an input pointing to the lighting control element, the computer system optionally sets the lighting setting for the first virtual environment to the second value. Similarly, in response to detecting an input pointing to the second lighting control element, the computer system optionally sets the lighting setting for the first virtual environment to the first value. Displaying user-selectable options for switching lighting settings increases user control of the virtual reality experience during the virtual reality experience by reducing the input involved in setting the lighting setting.

[0245] In some embodiments, the lighting control element that can be selected to change the lighting setting of the first virtual environment can be selected to set the lighting setting of the first virtual environment to the second value and the lighting control element is displayed with an appearance independent of the characteristics of the first virtual environment in the case where the lighting setting has the second value (824a), such as Figure 7C display mode 2 option 736b, and the second lighting control element that can be selected to set the lighting setting of the first virtual environment to the first value is displayed with an appearance independent of the characteristics of the first virtual environment in the case where the lighting setting has the first value (824b), such as Figure 7CDisplay mode 1 option 736a (e.g., the lighting control element displayed in the system user interface is displayed as a visual indication (e.g., glyph) that does not include a representation of the currently active or currently displayed virtual environment in the case where the lighting setting has a second value or a first value). As another example, prior to the display generation component displaying the first virtual environment, a system user interface including the lighting control element is optionally displayed. The lighting control element is optionally displayed without a preview of the virtual environment (optionally because no virtual environment is currently selected, active, and / or being displayed by the display generation component). In some embodiments, when the display generation component displays the first virtual environment and / or when the first virtual environment is selected for display in a virtual reality experience, the lighting control element is optionally displayed without a preview (or aspect) of the virtual environment. For example, the lighting control element optionally has the same visual appearance regardless of whether the computer system is displaying a virtual environment and / or regardless of whether the computer system is displaying a first virtual environment or a second virtual environment. Displaying user-selectable options for switching lighting settings with a consistent visual appearance reduces the likelihood of errors in the use of the computer system.

[0246] In some embodiments, the system user interface includes (826a): a lighting control element that can be selected to change the lighting setting of the first virtual environment, wherein the lighting control element that can be selected to change the lighting setting of the first virtual environment can be selected to set the lighting setting of the first virtual environment to a second value (e.g., a nighttime lighting setting), and wherein the lighting control element that can be selected to set the lighting setting of the first virtual environment to the second value includes a visual representation (826b) of the first virtual environment in the case where the lighting setting has the second value, such as according to Figure 7C display mode 2 option 736b; and a second lighting control element that can be selected to set the lighting setting of the first virtual environment to a first value, wherein the second lighting control element that can be selected to set the lighting setting of the first virtual environment to the first value (e.g., a daytime lighting setting) includes a visual representation (826c) of the first virtual environment in the case where the lighting setting has the first value, such as according to Figure 7CThe display mode 1 option 736a. In some embodiments, the automatic light control element discussed above includes a visual representation of a first virtual environment in a first visual portion of the automatic light control element when the light setting has a second value, and includes a visual representation of the first virtual environment in a second visual portion of the automatic light control element when the light setting has a first value. In some embodiments, the automatic light control element includes a representation of a first virtual environment when the light setting has a value based on the current time of day and the location of the computer system. User-selectable options for switching the light setting are displayed together with a preview of the corresponding light setting applied to the current virtual reality experience, increasing user control over the virtual reality experience by reducing the inputs and potential errors involved in setting the light setting.

[0247] In some embodiments, when the first virtual environment is not displayed and the system user interface including the light control element is displayed, the computer system receives (828a) via one or more input devices a third input directed at the light control element corresponding to a request to change the light setting for the first virtual environment from a first value (e.g., daytime light setting) to a second value (e.g., nighttime light setting), such as Figure 7C the user attention 730d (e.g., the third input optionally includes one or more aspects of the input directed at the immersion control element discussed above with reference to step 802, such as an air gesture or a gaze input from the user, and / or another aspect of the input directed at the immersion control element and directed at the light control element). In some embodiments, in response to receiving the third input, the computer system at least partially displays (828b) the first virtual environment (e.g., displays a preview of the first virtual environment outside the system user interface in the three-dimensional environment with the light setting having the second value) via a display generation component (e.g., in at least a portion of the three-dimensional environment displayed via the display generation component) with the light setting having the second value, such as Fig.7DVirtual environments 722a-1, 722-2a. In some embodiments, displaying a preview of the first virtual environment corresponds to displaying at least a portion of the first virtual environment in a three-dimensional environment. In some embodiments, a third input is received while the first virtual environment is concurrently displayed and / or active with the system user interface, wherein the lighting setting of the first virtual environment has a first value. Accordingly, a preview of the lighting setting at a second value is optionally initiated on the first virtual environment that was already being displayed when the third input was received. In some embodiments, the immersion level at which the first virtual environment is displayed when the third input is detected is increased, such as discussed with reference to steps 802 and 806. Displaying a preview of the first virtual environment outside the system user interface with the lighting setting at a second value in response to a selection of a lighting control element provides feedback on the appearance of the first virtual environment with the lighting setting applied, thereby reducing errors in the use of the computer system and reducing the inputs involved in correcting such errors.

[0248] In some embodiments, when at least a portion of the first virtual environment is displayed after receiving the third input and with the lighting setting at a second value, the computer system automatically stops displaying the first virtual environment (830) based on determining that one or more criteria are met (such as when at least a portion of the first virtual environment is displayed with the lighting setting at a second value and / or a predetermined amount of time has elapsed since receiving the third input (e.g., 0.1 s, 0.5 s, 1 s, 2 s, 5 s, 10 s, 45 s)), such as Fig. 7E exemplified by the lack of the virtual environment 722 (and / or returning to the display state that the display generation component had when the third input was received, which optionally is a display state including the display of the first virtual environment with the lighting setting at a first value or a display state that does not include the display of the first virtual environment outside the system user interface (e.g., different from the representation of the first virtual environment in the system user interface)). In some embodiments, the immersion level corresponding to displaying at least a portion of the first virtual environment with the lighting setting at a second value is decreased, such as the decrease in immersion discussed with reference to steps 802 and 806. Stopping the display of the preview of the first virtual environment outside the system user interface with the lighting setting at a second value in response to certain criteria being met reduces the inputs for returning to a previous state of the computer system and reduces the disruption to the user experience of the computer system.

[0249] In some embodiments, the system user interface includes a system virtual environment control element (832a) that can be selected to initiate a process of changing the current system virtual environment from the first virtual environment to a second virtual environment, such as Figure 7CChanging the background option 736d (e.g., the system virtual environment control element is optionally displayed concurrently with the display of the lighting control element, the second lighting control element, and / or the automatic lighting control element discussed with reference to steps 820 and 822). In some embodiments, when the system user interface is displayed (and optionally when the display generation component is displaying the first virtual environment), the computer system receives (832b) a second input corresponding to the selection of the system virtual environment control element, such as user attention 730e (e.g., the second input optionally includes one or more aspects of the input discussed above with reference to step 802 for pointing to the immersion control element, such as an air gesture or a gaze input from the user, and / or another aspect of the input for pointing to the immersion control element and pointing to the system virtual environment control element).

[0250] In some embodiments, in response to receiving the second input, the computer system displays (832c) a system virtual environment control user interface via the display generation component, the system virtual environment control user interface including one or more selectable options for changing the current system virtual environment from the first virtual environment to the second virtual environment, such as Fig. 7E the control center user interface 724a (and / or for causing the second virtual environment to be displayed outside the boundary and / or region in which the system user interface or the system virtual environment control user interface is displayed), such as the system virtual environment described with reference to step 812. In some embodiments, a corresponding selectable option among the one or more selectable options can be selected to change the current system virtual environment from the first virtual environment to the corresponding virtual environment. The system virtual environment control user interface is optionally as Fig. 7E illustrated and / or as described with reference to steps 834 and 836. Displaying user-selectable options for switching the system virtual environment increases user control over the virtual reality experience by reducing the input involved in switching virtual environments.

[0251] In some embodiments, the system virtual environment control user interface includes one or more selectable options (834) for displaying one or more atmosphere effects (e.g., virtual reality, augmented reality, or another computer-aided reality simulation of sunlight, rain, dew, clouds, or another atmosphere effect) on one or more portions of the physical environment visible via the display generation component, such as Fig. 7EOptional options 742a, 742b, 742c (e.g., an atmosphere effect is optionally applied to the physical environment visible via the display generation component rather than to the virtual environment displayed via the display generation component). For example, the atmosphere effect optionally includes one or more of virtual reality (VR), augmented reality (AR), or another computer-aided reality effect applied to one or more portions of the physical environment to simulate an atmosphere effect. In some embodiments, the atmosphere effect includes a visual modification of at least a portion of the three-dimensional environment (e.g., not associated with an object in the three-dimensional environment), such as a portion of the three-dimensional environment corresponding to the physical environment and / or virtual content. For example, a display of an environmental lighting effect (e.g., sunrise, sunset, moonlight, starlight, or another environmental lighting effect), a fog effect, a haze effect, and / or a smoke / particle effect. In some embodiments, the atmosphere effect is an effect in which the air or empty space in the three-dimensional environment appears to be filled with a physical effect. One or more optional options for displaying one or more atmosphere effects optionally include a visual representation of the corresponding one or more atmosphere effects. Providing optional options for applying an atmosphere effect to the physical environment increases user control of the virtual reality / augmented reality experience by reducing the input involved in applying the atmosphere effect, and can reduce adverse health effects on the user caused by using the computer system or by the physical environment itself (e.g., by using the computer system to set certain atmosphere effects to reduce the amount of blue light incident on the user's eyes due to being in the physical environment) and / or reduce fatigue or discomfort caused by the user using the computer system.

[0252] In some embodiments, one or more optional options for changing the current system virtual environment from a first virtual environment to a second virtual environment include (836a): a first optional option (836b) that can be selected to set the current system virtual environment to the first virtual environment (e.g., the first optional option optionally includes a visual representation (e.g., a preview) of the first virtual environment, such as a picture or visual preview of a California beach when the first optional option corresponds to a virtual environment corresponding to a California beach), such as Fig. 7E optional option 740a; and a second optional option (836c) that can be selected to set the current system virtual environment to the second virtual environment, such as Fig. 7EOptional option 740b. The second optional option optionally includes a visual representation (e.g., a preview) of the second virtual environment, such as a picture or visual preview of a stream in Texas when the second virtual environment corresponds to a virtual environment of a stream corresponding to the state of Texas. In some embodiments, the method includes detecting: an input pointing to an optional option (such as pointing to the first optional option or the second optional option); and in response to detecting the input, the computer system optionally performs an operation corresponding to the selection of the optional option. For example, in response to detecting an input pointing to the first optional option, the computer system optionally sets the current system virtual environment to the first virtual environment. Similarly, in response to detecting an input pointing to the second optional option, the computer system optionally sets the current system virtual environment to the second virtual environment. Providing optional options for setting different system virtual environments gives the user more control over the virtual experience by reducing the input involved in setting the system virtual environment.

[0253] In some embodiments, the virtual content (optionally including the user interface of the application) is displayed within a three-dimensional environment, and the system user interface includes a first selectable option (836a) that can be selected to enable or disable the automatic de-emphasis of one or more portions of the three-dimensional environment outside of one or more portions of the virtual content (such as optionally outside of the user interface of an application different from the system user interface (e.g., a content playback application or a photo application)), such as Fig. 7A and Fig.7A1 the auto-dimming user interface element 728c. The automatic de-emphasis is optionally applied to one or more virtual contents, such as the virtual environment and the user interface of the application. When multiple user interfaces of the application are displayed via the display generation component, the automatic de-emphasis is optionally applied to the first user interface of the first application and not to the second user of the second application based on the computer system determining which user interface the user is focused on and / or which user interface is currently being interacted with and / or consumed (e.g., the automatic de-emphasis is applied to the virtual content and / or object that is not currently being interacted with and / or consumed). In some embodiments, the first selectable option is displayed concurrently with the display of the lighting control element discussed above.

[0254] In some embodiments, when displaying the virtual content (838b), based on determining that one or more criteria are met and the automatic de-emphasis of one or more portions of the three-dimensional environment outside of one or more portions of the virtual content (such as outside of the user interface of an application optionally different from the system user interface (e.g., a content playback application or a photo application)) is enabled, the computer system, via the display generation component, at a first visual emphasis level with respect to one or more portions of the three-dimensional environment outside of one or more portions of the virtual content (such as Figure 7Hdisplay (838c) virtual content with a level of visual emphasis relative to a three-dimensional environment 704 outside of a video application user interface 724a. One or more criteria optionally include criteria that are satisfied when user interface activity and / or is being interacted with of a certain type of application (such as a content playback application, or a photo application, or an email application, or another type of application), and / or criteria that are satisfied when the user interface of that type of application is currently being displayed and / or active (such as with respect to playing content (e.g., playing a video)). Optionally, the virtual content is set to be displayed with a first level of visual emphasis relative to one or more portions outside of one or more portions of the virtual content with respect to the three-dimensional environment by performing the following: not emphasizing one or more portions of the three-dimensional environment outside of one or more portions of the virtual content relative to the virtual content (e.g., reducing the brightness level of the virtual environment around the virtual content, reducing the opacity of the virtual environment, reducing the clarity of the virtual environment (e.g., increasing the blurriness of the virtual environment), reducing the color saturation of the virtual environment, changing the light settings of the virtual environment to a "dark" setting as discussed with respect to Figure 7H and / or changing another lighting setting that is applied to one or more portions of the three-dimensional environment outside of one or more portions of the virtual content as discussed above with respect to the first virtual environment), and / or emphasizing one or more portions of the virtual content (e.g., the user interface of the application) relative to portions outside of the virtual content (e.g., increasing the brightness, size, saturation, and / or another visual characteristic of the one or more portions).

[0255] In some embodiments, in accordance with determining that one or more criteria (optionally including criteria that are satisfied when a de-emphasis instruction is received at the computer system) are satisfied and automatic de-emphasis of one or more portions of the three-dimensional environment outside of one or more portions of the virtual content (such as outside of the user interface of an application (e.g., a content playback application or a photo application) optionally different from the system user interface) is disabled, the computer system displays (838d) the virtual content via a display generation component with a second level of visual emphasis relative to one or more portions of the three-dimensional environment outside of one or more portions of the virtual content (such as Fig. 7A and Fig.7A1 the level of visual emphasis relative to the three-dimensional environment 704 outside of the video application user interface 724a), where the second level of visual emphasis is less than the first level of visual emphasis. The second level of visual emphasis optionally is a default level of visual emphasis as if the one or more criteria were not satisfied. For example, portions of the three-dimensional environment around the virtual content optionally are not visually de-emphasized relative to the virtual content.

[0256] In some embodiments, based on determining that one or more criteria are not met (optionally including criteria that are met when user interface activity of a certain type of application (such as a content playback application, or a photo application, or an email application, or another type of application) and / or is being interacted with, and / or criteria that are met when the user interface of the type of application is currently being displayed and / or active (such as regarding playing content (e.g., playing a video))), the computer system via the display generation component at a second visual emphasis level for one or more portions outside of one or more portions of the virtual content relative to the three-dimensional environment (such as Fig. 7A and Fig.7A1Displays (838e) virtual content with a level of visual emphasis relative to a three-dimensional environment 704 outside of a video application user interface 724a. In some embodiments, a first selectable option may be selected to globally enable or disable automatic de-emphasis. In some embodiments, an individual application may be configured to enable or disable automatic de-emphasis of one or more portions of the three-dimensional environment outside of one or more portions of the virtual content independent of or overriding the selection or deselection of the first selectable option. In some embodiments, when a virtual environment is displayed with the virtual content of step 838, automatic de-emphasis causes the virtual environment to be displayed with a lighting setting that is set to a nighttime lighting setting (optionally as a supplement to or alternative to simply dimming (or equivalent treatment) the portion of the three-dimensional environment outside of the virtual content without changing the lighting setting of the virtual environment). In some embodiments, automatic de-emphasis causes the virtual environment to be dimmed (or equivalent treatment) without changing the lighting setting. In some embodiments, automatic de-emphasis occurs without a user input specifically for de-emphasis, such as a user input pointing to the first selectable option or another selectable option for configuring automatic de-emphasis on a per-application basis, as discussed above (e.g., the user input may be doing something else, like playing content). Thus, automatic de-emphasis optionally occurs in response to the computer system receiving a user input corresponding to a request to perform an action different from automatic de-emphasis. In some embodiments, when one or more criteria are subsequently no longer met, such as the closing of a content playback application, or the user's attention being directed outside of the content playback application for a certain period of time (e.g., 0.9s, 10s, 30s, or another period), optionally reduce or eliminate the various changes in de-emphasis described above with respect to one or more portions of the three-dimensional environment outside of one or more portions of the virtual content (optionally without a user input specifically or specifically for re-emphasizing one or more portions of the three-dimensional environment). Providing selectable options for automatically changing the visual emphasis outside of a portion of virtual content increases user control of the virtual experience by reducing the input involved in changing the visual emphasis, reducing distractions outside of the virtual content, and may reduce user fatigue or discomfort caused by using the computer system.

[0257] In some embodiments, when the system user interface is not being displayed and while a first virtual content (e.g., a first virtual environment, a first virtual environment at a first immersion level, a first set of user interfaces of an application, a first atmosphere effect, a first location, or other virtual content as described above with reference to step 802) is being displayed via a display generation component, the computer system receives (840a) a second input corresponding to a request to display the system user interface via one or more input devices (e.g., the second input optionally includes one or more aspects of the input discussed above with reference to step 802 that points to an immersion control element and corresponds to the request to display the system user interface), such as an input including Fig. 7A and Fig.7A1 one or more aspects of the user attention 730a. In some embodiments, in response to receiving the second input, the computer system displays (840b) the first virtual content and the system user interface via the display generation component (e.g., the system user interface optionally blurs the first virtual content or is displayed in front of the first virtual content (e.g., between the user's viewpoint and the first virtual content)), such as Fig. 7A and Fig.7A1 the control center user interface 724a.

[0258] In some embodiments, when the system user interface is not being displayed and while a second virtual content different from the first virtual content (e.g., a different virtual environment, a first virtual environment at a second immersion level different from the first immersion level, a second set of user interfaces of an application, a second atmosphere effect, a second location, or other virtual content as described above with reference to step 802) is being displayed via a display generation component, the computer system receives (840c) a third input corresponding to a request to display the system user interface via one or more input devices (e.g., the third input optionally includes one or more aspects of the second input and / or the input discussed above with reference to step 802 that points to an immersion control element and corresponds to the request to display the system user interface), such as an input including Fig. 7A and Fig.7A1 one or more aspects of the input from the hand 732. In some embodiments, in response to receiving the third input, the computer system displays (840d) the second virtual content and the system user interface, such as Fig. 7A and Fig.7A1The control center user interface 724a (e.g., the system user interface optionally obscures or is displayed in front of the second virtual content (e.g., between the user's viewpoint and the second virtual content)). In some embodiments, the system user interface is locked to the viewpoint when displayed, as previously described earlier in this disclosure. Providing access to the system control user interface from different virtual experiences presented by the display generation component provides consistent interaction with the computer system, thereby reducing errors in the use of the computer system.

[0259] In some embodiments, the second input and the third input correspond to a gaze input (842), such as an input including one or more aspects of the user's attention 730a including Fig. 7A and Fig.7A1 The second input and the third input optionally correspond to the user's attention of the computer system being directed to specific portions of the first virtual content or the second virtual content, such as portions of the first virtual content or the second virtual content that can be gazed at to initiate the display of the system user interface in the three-dimensional environment simulated and / or presented by the display generation component (optionally without input other than the user's attention (e.g., the user's attention directed to the gaze-selectable portion for a period longer than a time threshold such as 0.1 second, 0.3 second, 0.5 second, 1 second, 2 seconds, 3 seconds, 5 seconds, 10 seconds, 20 seconds, or 30 seconds or another time threshold)). In an example, the second input and the third input optionally include the user's attention directed to the top center region of the user's field of view upward in the three-dimensional environment, optionally for a predetermined period of time (e.g., 0.5 s, 1 s, 5 s, 20 s, or another predetermined period of time), which optionally causes the display of the system user interface, such as the control center user interface, as illustrated and discussed with reference to Fig. 7A and Figure 7B In some embodiments, in addition to hand gestures performed by the user, such as the air gestures of pointing to the immersion control element, pointing to the gaze-selectable portion of the first virtual content or the second virtual content as discussed above with reference to step 802, the second input and the third input correspond to the user's attention of the computer system as discussed above. In fact, in some embodiments, the portion of the first virtual content or the second virtual content that can be gazed at to cause the display of the system user interface can alternatively or additionally be selected via gaze and air gestures. Displaying the system control user interface in response to detecting the user's attention (e.g., the user's gaze) is an effective way to display the system user interface, which can reduce user fatigue or strain caused by the display of the system user interface and increase user control by reducing the input involved in accessing the control user interface.

[0260] It should be understood that the specific order in which the operations in method 800 are described is merely exemplary and is not intended to indicate that the described order is the only order in which these operations can be performed. Those of ordinary skill in the art will envision various ways to reorder the operations described herein.

[0261] 9A to 9E Illustrates examples for controlling audio settings of a virtual environment according to some embodiments.

[0262] Fig. 9A Illustrates a computer system 101 in a real-world environment 902 according to some embodiments, the computer system displaying a three-dimensional environment 904 via a display generation component (e.g., display generation component 120 of FIG. 1), the three-dimensional environment including a virtual environment 912a displayed at a first level of immersion as indicated by a current level of immersion indicator 916. As referred to above with reference to FIGS. 1 to Figure 6 As described, computer system 101 optionally includes a display generation component (e.g., a touchscreen) and a plurality of image sensors (e.g., Figure 3 image sensor 314). Additionally, computer system 101 is optionally as described with reference to FIGS. 1 to 7. In some embodiments, the user interface described below is implemented on a head-mounted display that includes: a display generation component that displays the user interface to the user, and sensors that detect movement of the physical environment and / or the user's hand (such as movement that is interpreted by the computer system as a gesture such as an air gesture), and / or sensors that detect the user's gaze (e.g., inward-facing sensors towards the user's face). The figures herein illustrate a top view 918 of the three-dimensional environment presented to the user (and displayed by the display generation component of computer system 101) and the physical environment and three-dimensional environment 904 associated with computer system 101, the top view being used to illustrate the relative positions of objects in the real-world environment and the positions of virtual objects in the three-dimensional environment.

[0263] As Fig. 9AAs shown, computer system 101 captures one or more images of the real-world environment 902 (e.g., operating environment 100) around computer system 101, including one or more objects in the real-world environment 902 around computer system 101. In some embodiments, computer system 101 displays a representation of the real-world environment 902 in a three-dimensional environment 904. For example, the three-dimensional environment 904 includes a room that includes a representation of a desk 914a (desk 914b in top view 916), which representation is optionally a photo-realistic representation of the desk 914a, a simplified representation, a cartoon, a comic, a see-through visibility of the desk through the display generation component 120, etc. Although not shown in the three-dimensional environment 904, the room includes a real desk 905b that is occluded by the virtual environment 912a, as shown in top view 918. Additionally, as shown in top view 918, a user 920 of computer system 101 is sitting on a couch 919 and is interacting with computer system 101 (e.g., if computer system 101 is a head-mounted device, the user is holding or wearing computer system 101).

[0264] In Fig. 9A the illustrated embodiment, computer system 101 displays a three-dimensional environment 904 that includes a first user interface (e.g., a system user interface) of a control center user interface 924a and a music application user interface 926a (e.g., a user interface of an application that optionally includes one or both of video and audio content for playback). As shown in top view 918, the control center user interface 924b and the video application user interface 926b are located at different positions in the three-dimensional environment 904. The first user interface of the control center user interface 924a includes an immersion slider user interface element 928a, a system environment settings user interface element 928b, an auto-dim user interface element 928c, a volume control user interface element 928d, and a focus mode control user interface element 928e. The immersion slider user interface element 928a is displayed at a first fill level corresponding to the current immersion level (e.g., the current immersion level shown in the immersion indicator 916). Further details regarding immersion are described with reference to method 1000. Similarly, the volume control user interface element 928d includes a display of a slider at a position corresponding to the current volume level of computer system 101. Further details regarding the control center user interface 924a are described with reference to methods 800, 1000, and / or 1200.

[0265] In Fig. 9AIn an exemplary embodiment, the computer system 101 is associated with audio parameters. For example, in the audio legend 934, an audio parameter A, optionally corresponding to a representative number of audio point sources simulated in the three-dimensional environment 904 (e.g., groups 930, 931 corresponding to audio point sources generated by the computer system 101 as part of the virtual environment 912a), is set to 8; an audio parameter 1 corresponding to the system environment volume or the virtual environment volume is set to a first level; an audio parameter 2 corresponding to the application volume level, such as the volume level of the user interface of an application in the three-dimensional environment (such as the music application user interface 926a), is set to the first level; and an audio parameter 3 corresponding to the volume associated with the virtual avatar (e.g., the virtual representations of persons 933a and 935a in the three-dimensional environment 904 and / or the virtual environment 912a simulated by the computer system 101) is set to the first level. Further details regarding the audio point sources are described with reference to methods 800, 1000, and / or 1200.

[0266] In Fig. 9A an exemplary embodiment, the user attention 930a, 930b and the input from the hand 932 of the user 920 alternatively point to the immersion slider user interface element 928a and the volume control user interface element 928d. In some embodiments, the user interface element can be selected via user attention or the input from the hand 932 or via a combination of both the user attention 930a, 930b and the input from the hand 932, and such characteristics of the input and the processes for detecting such input are described in more detail with reference to methods 800, 1000, and / or 1200.

[0267] Fig. 9B An example shows the three-dimensional environment 904 and the associated volume level in response to user input pointing to Fig. 9A the volume control user interface element 928d. In Fig. 9B an exemplary embodiment, the user input pointing to the volume control user interface element 928d is an input to decrease the volume level of the computer system 101. In response, an audio parameter A, corresponding to a representative number of audio point sources (e.g., groups 930, 931) in the three-dimensional environment 904 (e.g., generated by the computer system 101 as part of presenting the virtual environment 912a), is now set to 4; an audio parameter 1 corresponding to the system environment volume or the virtual environment volume is decreased to a second audio level lower than the first level it had in Fig. 9A ; an audio parameter 2 corresponding to the application volume level (such as the volume level of the user interface of an application in the three-dimensional environment such as the music application user interface 926a) is decreased to a level lower than the first level it had in Fig. 9Aa second audio level lower than the first level in Fig. 9A and the audio parameter 3 corresponding to the volume associated with the virtual avatar is reduced to a second audio level lower than the first level in FIG. 9A to FIG. 9B It is contemplated that, in response to an input pointing to the volume control user interface element 928d, one or more of the audio parameter A, audio parameter 1, audio parameter 2, and audio parameter 3 change in amplitude and / or frequency by a similar or different amount, optionally in the same direction (e.g., increased or decreased).

[0268] Fig. 9C illustrates a three-dimensional environment 904 in response to a user input pointing to Fig. 9A the volume control user interface element 928d. For example, such an input is optionally a gaze and dwell input, such as the user's gaze on the volume control user interface element 928d and its dwell for a threshold period of time, or another type of input, such as described in reference method 1000. In Fig. 9C the illustrated embodiment in, a second user interface of the display control center user interface 924a is displayed, the second user interface including a volume control user interface element 940a corresponding to the audio parameter 1, a volume control user interface element 940b corresponding to the audio parameter 2, a volume control user interface element 940c corresponding to the audio parameter 3, and selectable options for enabling the following movement settings: movement setting A 942a (e.g., a head-tracked spatial audio setting, where the audio associated with the virtual environment 912a is optionally generated based on the pose of the head of the user 920 of the computer system 101, which is optionally referenced or relative to a virtual object or other virtual content in the virtual environment 912a (such as the applied music user interface 926a and the representative audio point sources for which the computer system generates audio (e.g., groups 930, 931)); or movement setting B 942b (e.g., a non-head-tracked spatial audio setting, where the audio associated with the virtual environment 912a is optionally generated independently of the pose of the head of the user 920 of the computer system 101, which is optionally referenced or relative to a virtual object or other virtual content in the virtual environment 912a (such as the applied music user interface 926a and the representative audio point sources for which the computer system generates audio (e.g., groups 930, 931)). The slider levels of the volume control user interface element 940a corresponding to the audio parameter 1, the volume control user interface element 940b corresponding to the audio parameter 2, and the volume control user interface element 940c corresponding to the audio parameter 3 respectively correspond to the parameter levels in the audio legend 934.

[0269] In Fig. 9CIn the illustrated embodiments, the user attention 930c, 930d, 930e and the input from the hand 932 of the user 920 alternatively point to the volume control user interface element 940a corresponding to audio parameter 1, the volume control user interface element 940b corresponding to audio parameter 2, the volume control user interface element 940c corresponding to audio parameter 3, and the selectable option for enabling mobile setting A 942a or mobile setting B 942b. In some embodiments, the user interface element can be selected via user attention or the input from the hand 732 or via a combination of both the user attention 930c, 930d, 930e and the input from the hand 932, and such characteristics of the input and the processes for detecting such input are described in more detail with reference to methods 800, 1000 and / or 1200.

[0270] Fig.9D Illustrates a three-dimensional environment 904 in response to Fig. 9C the user input in, the user input alternatively pointing to the volume control user interface element 940a corresponding to audio parameter 1, the volume control user interface element 940b corresponding to audio parameter 2, the volume control user interface element 940c corresponding to audio parameter 3, and the selectable option for enabling mobile setting A 942a or mobile setting B 942b. In Fig.9D the illustrated embodiments in, the slider of the volume control user interface element 940a corresponding to the system environment volume or the virtual environment volume decreases in fill level from Fig. 9C in response to the user input corresponding to the request to decrease the slider, the slider of the volume control user interface element 940b corresponding to the application volume level such as the volume level of the user interface of the application in the three-dimensional environment (such as the music application user interface 926a) increases in fill level from Fig. 9C in response to the user input corresponding to the request to increase the slider, the slider of the volume control user interface element 940c corresponding to the volume associated with the avatar decreases in fill level from Fig. 9C in response to the user input corresponding to the request to decrease the slider, and the selectable option for enabling mobile setting A 942a or mobile setting B 942b is set to enable mobile setting B in response to the user input pointing to the selectable option for enabling mobile setting A942a or mobile setting B 942b. Thus, even if some or all of the slider elements 940a - 940c change together and in the same direction in response to the slider input pointing to the slider, such as in Fig. 9A the slider elements 940a - 940c are also capable of changing individually and / or changing in different directions / amounts in response to the input pointing to the individual sliders, such as in Fig. 9CIn addition, the computer system's representative audio point sources for which it generates audio (e.g., groups 930, 931) also (optionally) change in number (and optionally volume) from Fig. 9C 8 in Fig.9D to 3 in

[0271] Fig.9D1 illustrates concepts similar and / or identical to those Fig.9D shown (with many of the same reference numerals). It should be understood that unless otherwise indicated below, elements with the same reference numeral as those 9A to 9E shown have one or more or all of the same characteristics. Fig.9D1 shown Fig.9D1 includes computer system 101, which includes a display generation component 120 (or is the same as the display generation component). In some embodiments, computer system 101 and display generation component 120 respectively have 9A to 9E the characteristics of computer system 101 shown in Figure 3 and display generation component 120 shown in FIGS. 1 and 9A to 9E one or more of the characteristics, and in some embodiments, Fig.9D1 computer system 101 and display generation component 120 shown in

[0272] In Fig.9D1 ...

Claims

1. A method, comprising: at a computer system in communication with a display generation component and one or more input devices: when displaying virtual content at a first immersion level via the display generation component, displaying a system user interface of the computer system via the display generation component, wherein displaying the system user interface includes: displaying an immersion control element configured to control the computer system to the immersion level at which it displays the virtual content; when displaying the virtual content at the first immersion level and displaying the system user interface including the immersion control element, receiving, via the one or more input devices, an input pointing to the immersion control element; and in response to receiving the input pointing to the immersion control element, displaying the virtual content at a second immersion level different from the first immersion level via the display generation component according to the input.

2. The method according to claim 1, wherein: the virtual content is a virtual reality experience in which the physical environment of the display generation component is not visible, when displaying the virtual content at the first immersion level, the virtual content is displayed within an augmented reality experience in which the physical environment of the display generation component is visible, when displaying the virtual content at the first immersion level, the virtual content occupies a first proportion of the augmented reality experience, and in response to receiving the input pointing to the immersion control element, changing, according to the input, the proportion of the augmented reality experience occupied by the virtual content to a second proportion different from the first proportion.

3. The method according to any one of claims 1 to 2, wherein in response to receiving the input pointing to the immersion control element: based on determining that the input corresponds to a request to change the immersion level of the virtual content by a first amount, the virtual content is displayed at a third immersion level, and based on determining that the input corresponds to a request to change the immersion level of the virtual content by a second amount different from the first amount, the virtual content is displayed at a fourth immersion level different from the third immersion level.

4. The method according to any one of claims 1 to 3, wherein the virtual content includes a user interface of an application.

5. The method according to any one of claims 1 to 4, wherein the virtual content includes a first user interface of a first application and a second user interface of a second application different from the first application.

6. The method according to any one of claims 1 to 5, wherein the virtual content includes a system virtual environment.

7. The method according to any one of claims 1 to 6, wherein the virtual content includes a first virtual environment, and the system user interface includes a lighting control element that can be selected to change the lighting settings of the first virtual environment, and the method further includes: when displaying the first virtual environment with the lighting settings having a first value, receiving, via the one or more input devices, a second input pointing to the lighting control element; and In response to receiving the second input, display the first virtual environment via the display generation component according to the second input when the lighting setting has a second value different from the first value.

8. The method according to claim 7, the method further comprising: When displaying the first virtual environment when the lighting setting has a corresponding value, receive a third input corresponding to a request to display a second virtual environment different from the first virtual environment via the one or more input devices; and In response to receiving the third input: According to determining that the corresponding value is the first value, display the second virtual environment when the lighting setting has the first value; and According to determining that the corresponding value is the second value, display the second virtual environment when the lighting setting has the second value.

9. The method according to any one of claims 7 to 8, wherein the system user interface includes an automatic lighting control element that can be selected to automatically set the lighting setting of the first virtual environment at least based on the current date and time of the computer system.

10. The method according to claim 9, wherein the automatic lighting control element can be selected to further automatically set the lighting setting of the first virtual environment based on the type of the corresponding virtual content concurrently displayed with the first virtual environment.

11. The method according to any one of claims 7 to 10, wherein the system user interface includes: A lighting control element that can be selected to change the lighting setting of the first virtual environment, wherein the lighting control element that can be selected to change the lighting setting of the first virtual environment can be selected to set the lighting setting of the first virtual environment to the second value; and A second lighting control element that can be selected to set the lighting setting of the first virtual environment to the first value.

12. The method according to any one of claims 7 to 11, wherein: The lighting control element that can be selected to change the lighting setting of the first virtual environment can be selected to set the lighting setting of the first virtual environment to the second value, and the lighting control element is displayed in an appearance independent of the characteristics of the first virtual environment when the lighting setting has the second value, and The second lighting control element that can be selected to set the lighting setting of the first virtual environment to the first value is displayed in an appearance independent of the characteristics of the first virtual environment when the lighting setting has the first value.

13. The method according to any one of claims 7 to 12, wherein the system user interface includes: The lighting control element that can be selected to change the lighting setting of the first virtual environment, wherein the lighting control element that can be selected to change the lighting setting of the first virtual environment can be selected to set the lighting setting of the first virtual environment to the second value, and wherein the lighting control element that can be selected to set the lighting setting of the first virtual environment to the second value includes a visual representation of the first virtual environment when the lighting setting has the second value; and A second lighting control element that can be selected to set the lighting setting of the first virtual environment to the first value, wherein the second lighting control element that can be selected to set the lighting setting of the first virtual environment to the first value includes a visual representation of the first virtual environment when the lighting setting has the first value.

14. The method according to any one of claims 7 to 13, the method further comprising: When the first virtual environment is not displayed and when the system user interface including the lighting control element is displayed, receiving, via the one or more input devices, a third input pointing to the lighting control element corresponding to a request to change the lighting setting of the first virtual environment from the first value to the second value; and In response to receiving the third input, at least partially displaying, via the display generating component, the first virtual environment when the lighting setting has the second value.

15. The method according to claim 14, the method further comprising: After receiving the third input and when at least partially displaying the first virtual environment when the lighting setting has the second value, automatically stopping displaying the first virtual environment according to a determination that one or more criteria are met.

16. The method according to any one of claims 1 to 15, wherein the system user interface includes a system virtual environment control element that can be selected to initiate a process of changing the current system virtual environment from a first virtual environment to a second virtual environment, the method further comprising: When the system user interface is displayed, receiving, via the one or more input devices, a second input corresponding to a selection of the system virtual environment control element; and In response to receiving the second input, displaying, via the display generating component, a system virtual environment control user interface, the system virtual environment control user interface including one or more selectable options for changing the current system virtual environment from the first virtual environment to the second virtual environment.

17. The method according to claim 16, wherein the system virtual environment control user interface includes one or more selectable options for displaying one or more atmosphere effects on one or more portions of the physical environment visible via the display generating component.

18. The method according to any one of claims 16 to 17, wherein the one or more selectable options for changing the current system virtual environment from the first virtual environment to the second virtual environment include: A first selectable option that can be selected to set the current system virtual environment to the first virtual environment; and A second selectable option that can be selected to set the current system virtual environment to the second virtual environment.

19. The method according to any one of claims 1 to 18, wherein the virtual content is displayed within a three-dimensional environment, and the system user interface includes a first selectable option that can be selected to enable or disable automatic de-emphasis of one or more portions of the three-dimensional environment outside of one or more portions of the virtual content, and the method further includes: When displaying the virtual content: Based on determining that one or more criteria are met and the automatic de-emphasis of one or more portions of the three-dimensional environment outside of one or more portions of the virtual content is enabled, displaying the virtual content via the display generating component at a first visual emphasis level with respect to the one or more portions of the three-dimensional environment outside of the one or more portions of the virtual content; Based on determining that the one or more criteria are met and the automatic de-emphasis of one or more portions of the three-dimensional environment outside of one or more portions of the virtual content is disabled, displaying the virtual content via the display generating component at a second visual emphasis level with respect to the one or more portions of the three-dimensional environment outside of the one or more portions of the virtual content, wherein the second visual emphasis level is less than the first visual emphasis level; and Based on determining that the one or more criteria are not met, displaying the virtual content via the display generating component at the second visual emphasis level with respect to the one or more portions of the three-dimensional environment outside of the one or more portions of the virtual content.

20. The method according to any one of claims 1 to 19, the method further includes: When the system user interface is not displayed and when first virtual content is being displayed via the display generating component, receiving a second input corresponding to a request to display the system user interface via the one or more input devices; In response to receiving the second input, displaying the first virtual content and the system user interface via the display generating component; When the system user interface is not displayed and when second virtual content different from the first virtual content is being displayed via the display generating component, receiving a third input corresponding to a request to display the system user interface via the one or more input devices; and In response to receiving the third input, displaying the second virtual content and the system user interface via the display generating component.

21. The method according to claim 20, wherein the second input and the third input correspond to gaze input.

22. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: When displaying virtual content at a first immersion level via the display generation component, displaying the system user interface of the computer system via the display generation component, wherein displaying the system user interface includes: displaying an immersion control element configured to control the immersion level at which the computer system displays the virtual content; When displaying the virtual content at the first immersion level and displaying the system user interface including the immersion control element, receiving an input pointing to the immersion control element via the one or more input devices; and In response to receiving the input pointing to the immersion control element, displaying the virtual content via the display generation component at a second immersion level different from the first immersion level according to the input.

23. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, cause the computer system to perform a method including: When displaying virtual content at a first immersion level via the display generation component, a system user interface of the computer system is displayed via the display generation component, wherein displaying the system user interface includes: Displaying an immersion control element configured to control the immersion level at which the computer system displays the virtual content; When displaying the virtual content at the first immersion level and displaying the system user interface including the immersion control element, receiving an input pointing to the immersion control element via the one or more input devices; And In response to receiving the input pointing to the immersion control element, displaying the virtual content via the display generation component at a second immersion level different from the first immersion level according to the input.

24. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; Means for, when displaying virtual content at a first immersion level via the display generation component, displaying the system user interface of the computer system via the display generation component, wherein displaying the system user interface includes: displaying an immersion control element configured to control the immersion level at which the computer system displays the virtual content; Means for, when displaying the virtual content at the first immersion level and displaying the system user interface including the immersion control element, receiving an input pointing to the immersion control element via the one or more input devices; and A component for displaying the virtual content at a second immersion level different from the first immersion level via the display generation component according to the input in response to receiving the input pointing to the immersion control element.

25. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any one of the methods according to claims 1 to 21.

26. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 1 to 21.

27. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And A component for performing any one of the methods according to claims 1 to 21.

28. A method comprising: At a computer system communicating with a display generation component and one or more input devices: Displaying a system user interface of the computer system via the display generation component, wherein displaying the system user interface includes: displaying a volume control element; When displaying the system user interface including the volume control element, receiving an input pointing to the volume control element via the one or more input devices; and In response to receiving the input pointing to the volume control element: Adjusting a first volume level associated with a virtual environment associated with the computer system according to the input; and Adjusting a second volume level associated with a user interface of an application associated with the computer system according to the input.

29. The method according to claim 28, the method further comprising: In response to receiving the input pointing to the volume control element, adjusting a third volume level associated with one or more users other than the user of the computer system in a three-dimensional environment associated with the computer system, wherein the three-dimensional environment is shared by the one or more users and the user of the computer system.

30. The method according to any one of claims 28 to 29, wherein the system user interface includes a first volume control element and a second volume control element, and the method further comprises: When displaying the system user interface including the first volume control element and the second volume control element, receiving a second input pointing to the system user interface via the one or more input devices; And In response to receiving the second input: Determine that the second input points to the first volume control element, and adjust a third volume level associated with the virtual environment according to the second input without adjusting the volume level associated with the user interface of the application; and Determine that the second input points to the second volume control element, and adjust a fourth volume level associated with the user interface of the application according to the second input without adjusting the volume level associated with the virtual environment.

31. The method according to any one of claims 28 to 30, wherein adjusting the first volume level associated with the virtual environment comprises: Change the number of audio point sources associated with the virtual environment according to the input.

32. The method according to claim 31, wherein changing the number of audio point sources associated with the virtual environment according to the input includes: When the input corresponds to a request to decrease the first volume level associated with the virtual environment: Determine that the input is a first input for decreasing the first volume level, and reduce the number of audio point sources associated with the virtual environment by removing a first set of audio point sources from the audio associated with the virtual environment; And Determine that the input is a second input for decreasing the first volume level different from the first input, and reduce the number of audio point sources associated with the virtual environment by removing a second set of audio point sources different from the first set of audio point sources from the audio associated with the virtual environment; and When the input corresponds to a request to increase the first volume level associated with the virtual environment: Determine that the input is a first input for increasing the first volume level, and increase the number of audio point sources associated with the virtual environment by adding a third set of audio point sources to the audio associated with the virtual environment; And Determine that the input is a second input for increasing the first volume level different from the first input, and increase the number of audio point sources associated with the virtual environment by adding a fourth set of audio point sources different from the third set of audio point sources to the audio associated with the virtual environment.

33. The method according to any one of claims 31 to 32, the method further comprising: When the volume level of the virtual environment is the first volume level and the computer system is presenting audio associated with the virtual environment including a first number of audio point sources and the virtual environment is displayed at a first immersion level, receive, via the one or more input devices, a second input corresponding to a request to change the immersion level of the virtual environment away from the first immersion level; And In response to receiving the second input, determine that the second input corresponds to a request to display the virtual environment at a second immersion level different from the first immersion level: Display the virtual environment at the second immersion level; and Present the audio associated with the virtual environment at a third volume level different from the first volume level, including presenting the audio associated with the virtual environment including a second number of audio point sources different from the first number of audio point sources.

34. The method according to any one of claims 31 to 33, wherein changing the number of the audio point sources associated with the virtual environment according to the input comprises: Disable at least one audio point source of the audio associated with the virtual environment without disabling at least one audio point source of the audio associated with the virtual environment.

35. The method according to any one of claims 31 to 34, wherein adjusting the first volume level associated with the virtual environment comprises: Adjust one or more audio frequencies associated with one or more audio point sources associated with the virtual environment.

36. The method according to any one of claims 31 to 35, wherein adjusting the first volume level associated with the virtual environment comprises: Adjust the occurrence rate of one or more audio outputs of one or more audio point sources associated with the virtual environment.

37. The method according to any one of claims 28 to 36, wherein adjusting the first volume level associated with the virtual environment comprises: Adjust the volume level of the audio track for the virtual environment.

38. The method according to any one of claims 28 to 37, the method further comprising: When the volume level of the virtual environment is the first volume level and the virtual environment is displayed at a first immersion level, receiving, via the one or more input devices, a second input corresponding to a request to change the immersion level of the virtual environment away from the first immersion level; And In response to receiving the second input: Based on determining that the second input corresponds to a request to display the virtual environment at a second immersion level greater than the first immersion level: Display the virtual environment at the second immersion level; and Present the audio associated with the virtual environment at a third volume level greater than the first volume level; And Based on determining that the second input corresponds to a request to display the virtual environment at a third immersion level less than the first immersion level: Display the virtual environment at the third immersion level; and Present the audio associated with the virtual environment at a fourth volume level less than the first volume level.

39. The method according to any one of claims 28 to 38, wherein: When the input pointing to the volume control element is received, the audio associated with the virtual environment includes one or more audio point sources of a first type and does not include one or more audio point sources of a second type different from the first type, and In response to adjusting the first volume level associated with the virtual environment according to the input, the audio associated with the virtual environment includes one or more audio point sources of the second type.

40. The method according to any one of claims 28 to 39, wherein the system user interface concurrently includes: A first volume control element for adjusting a third volume level associated with the virtual environment, and A selectable option that can be selected to control whether the audio associated with the virtual environment is presented differently by the computer system based on the posture of the head of the user of the computer system.

41. The method according to any one of claims 28 to 40, wherein the system user interface includes a selectable option that can be selected to control whether the audio presented by the computer system is presented differently by the computer system based on the posture of the head of the user of the computer system.

42. The method according to any one of claims 28 to 41, wherein the system user interface includes a selectable option that can be selected to control the amount of noise cancellation performed on the audio presented by the computer system.

43. The method according to any one of claims 28 to 42, the method further comprising: Before displaying the system user interface, receiving, via the one or more input devices, a second input corresponding to a request to display the system user interface; And In response to receiving the second input: Based on determining that the computer system is displaying first content associated with a first application when receiving the second input, displaying the system user interface via the display generation component; And Based on determining that the computer system is displaying second content associated with a second application when receiving the second input, wherein the second content associated with the second application is different from the first content associated with the first application, displaying the system user interface via the display generation component.

44. The method according to claim 43, wherein the second input includes the attention of a user of the computer system.

45. A computer system, the computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: Displaying, via the display generation component, the system user interface of the computer system, wherein displaying the system user interface includes: displaying a volume control element; When displaying the system user interface including the volume control element, receiving, via the one or more input devices, an input pointing to the volume control element; and In response to receiving the input pointing to the volume control element: Adjusting, based on the input, a first volume level associated with a virtual environment associated with the computer system; and Adjusting, based on the input, a second volume level associated with a user interface of an application associated with the computer system.

46. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: Displaying a system user interface of the computer system via the display generation component, wherein displaying the system user interface includes: Displaying a volume control element; When displaying the system user interface including the volume control element, receiving, via the one or more input devices, an input pointing to the volume control element; And In response to receiving the input pointing to the volume control element: Adjusting, based on the input, a first volume level associated with a virtual environment associated with the computer system; and Adjusting, based on the input, a second volume level associated with a user interface of an application associated with the computer system.

47. A computer system, the computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; A component for displaying a system user interface of the computer system via the display generation component, wherein displaying the system user interface includes: displaying a volume control element; A component for receiving, via the one or more input devices, an input pointing to the volume control element when the system user interface including the volume control element is displayed; and A component for performing the following operations in response to receiving the input pointing to the volume control element: Adjusting a first volume level associated with a virtual environment associated with the computer system according to the input; and Adjusting a second volume level associated with a user interface of an application associated with the computer system according to the input.

48. A computer system, the computer system communicating with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one of the methods according to claims 28 to 44.

49. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 28 to 44.

50. A computer system, the computer system communicating with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; And A component for performing any one of the methods according to claims 28 to 44.

51. A method, the method including: At a computer system communicating with a display generation component and one or more input devices: When the computer system is operating in a first operation mode and displaying first virtual content, detecting a first event; In response to detecting the first event during the first operation mode: According to determining that the first event meets one or more first criteria, reducing the salience of at least a portion of the first virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating; And According to determining that the first event does not meet the one or more first criteria, continuing to display the first virtual content without reducing the salience of at least a portion of the first virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating; When the computer system is operating in the first operation mode, displaying a first selectable element via the display generation component; When the computer system is operating in the first operation mode and the first selectable element is being displayed, receiving a first input pointing to the first selectable element via the one or more input devices; In response to receiving the first input directed to the first selectable element, operate the computer system in a second operating mode different from the first operating mode; When operating the computer system in the second operating mode and displaying second virtual content, detect a second event; And In response to detecting the second event, and based on determining that the second event does not meet one or more second criteria different from the one or more first criteria, continue to display the second virtual content without reducing the salience of at least a portion of the second virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating, regardless of whether the second event meets the one or more first criteria.

52. The method according to claim 51, the method further comprising: In response to detecting the second event when operating the computer system in the second operating mode, and based on determining that the second event meets the one or more second criteria, reduce the salience of at least a portion of the second virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating.

53. The method according to any one of claims 51 to 52, wherein reducing the visual salience of the portion of the first virtual content includes: Reveal a portion of the physical environment in which the computer system is operating at a location of the portion of the first virtual content in the three-dimensional environment, wherein the location of the portion of the physical environment in the three-dimensional environment corresponds to the location of the portion of the first virtual content in the three-dimensional environment.

54. The method according to any one of claims 51 to 53, wherein the first event corresponds to one or more actions of a person in the physical environment, and wherein the one or more first criteria include criteria that are met based on the attention of the person.

55. The method according to any one of claims 51 to 54, wherein the one or more first criteria include criteria that are met based on the movement of a user of the computer system in the physical environment.

56. The method according to any one of claims 51 to 55, the method further comprising: When displaying third virtual content via the display generating component, detect a third event; And In response to detecting the third event, based on determining that the third event meets one or more third criteria, including criteria associated with alerting the user to a feature of the physical environment in an area where the user of the computer system may interact, and regardless of whether the computer system is operating in the first operating mode or the second operating mode, reduce the salience of at least a portion of the third virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating, regardless of whether the third event meets the one or more first criteria or the one or more second criteria.

57. The method according to any one of claims 51 to 56, the method further comprising: Detect a notification event; And In response to detecting the notification event: Based on determining that the computer system is operating in the first operating mode, Selectively generate a notification associated with the notification event based on whether one or more third criteria are met, regardless of whether one or more fourth criteria are met; And Based on determining that the computer system is operating in the second operating mode, Selectively generate the notification associated with the notification event based on whether the one or more fourth criteria are met, regardless of whether one or more third criteria are met.

58. The method according to any one of claims 51 to 57, the method further comprising: Detect a third event when displaying third virtual content via the display generating component; And In response to detecting the third event: Based on determining that the computer system is operating in the first operating mode, reduce the saliency of at least a portion of the third virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating; And Based on determining that the computer system is operating in the second operating mode, continue to display the third virtual content without reducing the saliency of at least a portion of the third virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating.

59. The method according to any one of claims 51 to 58, the method further comprising: When the first event meets the one or more first criteria: When the computer system is operating in the first operating mode and when displaying at least a portion of the first virtual content with reduced saliency relative to the representation of at least a portion of the physical environment, receive a second input corresponding to a request to transition from the first operating mode to the second operating mode via the one or more input devices; And In response to receiving the second input, operate the computer system in the second operating mode and increase the saliency of at least a portion of the first virtual content relative to the representation of at least a portion of the physical environment.

60. The method according to any one of claims 51 to 59, the method further comprising: When the first event does not meet the one or more first criteria: When the computer system is operating in the first operating mode and when not displaying at least a portion of the first virtual content with reduced saliency relative to the representation of at least a portion of the physical environment, receive a second input corresponding to a request to transition from the first operating mode to the second operating mode via the one or more input devices; And In response to receiving the second input, operate the computer system in the second operating mode and reduce the saliency of at least a portion of the first virtual content relative to the representation of at least a portion of the physical environment.

61. The method according to any one of claims 51 to 60, wherein during the first operation mode, the computer system operates according to a first set of values of a first set of settings, and the method further includes: When the computer system is operating in a third operation mode and displaying third virtual content, wherein during the third operation mode, the computer system operates according to a second set of values of the first set of settings that are different from the first set of values, detecting a third event; And In response to detecting the third event during the third operation mode: Based on determining that the third event meets the one or more first criteria, reducing the salience of at least a portion of the first virtual content generated by the computer system with respect to a representation of at least a portion of the physical environment in which the computer system is operating; And Based on determining that the third event does not meet the one or more first criteria, continuing to display the first virtual content without reducing the salience of at least a portion of the first virtual content generated by the computer system with respect to a representation of at least a portion of the physical environment in which the computer system is operating.

62. The method according to any one of claims 51 to 61, the method further includes: When the computer system is operating in the first operation mode, receiving, via the one or more input devices, a second input corresponding to a request to forego reducing the visual salience of at least a portion of the first virtual content in response to detecting a corresponding event that meets the one or more first criteria; After receiving the second input, detecting a third event that meets the one or more first criteria; And In response to detecting the third event: Based on determining that the third event is detected after a predetermined period of time since receiving the second input, reducing the salience of at least a portion of the first virtual content generated by the computer system with respect to the representation of at least a portion of the physical environment in which the computer system is operating; And Based on determining that the third event is detected before the predetermined period of time since receiving the second input, foregoing reducing the salience of at least a portion of the first virtual content generated by the computer system with respect to the representation of at least a portion of the physical environment in which the computer system is operating.

63. The method according to any one of claims 51 to 62, the method further includes: When the computer system is operating in the first operation mode, receiving, via the one or more input devices, a second input corresponding to a request to forego reducing the visual salience of at least a portion of the first virtual content in response to detecting a corresponding event that meets the one or more first criteria; After receiving the second input, detecting a third event that meets the one or more first criteria; And In response to detecting the third event: Based on determining that the third event has been detected after a pre-determined event has occurred after receiving the second input, reduce the salience of at least a portion of the first virtual content generated by the computer system with respect to the representation of at least a portion of the physical environment in which the computer system is operating; And Based on determining that the pre-determined event has not occurred since receiving the second input, forgo reducing the salience of at least a portion of the first virtual content generated by the computer system with respect to the representation of at least a portion of the physical environment in which the computer system is operating.

64. The method according to any one of claims 51 to 63, the method further comprising: When the computer system is operating in the first operating mode: When displaying the first virtual content, detect a third event; And In response to detecting the third event: Based on determining that the first setting is enabled and the third event meets one or more third criteria, regardless of whether the third event meets the one or more first criteria, reduce the salience of at least a portion of the first virtual content generated by the computer system with respect to the representation of at least a portion of the physical environment in which the computer system is operating; And Based on determining that the first setting is disabled and the third event meets the one or more first criteria, regardless of whether the third event meets the one or more third criteria, reduce the salience of at least a portion of the first virtual content generated by the computer system with respect to the representation of at least a portion of the physical environment in which the computer system is operating.

65. The method according to any one of claims 51 to 64, wherein the computer system generates an audio output when displaying the first virtual content, the method further comprising: In response to detecting the first event and based on determining that the first event meets the one or more first criteria, reduce the salience of at least a portion of the audio output generated by the computer system.

66. The method according to any one of claims 51 to 65, the method further comprising: Before displaying the first selectable element, receive a second input corresponding to a request to display the first selectable element via the one or more input devices; And In response to receiving the second input: Based on determining that the computer system is displaying first content associated with a first application when receiving the second input, display a system user interface including the first selectable element via the display generation component; And Based on determining that the computer system is displaying second content associated with a second application when receiving the second input, wherein the second content associated with the second application is different from the first content associated with the first application, display the system user interface including the first selectable element via the display generation component.

67. The method according to claim 66, wherein the second input includes the attention of a user of the computer system.

68. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: When the computer system is operating in a first operation mode and displaying first virtual content, detecting a first event; In response to detecting the first event during the first operation mode: Based on determining that the first event meets one or more first criteria, reducing the salience of at least a portion of the first virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating; And Based on determining that the first event does not meet the one or more first criteria, continuing to display the first virtual content without reducing the salience of at least a portion of the first virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating; When the computer system is operating in the first operation mode, displaying a first selectable element via the display generation component; When the computer system is operating in the first operation mode and the first selectable element is being displayed, receiving a first input pointing to the first selectable element via the one or more input devices; In response to receiving the first input pointing to the first selectable element, operating the computer system in a second operation mode different from the first operation mode; When the computer system is operating in the second operation mode and displaying second virtual content, detecting a second event; And In response to detecting the second event, and based on determining that the second event does not meet one or more second criteria different from the one or more first criteria, continuing to display the second virtual content without reducing the salience of at least a portion of the second virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating, regardless of whether the second event meets the one or more first criteria.

69. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, cause the computer system to perform a method including: When the computer system is operating in a first operation mode and displaying first virtual content, detecting a first event; In response to detecting the first event during the first operation mode: Based on determining that the first event satisfies one or more first criteria, reduce the salience of at least a portion of the first virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating; And Based on determining that the first event does not satisfy the one or more first criteria, continue to display the first virtual content without reducing the salience of at least a portion of the first virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating; When the computer system is operating in the first operation mode, display a first selectable element via the display generation component; When the computer system is operating in the first operation mode and the first selectable element is being displayed, receive a first input pointing to the first selectable element via the one or more input devices; In response to receiving the first input pointing to the first selectable element, operate the computer system in a second operation mode different from the first operation mode; When the computer system is operating in the second operation mode and displaying second virtual content, detect a second event; And In response to detecting the second event, and based on determining that the second event does not satisfy one or more second criteria different from the one or more first criteria, continue to display the second virtual content without reducing the salience of at least a portion of the second virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating, regardless of whether the second event satisfies the one or more first criteria.

70. A computer system, the computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; A component for detecting a first event when the computer system is operating in a first operation mode and displaying first virtual content; A component for performing the following operations in response to detecting the first event during the first operation mode: Based on determining that the first event satisfies one or more first criteria, reduce the salience of at least a portion of the first virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating; And Based on determining that the first event does not satisfy the one or more first criteria, continue to display the first virtual content without reducing the salience of at least a portion of the first virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating; A component for displaying a first selectable element via the display generation component when the computer system is operating in the first operation mode; means for receiving, via the one or more input devices, a first input directed to the first selectable element when the computer system is operating in the first operating mode and the first selectable element is being displayed; means for operating the computer system in a second operating mode different from the first operating mode in response to receiving the first input directed to the first selectable element; means for detecting a second event when the computer system is operating in the second operating mode and second virtual content is being displayed; and means for, in response to detecting the second event and in accordance with determining that the second event does not meet one or more second criteria different from the one or more first criteria, performing the following: continuing to display the second virtual content without reducing the salience of at least a portion of the second virtual content generated by the computer system relative to a representation of at least a portion of the physical environment in which the computer system is operating, regardless of whether the second event meets the one or more first criteria.

71. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods according to claims 51 to 67.

72. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, cause the computer system to perform any of the methods according to claims 51 to 67.

73. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and means for performing any of the methods according to claims 51 to 67.

74. A method, the method comprising: at a first computer system that communicates with a display generation component and one or more input devices: displaying, via the display generation component, a first selectable option in a three-dimensional environment that can be selected to display a representation of content from the second computer system when the second computer system is visible in the three-dimensional environment via the display generation component; receiving, via the one or more input devices, an input directed to the first selectable option when the first selectable option is being displayed in the three-dimensional environment; and in response to receiving the input directed to the first selectable option, displaying, via the display generation component and in the three-dimensional environment, the representation of the content from the second computer system. When displaying the representation of the content from the second computer system in the three-dimensional environment, detect, via the one or more input devices, one or more inputs pointing to the representation of the content from the second computer system; and In response to detecting the one or more inputs pointing to the representation of the content from the second computer system, perform one or more operations corresponding to the one or more inputs with respect to the content from the second computer system.

75. The method according to claim 74, the method comprising: In response to receiving the input pointing to the first selectable option, generate, via the display component, an animation including the conversion of the first selectable option into the representation of the content from the second computer system.

76. The method according to claim 74 or 75, wherein the first selectable option is displayed at a corresponding position in the three-dimensional environment with respect to the second computer system.

77. The method according to any one of claims 74 to 76, wherein displaying the representation of the content from the second computer system comprises: Display the representation of the content from the second computer system overlying the second computer system in the three-dimensional environment from the viewpoint of the user of the first computer system.

78. The method according to any one of claims 74 to 77, wherein the first selectable option is displayed in the first system user interface of the first computer system or the second computer system, and wherein the first system user interface is accessible when any user interface in the first plurality of user interfaces of the first plurality of applications is displayed by the first computer system or the second computer system.

79. The method according to any one of claims 74 to 78, wherein displaying the representation of the content from the second computer system comprises: Display the representation of the content from the second computer system at a first distance from the viewpoint of the user of the first computer system in the three-dimensional environment, wherein the first distance is independent of the distance from the viewpoint of the user to the second computer system in the three-dimensional environment.

80. The method according to any one of claims 74 to 79, wherein displaying the representation of the content from the second computer system comprises: Display the representation of the content from the second computer system at a first position with respect to the system user interface including the first selectable option.

81. The method according to any one of claims 74 to 80, wherein the representation showing the content from the second computer system comprises: Display the representation of the content from the second computer system behind the system user interface in the three-dimensional environment from the viewpoint of the user of the first computer system.

82. The method according to claim 81, the method comprising: When presenting a representation of content from the second computer system behind the system user interface as viewed from the user's perspective in the three-dimensional environment, user input including the user's attention directed to the representation of the content from the second computer system is received via the one or more input devices; and In response to receiving the user input including the user's attention, the representation of the content from the second computer system is presented in front of the system user interface as viewed from the user's perspective.

83. The method according to any one of claims 74 to 82, wherein in response to receiving the input directed to the first selectable option: Based on determining that the operating state of the second computer system is a first state of the second computer system when the input directed to the first selectable option is received, the operating state indicated by the representation of the content from the second computer system is a first state of the representation of the content from the second computer system based on the first state of the second computer system, and Based on determining that the operating state of the second computer system is a second state of the second computer system different from the first state of the second computer system when the input directed to the first selectable option is received, the operating state indicated by the representation of the content from the second computer system is a second state of the representation of the content from the second computer system different from the first state of the representation of the content from the second computer system, wherein the second state of the representation of the content from the second computer system is based on the second state of the second computer system.

84. The method according to any one of claims 74 to 83, wherein before detecting the one or more inputs directed to the representation of the content from the second computer system, the operating state of the second computer system is a first state, and wherein the method includes: In response to performing the one or more operations corresponding to the one or more inputs with respect to the content from the second computer system, changing the operating state of the second computer system to a second state different from the first state based on the one or more inputs directed to the representation of the content from the second computer system.

85. The method according to claim 84, wherein the first state corresponds to a first set of application windows being active, and wherein the second state corresponds to a second set of application windows different from the first set of applications being active.

86. The method according to claim 84 or 85, wherein the first state corresponds to a first set of application windows being active and in a first order, and wherein the second state corresponds to the first set of application windows being active and in a second order different from the first order.

87. The method according to any one of claims 84 to 86, wherein the first state corresponds to the application window being active and including a first visual characteristic, and the second state corresponds to the application window being active and including a second visual characteristic different from the first visual characteristic.

88. The method according to any one of claims 74 to 87, the method comprising: When displaying the first selectable option in the three-dimensional environment, display, via the display generation component, a second selectable option in the three-dimensional environment that can be selected to display content from a third computer system different from the second computer system.

89. The method according to claim 88, wherein: the first selectable option is displayed based on determining that the second computer system is within a threshold distance of the first computer system, and the second selectable option is displayed based on determining that the third computer system is within the threshold distance of the first computer system.

90. The method according to claim 88 or 89, the method comprising: receiving, via the one or more input devices, an input pointing to the second selectable option that can be selected to display content from the third computer system; and in response to receiving the input pointing to the second selectable option: based on determining that the representation of the content from the second computer system is being displayed when the input pointing to the second selectable option is received: stop displaying the representation of the content from the second computer system on the display generation component; and display, via the display generation component and in the three-dimensional environment, the second representation of the content from the third computer system.

91. The method according to any one of claims 88 to 90, the method comprising: receiving, via the one or more input devices, a second input pointing to the first selectable option that can be selected to display the representation of the content from the second computer system; and in response to receiving the second input pointing to the first selectable option: based on determining that the representation of the content from the second computer system is being displayed when the second input pointing to the first selectable option is received, stop displaying the representation of the content from the second computer system on the display generation component.

92. The method according to any one of claims 74 to 91, wherein: the representation of the content from the second computer system is bent around the user's viewpoint of the first computer system in the three-dimensional environment.

93. The method according to any one of claims 74 to 92, wherein: the representation of the content from the second computer system includes background content corresponding to the background content displayed by the second computer system when the input pointing to the first selectable option is received.

94. The method according to any one of claims 74 to 93, wherein: the representation of the content from the second computer system is displayed on a simulated glass material in the three-dimensional environment.

95. The method according to any one of claims 74 to 94, the method comprising: When displaying the representation of the content from the second computer system and while the second computer system is displaying content, detecting, via the one or more input devices, a change in the state of the second computer system, including a change in the state of the display device of the second computer system; And after detecting the change in the state of the second computer system, continuing to display, in the three-dimensional environment, the representation of the content from the second computer system.

96. The method according to any one of claims 74 to 95, the method further comprising: When displaying the representation of the content from the second computer system in the three-dimensional environment, receiving, via the one or more input devices, a first input corresponding to a request to transfer corresponding content from the first computer system to the second computer system; And In response to receiving the first input, initiating a process of transferring the corresponding content from the first computer system to the second computer system.

97. The method according to any one of claims 74 to 96, the method comprising: When displaying the representation of the content from the second computer system in the three-dimensional environment, receiving, via one or more input devices of the second computer system, a second input corresponding to a request to transfer corresponding content from the second computer system to the first computer system; And In response to receiving the second input, initiating a process of transferring the corresponding content from the second computer system to the first computer system.

98. A computer system, the computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: When a second computer system is visible in a three-dimensional environment via the display generation component, displaying, via the display generation component, in the three-dimensional environment a first selectable option that can be selected to display a representation of content from the second computer system; When displaying the first selectable option in the three-dimensional environment, receiving, via the one or more input devices, an input pointing to the first selectable option; And In response to receiving the input pointing to the first selectable option, displaying, via the display generation component and in the three-dimensional environment, the representation of the content from the second computer system; When displaying the representation of the content from the second computer system in the three-dimensional environment, detecting, via the one or more input devices, one or more inputs pointing to the representation of the content from the second computer system; And In response to detecting the one or more inputs pointing to the representation of the content from the second computer system, one or more operations corresponding to the one or more inputs are performed with respect to the content from the second computer system.

99. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: When a second computer system is visible in a three-dimensional environment via the display generation component, display, via the display generation component, in the three-dimensional environment a first selectable option that can be selected to display a representation of content from the second computer system; When the first selectable option is displayed in the three-dimensional environment, receive, via the one or more input devices, an input pointing to the first selectable option; And In response to receiving the input pointing to the first selectable option, display, via the display generation component and in the three-dimensional environment, the representation of the content from the second computer system; When the representation of the content from the second computer system is displayed in the three-dimensional environment, detect, via the one or more input devices, one or more inputs pointing to the representation of the content from the second computer system; And In response to detecting the one or more inputs pointing to the representation of the content from the second computer system, perform one or more operations corresponding to the one or more inputs with respect to the content from the second computer system.

100. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; Means for displaying, via the display generation component and in a three-dimensional environment, a first selectable option that can be selected to display a representation of content from a second computer system when the second computer system is visible in the three-dimensional environment via the display generation component; Means for receiving, via the one or more input devices, an input pointing to the first selectable option when the first selectable option is displayed in the three-dimensional environment; And Means for displaying, in response to receiving the input pointing to the first selectable option, via the display generation component and in the three-dimensional environment, the representation of the content from the second computer system; Means for detecting, via the one or more input devices, one or more inputs pointing to the representation of the content from the second computer system when the representation of the content from the second computer system is displayed in the three-dimensional environment; And Means for performing one or more operations corresponding to the one or more inputs with respect to the content from the second computer system in response to detecting the one or more inputs pointing to the representation of the content from the second computer system.

101. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one of the methods according to claims 74 to 97.

102. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 74 to 97.

103. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And Means for performing any one of the methods according to claims 74 to 97.

104. A method, the method comprising: At a first computer system that communicates with a display generation component and one or more input devices: When a second computer system is displaying a first user interface including first content and the first computer system is not displaying a representation of content from the second computer system, receiving, via the one or more input devices, a first input corresponding to a request to display, via the display generation component of the first computer system, a representation of content from the second computer system; And In response to receiving the first input corresponding to the request to display, via the display generation component of the first computer system, a representation of content from the second computer system, initiating the following processes: Displaying, via the display generation component of the first computer system, a representation of the first content from the second computer system; And De-emphasizing the first content in the first user interface displayed by the second computer system.

105. The method according to claim 104, wherein initiating the process of de-emphasizing the first content in the first user interface displayed by the second computer system comprises: Initiating a process to stop displaying the first content in the first user interface displayed by the second computer system.

106. The method according to claim 104 or 105, wherein initiating the process of de-emphasizing the first content in the first user interface displayed by the second computer system includes: Initiating a process to change a value of a visual characteristic by which the second computer system displays the first content in the first user interface.

107. The method according to any one of claims 104 to 106, wherein initiating the process of de-emphasizing the first content in the first user interface displayed by the second computer system includes: Initiating a process to obscure the first content in the first user interface displayed by the second computer system with second content.

108. The method according to any one of claims 104 to 107, wherein the second computer system displays the first user interface including the first content via a second display generation component communicating with the second computer system, and wherein initiating the process of de-emphasizing the first content in the first user interface displayed by the second computer system comprises: Initiating a process to stop displaying any content on the second display generation component of the second computer system.

109. The method according to any one of claims 104 to 108, the method further comprising: In response to receiving the first input corresponding to the request to display, via the display generation component of the first computer system, a representation of content from the second computer system, initiating a process to display, via a second display generation component of the second computer system, placeholder content different from the first content in the first user interface.

110. The method according to any one of claims 104 to 109, wherein the first input corresponding to the request to display a representation of the content from the second computer system via the display generation component of the first computer system comprises: Detection of the placement of the display generation component on a part of the user of the first computer system.

111. The method according to claim 110, the method comprising: When it is detected that the display generation component is placed on the part of the user and when the display generation component is displaying the representation of the first content from the second computer system, wherein the representation of the first content is in a first state, detecting that the display generation component is no longer placed on the part of the user; And In response to detecting that the display generation component is no longer placed on the part of the user, initiating a process of causing the first content to be displayed in the first state on the second display generation component of the second computer system.

112. The method according to any one of claims 104 to 109, wherein the first input corresponding to the request to display the representation of the content from the second computer system via the display generation component of the first computer system is received when it is detected that the display generation component is placed on a predefined part of the user of the first computer system.

113. The method according to claim 112, the method comprising: Before receiving the first input, displaying, via the display generation component of the first computer system, a user interface element that can be selected to initiate the process of displaying the representation of the first content from the second computer system via the display generation component of the first computer system, wherein the first input corresponding to the request to display the representation of the content from the second computer system via the display generation component of the first computer system points to the user interface element, and wherein the user interface element is displayed at a position in the display area of the display generation component, the position being based on the relative position of the display generation component with respect to the second computer system; In response to receiving the first input that points to the user interface element and corresponds to the request to display the representation of the content from the second computer system via the display generation component of the first computer system, stopping the display of the user interface element displayed via the display generation component of the first computer system; And In response to receiving a second input corresponding to the request to stop displaying the representation of the first content from the second computer system via the display generation component of the first computer system, redisplaying the user interface element at a position in the display area of the display generation component via the display generation component of the first computer system, the position being based on the relative position of the display generation component with respect to the second computer system.

114. The method according to any one of claims 104 to 113, wherein the first input corresponding to the request to display the representation of the content is received by the first computer system from the second computer system.

115. The method according to claim 114, wherein in response to the first computer system receiving the first input corresponding to the request for the representation of displaying the content from the second computer system via the display generation component of the first computer system, a user interface including a representation of a user instruction associated with activating the first computer system is displayed via a second display generation component of the second computer system.

116. The method according to claim 114 or 115, wherein when displaying the representation of the first content from the second computer system via the display generation component of the first computer system in response to receiving the first input from the second computer system, a user interface element that can be selected to stop displaying the representation of the first content from the second computer system on the display generation component of the first computer system is displayed via a second display generation component of the second computer system.

117. The method according to any one of claims 104 to 116, wherein the representation of the first content is displayed via the display generation component of the first computer system at a predetermined resolution.

118. The method according to any one of claims 104 to 117, the method comprising: displaying, via the display generation component of the first computer system, the representation of the first content from the second computer system in a first size and at a first resolution; and in response to receiving a second input corresponding to a request to change the size of the representation of the first content from the second computer system, displaying the representation of the first content from the second computer system at the first resolution and in a second size different from the first size.

119. The method according to any one of claims 104 to 118, the method comprising: displaying, via the display generation component of the first computer system, the representation of the first content from the second computer system at a first location in the three-dimensional environment; when displaying the representation of the first content at the first location in the three-dimensional environment, receiving, via the one or more input devices, a second input corresponding to a request to move the representation of the first content from the second computer system from the first location to a second location different from the first location in the three-dimensional environment; and in response to receiving the second input corresponding to the request to move the representation of the first content from the second computer system, moving the representation of the first content from the second computer system from the first location to the second location in the three-dimensional environment.

120. The method according to claim 119, wherein displaying, via the display generation component of the first computer system, a representation of the first content from the second computer system in response to receiving the first input corresponding to the request for displaying the representation of the content from the second computer system includes: Determining that the second computer system is at a first position relative to the field of view as viewed from the user's viewpoint, and displaying, at a second position relative to the field of view as viewed from the user's viewpoint, a representation of the first content from the second computer system, wherein the second position has a predetermined spatial relationship with respect to the first position; And Determining that the second computer system is at a third position different from the first position relative to the field of view as viewed from the user's viewpoint, and displaying, at a fourth position different from the second position relative to the field of view as viewed from the user's viewpoint, a representation of the first content from the second computer system, wherein the fourth position has the predetermined spatial relationship with respect to the second position.

121. The method according to any one of claims 104 to 120, wherein the size of the representation of the first content from the second computer system displayed via the display generation component of the first computer system is larger than the display area of the second display generation component of the second computer system.

122. The method according to claim 121, wherein: The first user interface including the first content includes: A first application window; and A second application window, the second application window being displayed at a first position relative to the first application window in the first user interface; and The process of initiating the display of the representation of the first content from the second computer system via the display generation component of the first computer system includes: Initiating a process of displaying the representation of the second application window to a position different from the first position in the display of the representation of the first content from the second computer system relative to the representation of the first application window.

123. The method according to any one of claims 104 to 122, wherein when an event corresponding to terminating the display of the representation of the first content from the second computer system via the display generation component of the first computer system is detected, the representation of the first content from the second computer system is in a first state; and In response to detecting the event corresponding to terminating the display of the representation of the first content from the second computer system, initiating a process of displaying the first content in the first state via the second display generation component of the second computer system.

124. The method according to claim 123, wherein detecting the event corresponding to terminating the display of the representation of the first content from the second computer system by the display generation component of the first computer system comprises: Detecting that the display generation component is no longer placed on a predefined portion of the user of the first computer system.

125. The method according to claim 123 or 124, wherein detecting the event corresponding to terminating the display of the representation of the first content from the second computer system by the display generating component of the first computer system includes: Detect a user selection of a user interface element that is concurrently displayed with the representation of the first content from the second computer system, wherein the user interface element can be selected to initiate a process of stopping the display of the representation of the first content from the second computer system via the display generation component of the first computer system.

126. The method according to any one of claims 123 to 125, wherein: When the first input is received, the state of the representation of the first content from the second computer system is a second state different from the first state; and When the representation of the first content from the second computer system is displayed in the second state, one or more operations corresponding to one or more inputs pointing to the representation of the first content from the second computer system are performed on the first content from the second computer system, such that in response to performing the one or more operations, the representation of the first content from the second computer system is in the first state.

127. The method according to any one of claims 123 to 126, wherein displaying the representation of the first content from the second computer system in the first state includes displaying: A first representation of a first application window of the representation of the first content; and A second representation of a second application window of the representation of the first content, wherein the first representation of the first application window is displayed via the display generation component of the first computer system in a first relative placement with respect to the second representation of the second application window, and wherein when the first content is displayed in the first state via the second display generation component of the second computer system, the first relative placement of the first application window with respect to the second application window is maintained.

128. The method according to any one of claims 104 to 127, wherein before the first input is received, a second display generation component of the second computer system is displaying a user interface object including one or more selectable options for displaying one or more application user interfaces, and wherein the method includes: In response to receiving the first input, displaying via the display generation component of the first computer system: A representation of the user interface object, the representation of the user interface object including representations of the one or more selectable options for displaying the one or more representations of the application user interfaces, the representation of the user interface object being displayed visually separately from the representation of the first content from the second computer system.

129. A computer system, the computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; Memory; And One or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for the following operations: When a second computer system is displaying a first user interface including first content and the first computer system is not displaying a representation of content from the second computer system, receive, via the one or more input devices, a first input corresponding to a request to display, via the display generating component of the first computer system, a representation of content from the second computer system; And In response to receiving the first input corresponding to the request to display, via the display generating component of the first computer system, a representation of content from the second computer system, initiate the following process: Display, via the display generating component of the first computer system, a representation of the first content from the second computer system; And De-emphasize the first content in the first user interface displayed by the second computer system.

130. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generating component and one or more input devices, cause the computer system to perform a method including the following operations: While a second computer system is displaying a first user interface that includes first content and while the first computer system is not displaying a representation of content from the second computer system, receive a first input via the one or more input devices that corresponds to a request to display, via a display generation component of the first computer system, a representation of content from the second computer system; And In response to receiving the first input corresponding to the request to display, via the display generating component of the first computer system, a representation of content from the second computer system, initiate the following process: Display, via the display generating component of the first computer system, a representation of the first content from the second computer system; And De-emphasize the first content in the first user interface displayed by the second computer system.

131. A computer system in communication with a display generating component and one or more input devices, the computer system including: One or more processors; A memory; Means for receiving, via the one or more input devices, a first input corresponding to a request to display, via the display generating component of the first computer system, a representation of content from the second computer system when a second computer system is displaying a first user interface including first content and the first computer system is not displaying a representation of content from the second computer system; And Means for, in response to receiving the first input corresponding to the request to display, via the display generating component of the first computer system, a representation of content from the second computer system, initiate the following process: Display, via the display generating component of the first computer system, a representation of the first content from the second computer system; And De-emphasize the first content in the first user interface displayed by the second computer system.

132. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one of the methods according to claims 104 to 128.

133. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 104 to 128.

134. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And Components for performing any one of the methods according to claims 104 to 128.

135. A method, the method comprising: At a first computer system that communicates with a display generation component, one or more input devices, and one or more cameras: Visually detecting, via the one or more cameras, a second computer system in a physical environment corresponding to a three-dimensional environment visible via the display generation component; and In response to visually detecting the second computer system: Displaying, in the three-dimensional environment, a first selectable option that can be selected to initiate a process of establishing a connection between the first computer system and the second computer system, based on determining that the second computer system meets one or more connection criteria; And Abstaining from displaying the first selectable option in the three-dimensional environment, based on determining that the second computer system does not meet the one or more connection criteria.

136. The method according to claim 135, wherein satisfaction of the one or more connection criteria is based on a wireless connection between the first computer system and the second computer system.

137. The method according to any one of claims 135 to 136, wherein satisfaction of the one or more connection criteria is based on determining that the first computer system and the second computer system are associated with the same user account.

138. The method according to any one of claims 135 to 137, wherein satisfaction of the one or more connection criteria is based on determining that the second computer system is in a wake state.

139. The method according to any one of claims 135 to 138, wherein satisfaction of the one or more connection criteria is based on determining that the second computer system is in an unlocked state.

140. The method according to any one of claims 135 to 139, wherein: The physical environment of the first computer system includes a plurality of computer systems, the plurality of computer systems including the second computer system; and Satisfaction of the one or more connection criteria is based on determining that the second computer system is the computer system closest to the first computer system among the multiple computer systems.

141. The method according to claim 140, wherein determining that the second computer system is the closest computer system among the multiple computer systems is based on the strength of a wireless signal transmitted by the second computer system.

142. The method according to any one of claims 135 to 141, wherein satisfaction of the one or more connection criteria is based on determining that the first computer system has not detected a corresponding event for stopping the display of the first selectable option.

143. The method according to any one of claims 135 to 142, wherein the physical environment of the first computer system includes a third computer system different from the second computer system, and the method further includes: In response to visually detecting the second computer system and the third computer system: Based on determining that the second computer system satisfies the one or more connection criteria and the third computer system satisfies the one or more connection criteria, concurrently display via the display generating component: The first selectable option; And A second selectable option that can be selected to initiate a process of establishing a connection between the first computer system and the third computer system.

144. The method according to any one of claims 135 to 143, wherein the first selectable option is displayed at a first position in the three-dimensional environment, and the first position is within a threshold distance of a part of the second computer system.

145. The method according to claim 144, the method further includes: When displaying the first selectable option based on determining that the second computer system satisfies the one or more connection criteria in response to visually detecting the second computer system, detecting the movement of the second computer system in the physical environment; And In response to detecting the movement: Stop displaying the first selectable option in the three-dimensional environment.

146. The method according to claim 144, the method further includes: When displaying the first selectable option based on determining that the second computer system satisfies the one or more connection criteria in response to visually detecting the second computer system, detecting the movement of the second computer system in the physical environment; And In response to detecting the movement: Display the first selectable option at a second position different from the first position in the three-dimensional environment via the display generating component, and the second position is within a threshold distance of the part of the second computer system.

147. The method according to any one of claims 135 to 146, wherein the second computer system communicates with a second display generating component different from the display generating component.

148. The method according to any one of claims 135 to 147, wherein the second computer system communicates with a keyboard.

149. The method according to any one of claims 135 to 148, wherein the second computer system communicates with a touchpad.

150. The method according to any one of claims 135 to 149, wherein the physical environment of the first computer system includes a third computer system different from the second computer system, and the method further includes: When displaying the first selectable option in response to visually detecting the second computer system and determining that the second computer system meets the one or more connection criteria, detecting, via the one or more input devices, an input corresponding to a selection of the first selectable option; And In response to detecting the input: Providing an indication of a request for the user of the first computer system to provide disambiguation input.

151. The method according to claim 150, wherein: The second computer system communicates with one or more second input devices different from the one or more input devices, the one or more second input devices including one or more physical buttons; and the disambiguation input includes a selection of a first button among the one or more physical buttons.

152. The method according to any one of claims 150 to 151, wherein: The second computer system communicates with one or more second input devices different from the one or more input devices, the one or more second input devices including a touch-sensitive surface; and The disambiguation input includes an input of a contact object on the touch-sensitive surface.

153. The method according to any one of claims 150 to 152, wherein providing the indication of the request for the user of the first computer system to provide the disambiguation input comprises: Displaying, via the display generation component, a visual prompt for the request.

154. The method according to any one of claims 135 to 153, wherein the physical environment of the first computer system includes a first input device different from the one or more input devices, and the method further includes: In response to visually detecting the first input device: According to determining that the first input device meets one or more second connection criteria, displaying, via the display generation component, a second selectable option that can be selected to initiate a process of establishing a connection between the first computer system and the first input device.

155. The method according to claim 154, the method further includes: When displaying the second selectable option in response to visually detecting the first input device and determining that the first input device meets the one or more second connection criteria, detecting, via the one or more input devices, an input corresponding to a selection of the second selectable option; And In response to detecting the input: Initiating a process of establishing the connection between the first computer system and the first input device.

156. The method according to claim 155, the method further includes: In response to detecting the input: Displaying, via the display generation component, a visual animation indicating the connection at one or more positions within a threshold distance of a part of the first input device in the three-dimensional environment.

157. The method according to any one of claims 135 to 156, the method further includes: When displaying the first selectable option in response to visually detecting the second computer system and determining that the second computer system meets the one or more connection criteria, detect an input corresponding to the selection of the first selectable option via the one or more input devices; And In response to detecting the input: Establish the connection between the first computer system and the second computer system, including displaying a representation of content from the second computer system in the three-dimensional environment via the display generation component.

158. The method according to claim 157, wherein: The second computer system communicates with a second display generation component different from the display generation component; And Displaying the representation of content from the second computer system in the three-dimensional environment includes: displaying the representation of content from the second computer system at a position a certain distance behind the second display generation component in the three-dimensional environment.

159. The method according to any one of claims 157 to 158, wherein according to determining that the first selectable option is displayed in the system user interface when the input is detected, displaying the representation of content from the second computer system in the three-dimensional environment includes: Displaying the representation of content from the second computer system at a position based on the position of the viewpoint of the user of the first computer system.

160. A computer system, the computer system communicates with a display generation component and one or more input devices, the computer system includes: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs include instructions for the following operations: Visually detect a second computer system in a physical environment corresponding to a three-dimensional environment visible via the display generation component via one or more cameras; and In response to visually detecting the second computer system: According to determining that the second computer system meets one or more connection criteria, display a first selectable option in the three-dimensional environment that can be selected to initiate a process of establishing a connection between the first computer system and the second computer system; And According to determining that the second computer system does not meet the one or more connection criteria, refrain from displaying the first selectable option in the three-dimensional environment.

161. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs include instructions that, when executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: Visually detect a second computer system in a physical environment corresponding to a three-dimensional environment visible via the display generation component via one or more cameras; and In response to visually detecting the second computer system: Based on determining that the second computer system meets one or more connection criteria, display a first selectable option in the three-dimensional environment that can be selected to initiate a process of establishing a connection between the first computer system and the second computer system; and Based on determining that the second computer system does not meet the one or more connection criteria, refrain from displaying the first selectable option in the three-dimensional environment.

162. A computer system, the computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; A component for visually detecting, via one or more cameras, a second computer system in a physical environment corresponding to a three-dimensional environment visible via the display generation component; And A component for performing the following operations in response to visually detecting the second computer system: Based on determining that the second computer system meets one or more connection criteria, display a first selectable option in the three-dimensional environment that can be selected to initiate a process of establishing a connection between the first computer system and the second computer system; And Based on determining that the second computer system does not meet the one or more connection criteria, refrain from displaying the first selectable option in the three-dimensional environment.

163. A computer system, the computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one of the methods according to claims 135 to 159.

164. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 135 to 159.

165. A computer system, the computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And A component for performing any one of the methods according to claims 135 to 159.

166. A method, the method comprising: At a first computer system communicating with a display generation component and one or more input devices and one or more cameras: Detect, via the one or more input devices, a request to establish a connection with a corresponding computer system different from the first computer system, the corresponding computer system being within a corresponding area of the physical environment of the first computer system; and In response to detecting the request: Establish a connection between the first computer system and the second computer system among the multiple computer systems when it is determined that the second computer system among the multiple computer systems meets one or more criteria while the multiple computer systems are within the corresponding area, and do not establish a connection between the first computer system and other computer systems among the multiple computer systems; and Establish a connection between the first computer system and the third computer system among the multiple computer systems when it is determined that the third computer system different from the second computer system among the multiple computer systems meets the one or more criteria while the multiple computer systems are within the corresponding area, and do not establish a connection between the first computer system and other computer systems among the multiple computer systems including the second computer system.

167. The method according to claim 166, the method further comprising: In response to detecting the request: Establish the connection between the first computer system and the second computer system regardless of whether the second computer system meets the one or more criteria based on determining that the corresponding area includes the second computer system and does not include other computer systems.

168. The method according to any one of claims 166 to 167, wherein: Establishing the connection between the first computer system and the second computer system and not establishing the connection between the first computer system and other computer systems among the multiple computer systems includes: displaying a representation of content from the second computer system in the three-dimensional environment via the display generation component without displaying a representation of content from other computer systems among the multiple computer systems; and Establishing the connection between the first computer system and the third computer system and not establishing a connection between the first computer system and other computer systems among the multiple computer systems including the second computer system includes: displaying a representation of content from the third computer system in the three-dimensional environment without displaying a representation of content from other computer systems among the multiple computer systems, including the representation of content from the second computer system.

169. The method according to any one of claims 166 to 168, wherein the corresponding area further includes one or more corresponding input devices, and the method further comprises: In response to detecting the request: Establish a connection between the first computer system and the first input device among the one or more corresponding input devices when it is determined that the first input device among the one or more corresponding input devices meets one or more second criteria while the one or more corresponding input devices are within the corresponding area, including configuring the first input device to operate as an input device for the first computer system.

170. The method according to any one of claims 166 to 169, wherein: Determining that the second computer system meets the one or more criteria is based on detecting an indication that the second computer system has detected a selection of a first button associated with the second computer system; and Determining that the third computer system meets the one or more criteria is based on detecting an indication that the third computer system has detected a selection of a second button associated with the third computer system.

171. The method according to any one of claims 166 to 170, wherein: Determining that the second computer system meets the one or more criteria is based on detecting a first audio output from the second computer system; and Determining that the third computer system meets the one or more criteria is based on detecting an audio output from the third computer system.

172. The method according to any one of claims 166 to 171, wherein: Determining that the second computer system meets the one or more criteria is based on detecting an indication of a first image captured by the second computer system; and determining that the third computer system meets the one or more criteria is based on detecting an indication of a second image captured by the third computer system.

173. The method according to any one of claims 166 to 172, wherein: Determining that the second computer system meets the one or more criteria is based on visually detecting, via the one or more cameras, a first content displayed by the second computer system; and Determining that the third computer system meets the one or more criteria is based on visually detecting a second content displayed by the third computer system.

174. The method according to claim 173, wherein: The first content is displayed by the second computer system when the second computer system detects first data sent by the first computer system; and The second content is displayed by the third computer system when the third computer system detects second data sent by the first computer system.

175. The method according to any one of claims 173 to 174, wherein: The first content and the second content are respectively displayed concurrently by the second computer system and the third computer system; and The first content is different from the second content.

176. The method according to any one of claims 166 to 175, wherein: Determining that the second computer system meets the one or more criteria is based on visually detecting, via the one or more cameras, a change in the visual appearance of a first input device communicating with the second computer system; and Determining that the third computer system meets the one or more criteria is based on visually detecting a change in the visual appearance of a second input device communicating with the third computer system.

177. The method according to any one of claims 166 to 176, wherein: Determining that the second computer system meets the one or more criteria is based on detecting, via the one or more input devices, first light emitted from a first light source associated with the second computer system; and Determining that the third computer system meets the one or more criteria is based on detecting second light emitted from a second light source associated with the third computer system.

178. The method according to any one of claims 166 to 177, wherein: Determining that the second computer system meets the one or more criteria is based on visually detecting, via the one or more cameras, a first light pattern emitted from a first light source associated with the second computer system; And Determining that the third computer system meets the one or more criteria is based on visually detecting a second light pattern emitted from a second light source associated with the third computer system.

179. The method according to any one of claims 166 to 178, wherein: Determining that the second computer system meets the one or more criteria is based on detecting, via the one or more input devices, a first signal generated by the second computer system; And Determining that the third computer system meets the one or more criteria is based on detecting a second signal generated by the third computer system.

180. The method according to any one of claims 166 to 179, wherein detecting the request to establish the connection with the corresponding computer system comprises: Detecting, via the one or more input devices, a respective input associated with a position within a threshold distance of a portion of the respective computer system.

181. The method according to any one of claims 166 to 180, wherein detecting the request to establish the connection with the corresponding computer system comprises: Detecting, via the one or more input devices, a respective input corresponding to the request.

182. The method according to any one of claims 166 to 181, wherein detecting the request to establish the connection with the corresponding computer system comprises: Detecting that the respective computer system has detected an indication of a respective input corresponding to the request.

183. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: Detecting, via the one or more input devices, a request to establish a connection with a respective computer system different from the first computer system, the respective computer system being within a respective region of the physical environment of the first computer system; and In response to detecting the request: Establishing a connection between the first computer system and the second computer system and not establishing a connection between the first computer system and other computer systems among the plurality of computer systems, based on determining that the second computer system among the plurality of computer systems meets one or more criteria when the plurality of computer systems are within the respective region; And Based on determining that a third computer system different from the second computer system among the multiple computer systems satisfies the one or more criteria when the multiple computer systems are within the corresponding region, establish a connection between the first computer system and the third computer system, and do not establish a connection between the first computer system and other computer systems among the multiple computer systems that include the second computer system.

184. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: Detect, via the one or more input devices, a request to establish a connection with a corresponding computer system different from the first computer system, the corresponding computer system being within a corresponding region of the physical environment of the first computer system; and In response to detecting the request: Based on determining that a second computer system among the multiple computer systems satisfies one or more criteria when the multiple computer systems are within the corresponding region, establish a connection between the first computer system and the second computer system, and do not establish a connection between the first computer system and other computer systems among the multiple computer systems; and Based on determining that a third computer system different from the second computer system among the multiple computer systems satisfies the one or more criteria when the multiple computer systems are within the corresponding region, establish a connection between the first computer system and the third computer system, and do not establish a connection between the first computer system and other computer systems among the multiple computer systems that include the second computer system.

185. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; Means for detecting, via the one or more input devices, a request to establish a connection with a corresponding computer system different from the first computer system, the corresponding computer system being within a corresponding region of the physical environment of the first computer system; and Means for performing the following operations in response to detecting the request: Based on determining that a second computer system among the multiple computer systems satisfies one or more criteria when the multiple computer systems are within the corresponding region, establish a connection between the first computer system and the second computer system, and do not establish a connection between the first computer system and other computer systems among the multiple computer systems; And Based on determining that a third computer system different from the second computer system among the multiple computer systems meets the one or more criteria when the multiple computer systems are within the corresponding region, establish a connection between the first computer system and the third computer system, and do not establish a connection between the first computer system and other computer systems among the multiple computer systems that include the second computer system.

186. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any of the methods according to claims 166 to 182.

187. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, cause the computer system to perform any of the methods according to claims 166 to 182.

188. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And Components for performing any of the methods according to claims 166 to 182.