Method for managing overlapping windows and applying visual effects

By detecting overlap and position changes of virtual objects and dynamically adjusting visual effects, the problem of inefficiency in interaction in virtual reality and augmented reality environments is solved, and more efficient user input and energy savings are achieved.

CN120303636APending Publication Date: 2025-07-11APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480005202.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-02
Filing Date
2024-06-04
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing methods of interacting with virtual reality and augmented reality environments are inefficient, user input is cumbersome and error-prone, resulting in waste of energy in computer systems and increased user cognitive burden.

Method used

By detecting the amount of overlap and position changes between virtual objects, dynamically adjust the visual prominence of virtual objects, and respond to transparent visibility events and background states, visual effects are applied to simplify user input and improve interaction efficiency.

Benefits of technology

Reduces the number and complexity of user input, improves the efficiency of the human-computer interface, saves the energy consumption of the computer system, extends the battery life, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303636A_ABST
    Figure CN120303636A_ABST
Patent Text Reader

Abstract

In some implementations, a computer system changes visual saliency of a respective virtual object in response to detecting a threshold amount of overlap between a first virtual object and a second virtual object. In some implementations, the computer system changes the visual saliency of the respective virtual object based on a change in the spatial position of the first virtual object relative to the second virtual object. In some embodiments, a computer system applies a visual effect to a physical object, a virtual environment, and / or a representation of a physical environment. In some implementations, a computer system changes the visual saliency of a virtual object relative to a three-dimensional environment based on the display of different types of overlapping objects in the three-dimensional environment. In some embodiments, a computer system changes an opacity level of a first virtual object that overlaps a second virtual object in response to movement of the first virtual object.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 587,442, filed Oct. 2, 2023; U.S. Provisional Application No. 63 / 515,119, filed Jul. 23, 2023; U.S. Provisional Application No. 63 / 506,128, filed Jun. 4, 2023; and U.S. Provisional Application No. 63 / 506,109, filed Jun. 4, 2023, the entire contents of which are incorporated herein by reference for all purposes. Technical Field

[0003] The present invention generally relates to computer systems for providing computer-generated experiences, including but not limited to electronic devices for providing virtual reality and mixed reality experiences via a display. Background Art

[0004] In recent years, the development of computer systems for augmented reality has increased significantly. Example augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays) for computer systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Example virtual elements include virtual objects such as digital images, videos, text, icons, and control elements (such as buttons and other graphics). Summary of the Invention

[0005] Some methods and interfaces for interacting with an environment that includes at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems in which virtual object manipulation is complex, cumbersome, and error-prone impose a significant cognitive burden on the user and detract from the experience of the virtual / augmented reality environment. In addition, these methods take longer than necessary, thereby wasting the energy of the computer system. This latter consideration is particularly important in battery-powered devices.

[0006] Accordingly, there is a need for computer systems with improved methods and interfaces for providing computer-generated experiences to users such that the interaction of the users with the computer systems is more effective and intuitive for the users. Such methods and interfaces optionally supplement or replace conventional methods for providing extended reality experiences. Such methods and interfaces form a more effective human-machine interface by helping the user understand the connection between the inputs provided and the device's response to those inputs, thereby reducing the quantity, degree, and / or nature of the inputs from the user.

[0007] The above-described deficiencies and other problems associated with the user interface of a computer system are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a laptop computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a watch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has (e.g., includes or communicates with) a display generation component (e.g., a display device such as a head-mounted device (HMD), a monitor, a projector, a touch-sensitive display (also referred to as a “touchscreen” or “touchscreen display”), or other devices or components that present visual content to the user on or within the display generation component itself or that generate and are visible elsewhere). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, in addition to the display generation component, the computer system has one or more output devices that include one or more haptic output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in the memory for performing multiple functions. In some embodiments, the user interacts with the GUI through contact and gestures of a stylus and / or finger on a touch-sensitive surface, movement of the user's eyes and hands in space relative to the GUI (and / or the computer system) or the user's body (as captured by cameras and other motion sensors), and / or voice input (as captured by one or more audio input devices). In some embodiments, the functions performed through the interaction optionally include image editing, drawing, presentation, word processing, spreadsheet creation, gaming, making phone calls, video conferencing, sending and receiving emails, instant messaging, test support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. The executable instructions for performing these functions are optionally included in a transient and / or non-transient computer-readable storage medium or other computer program product configured to be executed by one or more processors.

[0008] There is a need for electronic devices with improved methods and interfaces for interacting with three-dimensional environments. Such methods and interfaces can supplement or replace conventional methods for interacting with three-dimensional environments. Such methods and interfaces reduce the amount, degree, and / or nature of input from the user and result in a more efficient human-machine interface. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charges.

[0009] In some embodiments, the computer system changes the visual prominence of a corresponding virtual object in response to detecting a threshold amount of overlap between a first virtual object and a second virtual object. In some embodiments, the computer system changes the visual prominence of a corresponding virtual object based on a change in the spatial position of the first virtual object relative to the second virtual object. In some embodiments, the computer system applies a visual effect to a real-world object in response to detecting a passthrough visibility event (e.g., an event in which a real-world object becomes visible via the computer system). In some embodiments, the computer system applies a visual effect to the background based on the state of the background. In some embodiments, the computer system applies a visual effect associated with a virtual object based on the state of the virtual object. In some embodiments, the computer system changes the visual prominence of a virtual object relative to the three-dimensional environment based on the display of different types of overlapping objects in the three-dimensional environment. In some embodiments, the computer system changes the opacity level of a first virtual object in response to the movement of the first virtual object that overlaps a second virtual object.

[0010] Note that the various embodiments described above can be combined with any other embodiment described herein. The features and advantages described in this specification are not exhaustive, and in particular, many additional features and advantages will be apparent to those of ordinary skill in the art from the drawings, the specification, and the claims. Additionally, it should be noted that for readability and guidance purposes, the language used in this specification has been selected in principle and may not have been so selected to depict or define the subject matter of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To better understand the various described embodiments, reference should be made to the following detailed description in conjunction with the following drawings, in which like reference numerals indicate corresponding parts in all the drawings.

[0012] Figure 1A is a block diagram showing the operating environment of a computer system for providing an XR experience according to some embodiments.

[0013] Figures 1B to 1P is for providing an example of a computer system for an XR experience in the Figure 1A operating environment.

[0014] Figure 2 is a block diagram of a controller configured to manage and coordinate a user's XR experience in a computer system according to some embodiments.

[0015] Figure 3 is a block diagram of a display generation component configured to provide a visual component of an XR experience to a user in a computer system according to some embodiments.

[0016] Figure 4 is a block diagram of a hand tracking unit configured to capture a user's gesture input in a computer system according to some embodiments.

[0017] Figure 5 is a block diagram of an eye tracking unit configured to capture a user's gaze input in a computer system according to some embodiments.

[0018] Figure 6 is a flowchart of a flash-assisted gaze tracking pipeline according to some embodiments.

[0019] Figures 7A to 7EE shows an example of changing the visual prominence of a corresponding virtual object in a three-dimensional environment.

[0020] Figure 8 is a flowchart of an exemplary method of changing the visual prominence of a corresponding virtual object in response to a threshold amount of overlap between a first virtual object and a second virtual object.

[0021] Figure 9 is a flowchart of an exemplary method of changing the visual prominence of a corresponding virtual object based on a change in the spatial position of a first virtual object relative to a second virtual object.

[0022] Figures 10A to 10N1 shows an example of applying a visual effect to a real-world object.

[0023] Figure 11 is a flowchart of an exemplary method of applying a visual effect to a real-world object.

[0024] Figures 12A to 12Q1 shows an example of applying a visual effect to a background.

[0025] Figure 13 is a flowchart of an exemplary method of applying a visual effect to a background.

[0026] Figures 14A to 14K shows an example of applying a visual effect based on the state of a virtual object.

[0027] Figure 15 is a flowchart of a method of applying a visual effect based on the state of a virtual object.

[0028] Figures 16A to 16K Shows an example of a computer system that changes the visual prominence of a virtual object based on the display of different types of overlapping objects in a three - dimensional environment according to some embodiments.

[0029] Figure 17 Is a flowchart showing a method for changing the visual prominence of a virtual object based on the display of different types of overlapping objects according to some embodiments.

[0030] Figures 18A to 18T Shows an example of a computer system that changes the visual prominence of a virtual object to resolve a simulated overlap with another virtual object according to some embodiments.

[0031] Figure 19 Is a flowchart showing a method for changing the visual prominence of a virtual object to resolve a simulated overlap with another virtual object according to some embodiments. Detailed Description

[0032] According to some embodiments, the present disclosure relates to a user interface for providing an extended reality (XR) experience to a user.

[0033] The systems, methods, and GUIs described herein improve user interface interactions with virtual / augmented reality environments in a variety of ways.

[0034] In some embodiments, the computer system changes the visual prominence of a corresponding virtual object in a three - dimensional environment in response to detecting that at least a portion of a first virtual object overlaps a second virtual object by more than a threshold amount from the user's current viewpoint.

[0035] In some embodiments, the computer system reduces the visual prominence of a portion of a corresponding virtual object and changes the visual prominence of that portion of the corresponding virtual object based on a change in the spatial position of the first virtual object relative to the second virtual object during movement of the first virtual object in the three - dimensional environment.

[0036] In some embodiments, the computer system applies a visual effect (such as a dimming effect or a coloring effect) to a real - world object in response to detecting a see - through visibility event in which the real - world object becomes visible in the three - dimensional environment presented by the computer system.

[0037] In some embodiments, when virtual content is displayed in a three - dimensional environment and when a background is visible in the three - dimensional environment (e.g., a background optionally including a representation of a virtual environment and / or a physical environment), the computer system applies (or refrains from applying) a visual effect to the background based on the state of the background (such as a state associated with the current time - of - day setting).

[0038] In some embodiments, the computer system applies (or forgoes applying) visual effects associated with a virtual object (e.g., a virtual application window) based on whether the virtual object is active or inactive.

[0039] In some embodiments, the computer system changes the visual prominence of a virtual object in response to detecting an event that causes a user interface element to be displayed overlapping the virtual object in a three-dimensional environment, such as by changing the brightness and / or translucency of the virtual object.

[0040] In some embodiments, the computer system changes the opacity level of a first virtual object in response to movement of the first virtual object that overlaps a second virtual object.

[0041] Figures 1A to 6 A description of an exemplary computer system for providing an XR experience to a user is provided (such as described below with respect to methods 800, 900, 1100, 1300, and / or 1500). Figures 7A to 7EE An example of a computer system that changes the visual prominence of a corresponding virtual object relative to a three-dimensional environment is shown, according to some embodiments. Figure 8 is a flowchart showing an exemplary method of changing the visual prominence of a corresponding virtual object relative to a three-dimensional environment in response to detecting a threshold amount of overlap between a first virtual object and a second virtual object in the three-dimensional environment. Figures 7A to 7EE The user interface in Figure 8 is used to show the Figure 9 is a flowchart showing a method of changing the visual prominence of a corresponding virtual object based on a change in the spatial position of a first virtual object relative to a second virtual object in a three-dimensional environment, according to some embodiments. Figures 7A to 7EE The user interface in Figure 9 is used to show the Figures 10A to 10N An example technique for applying visual effects to real-world objects is shown, according to some embodiments. Figure 11 is a flowchart of a method of applying visual effects to real-world objects, according to various embodiments. Figures 10A to 10F The user interface in Figure 11 is used to show the Figures 12A to 12Q An example technique for applying visual effects to a background is shown, according to some embodiments. Figure 13 is a flowchart of a method of applying visual effects to a background, according to various embodiments. Figures 12A to 12Q The user interface in Figure 13 is used to show the Figures 14A to 14K An example technique for applying visual effects based on the state of a virtual object is shown, according to some embodiments. Figure 15Flowchart of a method for applying visual effects based on the state of a virtual object according to various embodiments. Figures 14A to 14K The user interface in Figure 15 is used to show the Figures 16A to 16K process in Figure 17 shows example techniques for changing the visual prominence of a virtual object based on the display of different types of overlapping objects in a three - dimensional environment according to various embodiments. Figures 16A to 16K The user interface in Figure 17 is used to show the Figures 18A to 18T process in Figure 19 Flowchart of a method for changing the visual prominence of a virtual object to resolve simulated overlap with another virtual object according to some embodiments. Figures 18A to 18T The user interface in Figure 19 is used to show the

[0042] The processes described below enhance the operability of the device and make the user - device interface more efficient through various techniques (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device). These techniques include providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional display controls, performing an operation when a set of conditions has been met without further user input, improving privacy and / or security, providing a more diverse, detailed, and / or realistic user experience while saving storage space, and / or additional techniques. These techniques also reduce power usage and extend the battery life of the device by enabling the user to use the device faster and more efficiently. Saving battery power, and thus weight, improves the ergonomics of the device. These techniques also enable real - time communication, allow the use of fewer and / or less precise sensors, resulting in a more compact, lighter, and cheaper device, and enable the device to be used under various lighting conditions. These techniques reduce energy usage, thereby reducing the heat emitted by the device, which is particularly important for wearable devices, where wearing the device can become uncomfortable for the user if the device generates too much heat within the operating parameters of its components.

[0043] In addition, in a method where one or more of the steps described herein depend on one or more conditions being satisfied, it should be understood that the method can be repeated in multiple iterations such that, during the repetition, all conditions that determine the steps in the method are satisfied in different iterations of the method. For example, if a method requires performing a first step (if a condition is satisfied) and a second step (if the condition is not satisfied), one of ordinary skill in the art will know to repeat the stated steps until both the condition being satisfied and the condition not being satisfied (in no particular order) occur. Thus, a method described as having one or more steps that depend on one or more conditions being satisfied can be rewritten as a method that repeats until each condition described in the method is satisfied. However, this does not require the system or computer-readable medium to state that the system or computer-readable medium includes instructions for performing conditional operations based on the satisfaction of the corresponding one or more conditions and is thus capable of determining whether the possible conditions have been satisfied without explicitly repeating the steps of the method until all conditions that determine the steps in the method are satisfied. One of ordinary skill in the art will also understand that, similar to a method with conditional steps, a system or computer-readable storage medium can repeat the steps of the method as many times as needed to ensure that all conditional steps have been performed.

[0044] In some embodiments, as Figure 1A shown, an XR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a haptic sensor, an orientation sensor, a proximity sensor, a temperature sensor, a position sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., household appliances, wearable devices, etc.). In some embodiments, one or more of the input devices 125, output devices 155, sensors 190, and peripheral devices 195 are integrated with the display generation component 120 (e.g., in a head-mounted device or a handheld device).

[0045] When describing XR experiences, various terms are used to distinctively refer to several related but different environments that a user can sense and / or with which the user can interact (e.g., interact using inputs detected by the computer system 101 that generates the XR experience, where these inputs cause the computer system generating the XR experience to generate audio, visual, and / or haptic feedback corresponding to the various inputs provided to the computer system 101). The following is a subset of these terms:

[0046] Physical environment: The physical environment refers to the physical world that people can sense and / or interact with without the assistance of an electronic system. Physical environments such as a physical park include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment, such as through vision, touch, hearing, taste, and smell.

[0047] Extended reality: In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via an electronic system. In XR, a subset of a person's physical movements or representations thereof are tracked, and in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner consistent with at least one physical law. For example, an XR system can detect a person's head rotation, and in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds would change in the physical environment. In some cases (e.g., for accessibility reasons), the adjustment of the characteristics of virtual objects in the XR environment can be made in response to a representation of a physical movement (e.g., a voice command). A person can use any of their senses to sense and / or interact with XR objects, including vision, hearing, touch, taste, and smell. For example, a person can sense and / or interact with an audio object that creates a 3D or spatial audio environment that provides the perception of point audio sources in 3D space. As another example, an audio object can enable audio transparency that selectively introduces ambient sounds from the physical environment with or without computer-generated audio. In certain XR environments, a person can sense and / or interact only with audio objects.

[0048] Examples of XR include virtual reality and mixed reality.

[0049] Virtual Reality: A virtual reality (VR) environment is an environment designed to be a simulated environment that is completely computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with the virtual objects in a VR environment through the simulation of the person's presence within the computer-generated environment and / or through the simulation of a subset of the person's physical movements within the computer-generated environment.

[0050] Mixed Reality: Compared with a VR environment designed to be completely computer-generated sensory input, a mixed reality (MR) environment is an environment designed to include, in addition to computer-generated sensory input (e.g., virtual objects), sensory input from the physical environment or a representation thereof. On the virtual continuum, a mixed reality environment is any condition between a completely physical environment at one end and a virtual reality environment at the other end, but excluding these two ends. In some MR environments, the computer-generated sensory input can respond to changes in the sensory input from the physical environment. Additionally, some electronic systems for presenting an MR environment can track the position and / or orientation relative to the physical environment so that virtual objects can interact with real objects (i.e., physical items from the physical environment or a representation thereof). For example, the system can cause movement so that a virtual tree appears stationary relative to the physical ground.

[0051] Examples of mixed reality include augmented reality and augmented virtuality.

[0052] Augmented Reality: An augmented reality (AR) environment is a simulated environment in which one or more virtual objects are superimposed over a physical environment or a representation of a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or semi-transparent display through which a person can directly view the physical environment. The system can be configured to present virtual objects on the transparent or semi-transparent display such that a person using the system perceives the virtual objects superimposed over the physical environment. Alternatively, the system can have an opaque display and one or more imaging sensors that capture images or video of the physical environment, which are representations of the physical environment. The system combines the images or video with the virtual objects and presents the combination on the opaque display. A person using the system indirectly views the physical environment via the images or video of the physical environment and perceives the virtual objects superimposed over the physical environment. As used herein, the video of the physical environment displayed on the opaque display is referred to as “passthrough video,” meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Further alternatively, the system can have a projection system that projects virtual objects into the physical environment, such as as a hologram or on a physical surface, such that a person using the system perceives the virtual objects superimposed over the physical environment. An augmented reality environment is also a simulated environment in which a representation of the physical environment is transformed by computer-generated sensory information. For example, in providing passthrough video, the system can transform one or more sensor images to impose a selected perspective (e.g., a viewpoint) that is different from the perspective captured by the imaging sensors. As another example, the representation of the physical environment can be transformed by graphically modifying (e.g., magnifying) portions thereof such that the modified portions can be a representative but not a true version of the originally captured image. As yet another example, the representation of the physical environment can be transformed by graphically removing portions thereof or blurring portions thereof.

[0053] Augmented Virtuality: An augmented virtuality (AV) environment is a simulated environment in which a virtual environment or computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory inputs can be representations of one or more characteristics of the physical environment. For example, an AV park can have virtual trees and virtual buildings, but a person's face is a realistic reproduction from an image of a physical person. As another example, a virtual object can adopt the shape or color of a physical item imaged by one or more imaging sensors. As yet another example, a virtual object can adopt a shadow that conforms to the position of the sun in the physical environment.

[0054] In an augmented reality, mixed reality, or virtual reality environment, a view of a three-dimensional environment is visible to a user. The view of the three-dimensional environment is typically visible to the user through a virtual viewport via one or more display generation components (e.g., a display or a pair of display modules that provide stereoscopic content to different eyes of the same user), the virtual viewport having a viewport boundary that defines the extent of the three-dimensional environment visible to the user via the one or more display generation components. In some embodiments, the region defined by the viewport boundary is smaller than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size, optical properties, or other physical characteristics of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user's eyes). In some embodiments, the region defined by the viewport boundary is larger than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size, optical properties, or other physical characteristics of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user's eyes). The viewport and the viewport boundary typically move as the one or more display generation components move (e.g., for a head-mounted device as the user's head moves, or for a handheld device such as a tablet or smartphone as the user's hand moves). The user's viewpoint determines the content visible in the viewport, the viewpoint typically specifying a position and orientation relative to the three-dimensional environment, and as the viewpoint moves, the view of the three-dimensional environment will also move in the viewport. For a head-mounted device, the viewpoint is typically based on the position and orientation of the user's head, face, and / or eyes to provide a perceptually accurate view of the three-dimensional environment and an immersive experience while the user is using the head-mounted device. For a handheld or stationary device, the viewpoint moves as the handheld or stationary device moves and / or as the user's positioning relative to the handheld or stationary device changes (e.g., the user moves towards, away from, up, down, right, and / or left). For a device that includes a display generation component with virtual passthrough, the portions of the physical environment visible (e.g., displayed and / or projected) via the one or more display generation components are based on the field of view of one or more cameras in communication with the display generation component, the one or more cameras typically moving as the display generation component moves (e.g., for a head-mounted device as the user's head moves, or for a handheld device such as a tablet or smartphone as the user's hand moves), because the user's viewpoint moves as the field of view of the one or more cameras moves (and the appearance of one or more virtual objects displayed via the one or more display generation components is updated based on the user's viewpoint (e.g., the display position and pose of the virtual object are updated based on the movement of the user's viewpoint)).For a display generation component with optical see-through, portions of the physical environment that are visible via one or more display generation components (e.g., optically visible through one or more portions or completely transparent portions of the display generation component) are based on the user's field of view through the partial or completely transparent portion of the display generation component (e.g., for a head-mounted device, it moves as the user's head moves, or for a handheld device such as a tablet or smartphone, it moves as the user's hand moves), because the user's viewpoint moves as the user's field of view through the partial or completely transparent portion of the display generation component moves (and the appearance of one or more virtual objects is updated based on the user's viewpoint).

[0055] In some embodiments, the representation of the physical environment (e.g., via virtual passthrough or optical passthrough display) may be partially or fully occluded by the virtual environment. In some embodiments, the amount of the virtual environment displayed (e.g., the amount of the physical environment not displayed) is based on the immersion level of the virtual environment (e.g., relative to the representation of the physical environment). For example, increasing the immersion level optionally causes more of the virtual environment to be displayed, replacing and / or occluding more of the physical environment, and decreasing the immersion level optionally causes less of the virtual environment to be displayed, thereby revealing portions of the physical environment that were previously not displayed and / or occluded. In some embodiments, at a particular immersion level, one or more first background objects (e.g., in the representation of the physical environment) are visually de-emphasized (e.g., dimmed, blurred, displayed with increased transparency) more than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, the immersion level includes the associated degree to which virtual content (e.g., virtual environment and / or virtual content) displayed by a computer system occludes background content (e.g., content other than the virtual environment and / or virtual content) around / behind the virtual environment, optionally including the number of items of the displayed background content and / or the displayed visual characteristics (e.g., color, contrast, and / or opacity) of the background content, the angular range of the virtual content displayed by a display generation component (e.g., 60 degrees for content displayed at low immersion, 120 degrees for content displayed at medium immersion, or 180 degrees for content displayed at high immersion), and / or the proportion of the field of view displayed by a display generation component occupied by the virtual content (e.g., 33% of the field of view occupied by the virtual content at low immersion, 66% of the field of view occupied by the virtual content at medium immersion, or 100% of the field of view occupied by the virtual content at high immersion). In some embodiments, the background content is included in the background on which the virtual content is displayed (e.g., the background content in the representation of the physical environment). In some embodiments, the background content includes a user interface (e.g., a user interface generated by a computer system corresponding to an application program), virtual objects not associated with and / or not included in the virtual environment and / or virtual content (e.g., files or representations of other users generated by a computer system, etc.), and / or real objects (e.g., passthrough objects representing real objects in the physical environment around the user, these passthrough objects being visible such that they are displayed by a display generation component and / or visible via a transparent or translucent component of the display generation component because the computer system does not occlude / hinder their visibility through the display generation component). In some embodiments, at a low immersion level (e.g., a first immersion level), the background, virtual, and / or real objects are displayed in an unoccluded manner. For example, a virtual environment with a low immersion level is optionally displayed simultaneously with the background content, which is optionally displayed at full brightness, color, and / or semi-transparency.In some embodiments, at higher levels of immersion (e.g., a second level of immersion higher than a first level of immersion), background, virtual, and / or real objects are displayed in an occluded manner (e.g., dimmed, blurred, or removed from the display). For example, a corresponding virtual environment with a high level of immersion is displayed without simultaneously displaying background content (e.g., in full-screen or fully immersive mode). As another example, a virtual environment displayed at a medium level of immersion is displayed simultaneously with background content that is dimmed, blurred, or otherwise de-emphasized. In some embodiments, the visual characteristics of background objects vary among the background objects. For example, at a particular level of immersion, one or more first background objects are more visually de-emphasized (e.g., dimmed, blurred, and / or displayed with increased transparency) than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, zero immersion or a zero level of immersion corresponds to a virtual environment that ceases to be displayed, and instead a representation of the physical environment is displayed (optionally with one or more virtual objects, such as applications, windows, or virtual three-dimensional objects), and the representation of the physical environment is not occluded by the virtual environment. Adjusting the level of immersion using physical input elements provides a quick and efficient way to adjust the degree of immersion, which enhances the operability of the computer system and makes the user-device interface more efficient.

[0056] Viewpoint-locked virtual objects: When a computer system displays a virtual object at the same position and / or orientation in the user's viewpoint, the virtual object is viewpoint-locked even if the user's viewpoint shifts (e.g., changes). In embodiments in which the computer system is a head-mounted device, the user's viewpoint is locked to the forward direction of the user's head (e.g., when the user looks straight ahead, the user's viewpoint is at least a portion of the user's field of view); thus, without moving the user's head, the user's viewpoint remains fixed even when the user's gaze shifts. In embodiments in which the computer system has a display generation component (e.g., a display screen) that is repositionable relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the display generation component of the computer system. For example, a viewpoint-locked virtual object that is displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's viewpoint even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the position and / or orientation at which the viewpoint-locked virtual object is displayed in the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments in which the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, such that the virtual object is also referred to as a "head-locked virtual object".

[0057] Environment-Locked Visual Objects: When a computer system displays a virtual object at a location and / or orientation within a user's line of sight, the virtual object is environment-locked (alternatively, "world-locked"), where the location and / or orientation is based on a location and / or object within a three-dimensional environment (e.g., a physical or virtual environment), such that the virtual object is selected and / or anchored with reference to the location and / or object. As the user's line of sight moves, the location and / or object within the environment relative to the user's line of sight changes, which causes the environment-locked virtual object to be displayed at a different location and / or orientation within the user's line of sight. For example, an environment-locked virtual object locked to a tree directly in front of the user is displayed at the center of the user's line of sight. When the user's line of sight shifts to the right (e.g., the user's head turns to the right) such that the tree is now to the left of center within the user's line of sight (e.g., the orientation of the tree within the user's line of sight is shifted), the environment-locked virtual object locked to the tree is displayed to the left of center within the user's line of sight. In other words, the location and / or orientation at which the environment-locked virtual object is displayed within the user's line of sight depends on the location and / or orientation of the location and / or object to which the virtual object is locked within the environment. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system anchored to a fixed location and / or object within a physical environment) in order to determine the orientation at which the environment-locked virtual object is displayed within the user's line of sight. The environment-locked virtual object may be locked to a stationary portion of the environment (e.g., a floor, wall, table, or other stationary object), or may be locked to a movable portion of the environment (e.g., a vehicle, animal, person, or even a representation of a part of the user's body such as the user's hand, wrist, arm, or foot that moves independently of the user's line of sight) such that the virtual object moves as the line of sight or that portion of the environment moves to maintain a fixed relationship between the virtual object and that portion of the environment.

[0058] In some embodiments, an environment-locked or view-locked virtual object exhibits lazy follow behavior, which reduces or delays the movement of the environment-locked or view-locked virtual object relative to the movement of a reference point that the virtual object follows. In some embodiments, when exhibiting lazy follow behavior, when the computer system detects movement of a reference point that the virtual object is following (e.g., a part of the environment, the view point, or a point fixed relative to the view point, such as a point between 5 cm and 300 cm from the view point), the computer system intentionally delays the movement of the virtual object. For example, when the reference point (e.g., the part of the environment or the view point) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point, but at a second speed that is slower than the first speed (e.g., until the reference point stops moving or slows down, at which point the virtual object begins to catch up to the reference point). In some embodiments, when the virtual object exhibits lazy follow behavior, the device ignores small movements of the reference point (e.g., ignores movements of the reference point below a threshold movement amount, such as moving 0 degrees to 5 degrees or moving 0 cm to 50 cm). For example, when the reference point (e.g., the part of the environment or the view point to which the virtual object is locked) moves a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed to maintain a fixed or substantially fixed orientation relative to a different view point or part of the environment than the reference point to which the virtual object is locked), and when the reference point (e.g., the part of the environment or the view point to which the virtual object is locked) moves a second amount that is greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is being displayed to maintain a fixed or substantially fixed orientation relative to a different view point or part of the environment than the reference point to which the virtual object is locked), and then decreases when the movement amount of the reference point increases above a threshold (e.g., the "lazy follow" threshold), because the virtual object is moved by the computer system to maintain a fixed or substantially fixed orientation relative to the reference point. In some embodiments, the virtual object maintaining a substantially fixed orientation relative to the reference point includes the virtual object being displayed within a threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 15 cm, 20 cm, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward relative to the orientation of the reference point).

[0059] Hardware: There are many different types of electronic systems that enable a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields integrated with display capabilities, windows integrated with display capabilities, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablet devices, and desktop / laptop computers. A head-mounted system can have one or more speakers and an integrated opaque display. Alternatively, the head-mounted system can be configured to accept an external opaque display (e.g., a smartphone). The head-mounted system can incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. The head-mounted system can have a transparent or translucent display instead of an opaque display. The transparent or translucent display can have a medium through which light representing an image is directed to a person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light sources, or any combination of these technologies. The medium can be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display can be configured to selectively become opaque. A projection-based system can employ retinal projection technology that projects a graphical image onto a person's retina. The projection system can also be configured to project virtual objects into the physical environment, such as as a hologram or on a physical surface. In some embodiments, the controller 110 is configured to manage and coordinate a user's XR experience. In some embodiments, the controller 110 includes a suitable combination of software, firmware, and / or hardware. Below with respect to Figure 2Controller 110 is described in more detail. In some embodiments, controller 110 is a computing device that is local or remote relative to scene 105 (e.g., a physical environment). For example, controller 110 is a local server located within scene 105. As another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside of scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., an HMD, a display, a projector, a touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, controller 110 is included within the housing (e.g., a physical enclosure) of one or more of display generation component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors, etc.), input device 125, output device 155, one or more sensors of sensor 190, and / or one or more peripheral devices of peripheral device 195, or shares the same physical housing or support structure with one or more of the above devices.

[0060] In some embodiments, display generation component 120 is configured to provide a user with an XR experience (e.g., at least a visual component of the XR experience). In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Display generation component 120 is described in more detail below with respect to Figure 3 In some embodiments, the functionality of controller 110 is provided by and / or in combination with display generation component 120.

[0061] According to some embodiments, when a user is virtually and / or physically present within scene 105, display generation component 120 provides the user with an XR experience.

[0062] In some embodiments, the display generation component is worn on a part of the user's body (e.g., on his / her head, on his / her hand, etc.). Thus, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smart phone or a tablet device) configured to present XR content, and the user holds the device with a display facing the user's field of view and a camera facing the scene 105. In some embodiments, the handheld device is optionally placed in a housing worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is an XR chamber, housing, or room configured to present XR content, where the user does not wear or hold the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) can be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing an interaction with XR content triggered based on an interaction occurring in the space in front of a handheld device or a tripod-mounted device can be similarly implemented with an HMD, where the interaction occurs in the space in front of the HMD and the response to the XR content is displayed via the HMD. Similarly, a user interface showing an interaction with XR content triggered based on the movement of a handheld device or a tripod-mounted device relative to the physical environment (e.g., the scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)) can be similarly implemented with an HMD, where the movement is caused by the movement of the HMD relative to the physical environment (e.g., the scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)).

[0063] Although relevant features of the operating environment 100 are shown in Figure 1A those of ordinary skill in the art will understand from this disclosure that, for the sake of brevity and to not obscure more relevant aspects of the example embodiments disclosed herein, various other features are not illustrated.

[0064] Figures 1A to 1PShows various examples of computer systems for performing methods and providing audio, visual, and / or tactile feedback as part of the user interfaces described herein. In some embodiments, the computer system includes one or more display generation components (e.g., a first display component and a second display component 1-120a, 1-120b and / or a first optical module and a second optical module 11.1.1-104a and 11.1.1-104b) for displaying virtual elements and / or a representation of the physical environment to a user of the computer system, the virtual elements and / or the representation of the physical environment optionally being generated based on detected events and / or user input detected by the computer system. The user interface generated by the computer system is optionally corrected by one or more corrective lenses 11.3.2-216 such that users who would otherwise use glasses or contact lenses to correct their vision can more easily view the user interface, the one or more corrective lenses optionally being removably attached to one or more of the optical modules. Although many of the user interfaces shown herein show a single view of the user interface, the user interface in an HMD optionally uses two optical modules (e.g., a first display component 1-120a and a second display component 1-120b and / or a first optical module 11.1.1-104a and a second optical module 11.1.1-104b) to display, one optical module for the user's right eye and a different optical module for the user's left eye, and presenting slightly different images to the two different eyes to create an illusion of stereoscopic depth, the single view of the user interface is typically a right-eye view or a left-eye view, and the depth effect is explained in the text or using other schematic diagrams or views. In some embodiments, the computer system includes one or more external displays (e.g., display component 1-108) for displaying status information of the computer system to a user of the computer system (when the computer system is not being worn) and / or to others in the vicinity of the computer system, the status information optionally being generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more audio output components (e.g., electronic component 1-112) for generating audio feedback, the audio feedback optionally being generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting information about the physical environment of the device (e.g., sensor component 1-356 and / or Figure 1I one or more of the sensors), the one or more sensors being usable (optionally in conjunction with one or more illuminators, such as Figure 1Iin combination with the illuminator described therein) generate a digital see-through image, capture visual media (e.g., photos and / or videos) corresponding to the physical environment, or determine the pose (e.g., position and / or orientation) of physical objects and / or surfaces in the physical environment such that virtual objects can be placed based on the detected pose of the physical objects and / or surfaces. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting hand position and / or movement (e.g., sensor assemblies 1-356 and / or Figure 1I one or more of the sensors therein), which one or more sensors can be used (optionally in combination with one or more illuminators, such as Figure 1I the illuminator 6-124 described therein) to determine when one or more air gestures have been performed. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting eye movement (e.g., Figure 1I the eye tracking and gaze tracking sensors therein), which sensors can be used (optionally in combination with one or more lights, such as Figure 1OThe lights in (11.3.2 - 110) determine the attention or gaze positioning and / or gaze movement, which can optionally be used to detect only gaze input based on gaze movement and / or dwell. Combinations of the various sensors described above can be used to determine the user's facial expressions and / or hand movements for generating an avatar or representation of the user, such as an anthropomorphic avatar or representation for a real - time communication session, where the avatar has facial expressions, hand movements, and / or body movements based on or similar to the detected facial expressions, hand movements, and / or body movements of the user of the device. Gaze and / or attention information is optionally combined with hand - tracking information to determine the interaction between the user and one or more user interfaces based on direct and / or indirect input, such as air gestures or input using one or more hardware input devices, such as one or more buttons (e.g., first button 1 - 128, button 11.1.1 - 114, second button 1 - 132, and / or dial or button 1 - 328), knobs (e.g., first button 1 - 128, button 11.1.1 - 114, and / or dial or button 1 - 328), digital crowns (e.g., first button 1 - 128, button 11.1.1 - 114, and / or dial or button 1 - 328 that can be pressed and twisted or rotated), touchpads, touchscreens, keyboards, mice, and / or other input devices. One or more buttons (e.g., first button 1 - 128, button 11.1.1 - 114, second button 1 - 132, and / or dial or button 1 - 328) are optionally used to perform system operations, such as re - centering the content in the three - dimensional environment visible to the user of the device, displaying the main user interface for launching an application, starting a real - time communication session, or initiating the display of a virtual three - dimensional background. A knob or digital crown (e.g., first button 1 - 128, button 11.1.1 - 114, and / or dial or button 1 - 328 that can be pressed and twisted or rotated) is optionally rotatable to adjust parameters of the visual content, such as the immersion level of the virtual three - dimensional environment (e.g., the extent to which the virtual content occupies the user's viewport in the three - dimensional environment) or other parameters associated with the three - dimensional environment and the virtual content displayed via the optical module (e.g., first display component 1 - 120a and second display component 1 - 120b and / or first optical module 11.1.1 - 104a and second optical module 11.1.1 - 104b).

[0065] Figure 1BShows a front view, a top view, and a perspective view of an example of a head-mounted display (HMD) device 1-100 configured to be worn by a user and provide a virtual and augmented reality (VR / AR) experience. The HMD 1-100 may include a display unit 1-102 or component, an electronic strip assembly 1-104 connected to and extending from the display unit 1-102, and a strap assembly 1-106 fixed to the electronic strip assembly 1-104 at either end. The electronic strip assembly 1-104 and the strap 1-106 may be part of a retention assembly configured to wrap around the user's head to hold the display unit 1-102 against the user's face.

[0066] In at least one example, the strap assembly 1-106 may include a first strap 1-116 configured to wrap around the back of the user's head and a second strap 1-117 configured to extend over the top of the user's head. As shown, the second strap may extend between a first electronic strip 1-105a and a second electronic strip 1-105b of the electronic strip assembly 1-104. The strip assembly 1-104 and the strap assembly 1-106 may be part of a fixation mechanism that extends rearward from the display unit 1-102 and is configured to hold the display unit 1-102 against the user's face.

[0067] In at least one example, the fixation mechanism includes a first electronic strip 1-105a that includes a first proximal end 1-134 coupled to the display unit 1-102 (e.g., the housing 1-150 of the display unit 1-102) and a first distal end 1-136 opposite the first proximal end 1-134. The fixation mechanism may also include a second electronic strip 1-105b that includes a second proximal end 1-138 coupled to the housing 1-150 of the display unit 1-102 and a second distal end 1-140 opposite the second proximal end 1-138. The fixation mechanism may also include a first strap 1-116 and a second strap 1-117, the first strap including a first end 1-142 coupled to the first distal end 1-136 and a second end 1-144 coupled to the second distal end 1-140, and the second strap extending between the first electronic strip 1-105a and the second electronic strip 1-105b. The strips 1-105a to b and the strap 1-116 may be coupled via a connection mechanism or component 1-114. In at least one example, the second strap 1-117 includes a first end 1-146 coupled to the first electronic strip 1-105a between the first proximal end 1-134 and the first distal end 1-136 and a second end 1-148 coupled to the second electronic strip 1-105b between the second proximal end 1-138 and the second distal end 1-140.

[0068] In at least one example, the first and second electronic strips 1-105a to b comprise plastic, metal, or other structural materials forming the shape of substantially rigid strips 1-105a to b. In at least one example, the first and second bands 1-116, 1-117 are formed of an elastomeric flexible material (including woven textiles, rubber, etc.). The first band 1-116 and the second band 1-117 can be flexible to conform to the shape of the user's head when wearing the HMD 1-100.

[0069] In at least one example, one or more of the first and second electronic strips 1-105a to b can define an internal strip volume and include one or more electronic components disposed within the internal strip volume. In one example, as Figure 1B shown, the first electronic strip 1-105a can include an electronic component 1-112. In one example, the electronic component 1-112 can include a speaker. In one example, the electronic component 1-112 can include a computing component, such as a processor.

[0070] In at least one example, the housing 1-150 defines a first front opening 1-152. The front opening is Figure 1B marked as 1-152 in dashed lines in, because the display assembly 1-108 is arranged to occlude the first opening 1-152 from view when the HMD 1-100 is assembled. The housing 1-150 can also define a rear second opening 1-154. The housing 1-150 also defines an internal volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display assembly 1-108, which can include a front cover and a display screen (shown in other figures) disposed in or across the front opening 1-152 to occlude the front opening 1-152. In at least one example, the display screen of the display assembly 1-108 and generally the display assembly 1-108 have a curvature configured to follow the curvature of the user's face. The display screen of the display assembly 1-108 can be curved as shown to complement the user's facial features and the overall curvature from one side of the face to the other, e.g., from left to right and / or from top to bottom, where the display unit 1-102 is pressed.

[0071] In at least one example, the housing 1-150 may define a first aperture 1-126 between a first opening 1-152 and a second opening 1-154, and a second aperture 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 may also include a first button 1-126 disposed in the first aperture 1-128, and a second button 1-132 disposed in the second aperture 1-130. The first button 1-128 and the second button 1-132 can be pressed through the respective apertures 1-126, 1-130. In at least one example, the first button 1-126 and / or the second button 1-132 can be a twist dial and a depressible button. In at least one example, the first button 1-128 is a pressable and twistable dial button, and the second button 1-132 is a pressable button.

[0072] Figure 1C A rear perspective view of the HMD 1-100 is shown. The HMD 1-100 may include a light seal 1-110 that extends rearwardly from the housing 1-150 of the display assembly 1-108 around a perimeter of the housing 1-150, as shown. The light seal 1-110 may be configured to extend from the housing 1-150 to a user's face, around the user's eyes, to block external light from being visible. In one example, the HMD 1-100 may include a first display assembly 1-120a and a second display assembly 1-120b that are disposed at or within a rearward-facing second opening 1-154 defined by the housing 1-150 and / or disposed within an interior volume of the housing 1-150 and configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a-b may include a respective display screen 1-122a, 1-122b that is configured to project light in a rearward direction through the second opening 1-154 toward the user's eyes.

[0073] In at least one example, referring Figure 1B and Figure 1C both, the display assembly 1-108 can be a front-facing forward display assembly that includes a display screen configured to project light in a first forward direction, and the rear display screens 1-122a-b can be configured to project light in a second rearward direction that is opposite the first direction. As described above, the light seal 1-110 can be configured to block light external to the HMD 1-100 from reaching the user's eyes, including light projected by Figure 1B the forward display screen of the display assembly 1-108 shown in the front perspective view of. In at least one example, the HMD 1-100 may also include a curtain 1-124 that occludes the second opening 1-154 between the housing 1-150 and the rear display assemblies 1-120a-b. In at least one example, the curtain 1-124 can be elastic or at least partially elastic.

[0074] Figure 1B and Figure 1C Any one of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination in Figures 1D to 1F any other example of the devices, features, components, and parts shown and described herein. Similarly, with reference to Figures 1D to 1F any one of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination in Figure 1B and Figure 1C the examples of the devices, features, components, and parts shown.

[0075] Figure 1D A disassembled view of an example of the HMD 1-200 including its various parts or components is shown, and these parts or components are separated according to the modularity and selective coupling of these components. For example, the HMD 1-200 may include a strap 1-216, which may be selectively coupled to a first electronic strip 1-205a and a second electronic strip 1-205b. The first fixed strip 1-205a may include a first electronic component 1-212a, and the second fixed strip 1-205b may include a second electronic component 1-212b. In at least one example, the first and second strips 1-205a-b can be removably coupled to the display unit 1-202.

[0076] In addition, the HMD 1-200 may include a light seal 1-210 configured to be removably coupled to the display unit 1-202. The HMD 1-200 may also include a lens 1-218, which may be removably coupled to the display unit 1-202, for example, on a first component including a display screen and a second display component. The lens 1-218 may include a custom prescription lens configured to correct vision. As noted, each of the components shown in the Figure 1D disassembled view and described above can be removably coupled, attached, reattached, and replaced to update components or swap out components for different users. For example, straps such as strap 1-216, light seals such as light seal 1-210, lenses such as lens 1-218, and electronic strips such as electronic strips 1-205a-b can be swapped out according to the user so that these parts are customized to fit and correspond to a single user of the HMD 1-200.

[0077] Figure 1D Any one of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination in Figure 1B , Figure 1C and Figures 1E to 1Fin any other example of the devices, features, components, and parts shown and described herein. Similarly, reference Figure 1B , Figure 1C and Figures 1E to 1F any one of the features, components, and / or parts shown and described (including their arrangements and configurations) may be included individually or in any combination in Figure 1D the example of the devices, features, components, and parts shown.

[0078] Figure 1E FIG. shows an exploded view of an example of the display unit 1-306 of the HMD. The display unit 1-306 may include a front display assembly 1-308, a frame / housing assembly 1-350, and a curtain assembly 1-324. The display unit 1-306 may also include a sensor assembly 1-356, a logic board assembly 1-358, and a cooling assembly 1-360 disposed between the frame assembly 1-350 and the front display assembly 1-308. In at least one example, the display unit 1-306 may also include a rear display assembly 1-320, which includes a first rear display screen 1-322a and a second rear display screen 1-322b disposed between the frame 1-350 and the curtain assembly 1-324.

[0079] In at least one example, the display unit 1-306 may also include a motor assembly 1-362 configured as an adjustment mechanism for adjusting the positioning of the display screens 1-322a to b of the display assembly 1-320 relative to the frame 1-350. In at least one example, the display assembly 1-320 is mechanically coupled to the motor assembly 1-362, and each display screen 1-322a to b has at least one motor such that the motor can translate the display screens 1-322a to b to match the pupil spacing of the user's eyes.

[0080] In at least one example, the display unit 1-306 may include a dial or button 1-328 that can be pressed relative to the frame 1-350 and accessed by a user external to the frame 1-350. The button 1-328 may be electrically connected to the motor assembly 1-362 via a controller such that the button 1-328 can be manipulated by the user to cause the motors of the motor assembly 1-362 to adjust the positioning of the display screens 1-322a to b.

[0081] Figure 1E any one of the features, components, and / or parts shown (including their arrangements and configurations) may be included individually or in any combination in Figures 1B to 1D and Figure 1F any other example of the devices, features, components, and parts shown and described herein. Similarly, reference Figures 1B to 1D and Figure 1FAny one of the features, components, and / or parts shown and described (including their arrangements and configurations) may be included individually or in any combination in Figure 1E the examples of the devices, features, components, and parts shown.

[0082] Figure 1F A disassembled view of another example of the display unit 1-406 of an HMD device similar to other HMD devices described herein is shown. The display unit 1-406 may include a front display assembly 1-402, a sensor assembly 1-456, a logic board assembly 1-458, a cooling assembly 1-460, a frame assembly 1-450, a rear display assembly 1-421, and a curtain assembly 1-424. The display unit 1-406 may also include a motor assembly 1-462 for adjusting the positioning of a first display sub-assembly 1-420a and a second display sub-assembly 1-420b of the rear display assembly 1-421, including a first corresponding display screen and a second corresponding display screen for inter-pupillary adjustment, as described above.

[0083] Reference is made herein to Figures 1B to 1E and the subsequent figures referred to in this disclosure to describe in more detail Figure 1F the various parts, systems, and components shown in the disassembled view. Figure 1F The display unit 1-406 shown may be assembled and integrated with Figures 1B to 1E the shown fixing mechanism, which includes an electronic strip, a belt, and other components (including a light seal, a connection assembly, etc.).

[0084] Figure 1F Any one of the features, components, and / or parts shown and described (including their arrangements and configurations) may be included individually or in any combination in Figures 1B to 1E any other example among the other examples of the devices, features, components, and parts shown and described herein. Similarly, reference is made to Figures 1B to 1E Any one of the features, components, and / or parts shown and described (including their arrangements and configurations) may be included individually or in any combination in Figure 1F the examples of the devices, features, components, and parts shown.

[0085] Figure 1G A disassembled perspective view of the front cover assembly 3-100 of the HMD device described herein (e.g., Figure 1G the front cover of the HMD 3-100 shown or the front cover assembly 3-1 of any other HMD device shown and described herein) is shown. Figure 1GThe front cover assembly 3-100 shown may include a transparent or translucent cover 3-102, a shield 3-104 (or "cover"), an adhesive layer 3-106, a display assembly 3-108 including a lenticular lens panel or array 3-110, and a structural trim 3-112. The adhesive layer 3-106 may fix the shield 3-104 and / or the transparent cover 3-102 to the display assembly 3-108 and / or the trim 3-112. The trim 3-112 may fix the various components of the front cover assembly 3-100 to the frame or base of the HMD device.

[0086] In at least one example, as Figure 1G shown, the transparent cover 3-102, the shield 3-104, and the display assembly 3-108 including the lenticular lens array 3-110 may be bent to conform to the curvature of the user's face. The transparent cover 3-102 and the shield 3-104 may be bent in two or three dimensions, for example, bent vertically in the Z direction inside and outside the Z-X plane, and bent horizontally in the X direction inside and outside the Z-X plane. In at least one example, the display assembly 3-108 may include a lenticular lens array 3-110 and a display panel having pixels configured to project light through the shield 3-104 and the transparent cover 3-102. The display assembly 3-108 may be bent in at least one direction (e.g., the horizontal direction) to conform to the curvature of the user's face from one side (e.g., the left side) to the other side (e.g., the right side) of the face. In at least one example, each layer or component of the display assembly 3-108 (which will be shown and described in more detail in subsequent figures, but which may include the lenticular lens array 3-110 and the display layer) may be bent similarly or concentrically in the horizontal direction to conform to the curvature of the user's face.

[0087] In at least one example, the shield 3-104 may include a transparent or translucent material through which the display assembly 3-108 projects light. In one example, the shield 3-104 may include one or more opaque portions, such as an opaque ink printed portion or other opaque film portion on the back surface of the shield 3-104. When the HMD device is worn, the rear surface may be the surface of the shield 3-104 facing the user's eyes. In at least one example, the opaque portion may be on the front surface of the shield 3-104 opposite the rear surface. In at least one example, one or more opaque portions of the shield 3-104 may include a peripheral portion that visually hides any components around the outer periphery of the display screen of the display assembly 3-108. In this way, the opaque portions of the shield hide any other components of the HMD device that would otherwise be visible through the transparent or translucent cover 3-102 and / or the shield 3-104, including electronic components, structural components, etc.

[0088] In at least one example, the shield 3-104 may define one or more apertured transparent portions 3-120 through which the sensor may transmit and receive signals. In one example, portion 3-120 is an aperture through which the sensor may extend or through which the sensor may transmit and receive signals. In one example, portion 3-120 is a transparent portion, or a portion that is more transparent than the surrounding translucent or opaque portion of the shield, through which the sensor may transmit and receive signals through the shield and through the transparent cover 3-102. In one example, the sensor may include a camera, an IR sensor, a LUX sensor, or any other visual or non-visual environmental sensor of the HMD device.

[0089] Figure 1G Any one of the features, components, and / or parts shown (including their arrangement and configuration) may be included, either alone or in any combination, in any other example of the devices, features, components, and parts described herein. Similarly, any one of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included, either alone or in any combination, in Figure 1G the examples of the devices, features, components, and parts shown.

[0090] Figure 1H An exploded view of an example of an HMD device 6-100 is shown. The HMD device 6-100 may include a sensor array or system 6-102 that includes one or more sensors, cameras, projectors, etc. mounted to one or more components of the HMD 6-100. In at least one example, the sensor system 6-102 may include a bracket 1-338 to which one or more sensors of the sensor system 6-102 may be fixed / fastened.

[0091] Figure 1I A portion of the HMD device 6-100 including the front transparent cover 6-104 and the sensor system 6-102 is shown. The sensor system 6-102 may include a plurality of different sensors, transmitters, receivers, including cameras, IR sensors, projectors, etc. The transparent cover 6-104 is shown in front of the sensor system 6-102 to illustrate the relative positioning of the various sensors and transmitters and the orientation of each sensor / transmitter of the system 6-102. As used herein, the terms "beside", "side", "lateral", "horizontal", and other similar terms refer to the orientation or direction indicated by the Figure 1J X-axis as shown. Terms such as "vertical", "upward", "downward", and similar terms refer to the Figure 1JThe orientation or direction indicated by the Z-axis as shown. Terms such as "frontward", "rearward", "forward", "backward" and similar terms refer to the orientation or direction indicated by the Y-axis as shown by Figure 1J shown.

[0092] In at least one example, the transparent cover 6-104 may define the front outer surface of the HMD device 6-100, and the sensor system 6-102 including various sensors and their components may be disposed behind the cover 6-104 in the Y-axis / direction. The cover 6-104 may be transparent or translucent to allow light to pass through the cover 6-104, including both the light detected by the sensor system 6-102 and the light emitted therefrom.

[0093] As described elsewhere herein, the HMD device 6-100 may include one or more controllers that include processors for electrically coupling the various sensors and transmitters of the sensor system 6-102 to one or more motherboards, processing units, and other electronic devices such as display screens. Additionally, as will be shown in more detail with reference to other figures below, the various sensors, transmitters, and other components of the sensor system 6-102 may be coupled to various structural frame members, brackets, etc. of the HMD device 6-100 that are not shown in Figure 1I For illustrative clarity, Figure 1I the components of the sensor system 6-102 are shown unattached and unelectrically coupled to other components.

[0094] In at least one example, the device may include one or more controllers having processors configured to execute instructions stored on a memory component electrically coupled to the processors. The instructions may include or cause the processors to execute one or more algorithms for self-correcting the angles and positions of the various cameras described herein as the initial position, angle, or orientation of the camera changes over time due to accidental drop events or other events that cause collisions or deformations.

[0095] In at least one example, the sensor system 6-102 can include one or more scene cameras 6-106. The system 6-102 can include two scene cameras 6-102, respectively disposed on both sides of the bridge or arch structure of the HMD device 6-100, such that each of the two cameras 6-106 generally corresponds to the positioning of the user's left and right eyes behind the cover 6-103. In at least one example, the scene cameras 6-106 are generally oriented forward in the Y direction to capture images in front of the user during use of the HMD 6-100. In at least one example, the scene cameras are color cameras and, when the HMD device 6-100 is used, provide images and content for MR video passthrough to the display screen facing the user's eyes. The scene cameras 6-106 can also be used for environment and object reconstruction.

[0096] In at least one example, the sensor system 6-102 can include a first depth sensor 6-108 that is generally pointed forward in the Y direction. In at least one example, the first depth sensor 6-108 can be used for environment and object reconstruction and for tracking the user's hands and body. In at least one example, the sensor system 6-102 can include a second depth sensor 6-110 that is centered along the width of the HMD device 6-100 (e.g., along the X axis). For example, the second depth sensor 6-110 can be disposed above the central bridge of the nose or on an adapter structure above the nose when the user wears the HMD 6-100. In at least one example, the second depth sensor 6-110 can be used for environment and object reconstruction and for hand and body tracking. In at least one example, the second depth sensor can include a LIDAR sensor.

[0097] In at least one example, the sensor system 6-102 can include a depth projector 6-112 that is generally oriented forward to project electromagnetic waves (e.g., in the form of a pre-determined pattern of light points) into the field of view or within the field of view of the user and / or the scene cameras 6-106, or into a field of view that includes and extends beyond the field of view of the user and / or the scene cameras 6-106. In at least one example, the depth projector is capable of projecting electromagnetic waves of light in the form of a pattern of light points that are reflected from objects and return to the aforementioned depth sensors, including depth sensors 6-108, 6-110. In at least one example, the depth projector 6-112 can be used for environment and object reconstruction and for hand and body tracking.

[0098] In at least one example, the sensor system 6-102 may include a downward-facing camera 6-114, the field of view of which generally points downward with respect to the HDM device 6-100 on the Z axis. In at least one example, the downward camera 6-114 may be disposed on the left and right sides of the HMD device 6-100 as shown in the figure and is used for hand and body tracking, headset tracking, and face avatar detection and creation for displaying a user avatar on the forward display screen of the HMD device 6-100 described elsewhere herein. For example, the downward camera 6-114 may be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the cheeks, mouth, and chin.

[0099] In at least one example, the sensor system 6-102 may include a jaw camera 6-116. In at least one example, the jaw camera 6-116 may be disposed on the left and right sides of the HMD device 6-100 as shown in the figure and is used for hand and body tracking, headset tracking, and face avatar detection and creation for displaying a user avatar on the forward display screen of the HMD device 6-100 described elsewhere herein. For example, the jaw camera 6-116 may be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the user's jaw, cheeks, mouth, and chin. For hand and body tracking, headset tracking, and face avatar

[0100] In at least one example, the sensor system 6-102 may include a side camera 6-118. The side camera 6-118 may be oriented to capture left and right views in the X axis or with respect to the direction of the HMD device 6-100. In at least one example, the side camera 6-118 may be used for hand and body tracking, headset tracking, and face avatar detection and recreation.

[0101] In at least one example, the sensor system 6-102 may include a plurality of eye tracking and gaze tracking sensors for determining the identity, status, and gaze direction of the user's eyes during and / or before use. In at least one example, the eye / gaze tracking sensors may include a nose-eye camera 6-120, which is disposed on either side of the user's nose and adjacent to the user's nose when wearing the HMD device 6-100. The eye / gaze sensors may also include a bottom eye camera 6-122 disposed below the corresponding user's eye for capturing an image of the eye for face avatar detection and creation, gaze tracking, and iris identification functions.

[0102] In at least one example, the sensor system 6-102 can include an infrared illuminator 6-124 that points outward from the HMD device 6-100 to illuminate the external environment and any objects therein with IR light for IR detection using one or more IR sensors of the sensor system 6-102. In at least one example, the sensor system 6-102 can include a flicker sensor 6-126 and an ambient light sensor 6-128. In at least one example, the flicker sensor 6-126 can detect the top light refresh rate to avoid display flicker. In one example, the infrared illuminator 6-124 can include light-emitting diodes and can be particularly used in low-light environments to illuminate the user's hand and other objects in low light for detection by the infrared sensors of the sensor system 6-102.

[0103] In at least one example, multiple sensors (including the scene camera 6-106, the downward camera 6-114, the jaw camera 6-116, the side camera 6-118, the depth projector 6-112, and the depth sensors 6-108, 6-110) can be used in combination with an electrically coupled controller to combine depth data with camera data for hand tracking and for size determination to better perform the hand tracking and object recognition and tracking functions of the HMD device 6-100. In at least one example, the downward camera 6-114, the jaw camera 6-116, and the side camera 6-118 described above and shown in Figure 1I can be wide-angle cameras capable of operating in both the visible and infrared spectra. In at least one example, these cameras 6-114, 6-116, 6-118 can operate only in black-and-white light detection to simplify image processing and obtain sensitivity.

[0104] Figure 1I Any of the features, components, and / or parts shown (including their arrangement and configuration) can be included individually or in any combination in Figures 1J to 1L any other example of the devices, features, components, and parts shown and described herein. Similarly, any of the features, components, and / or parts shown and described with reference to Figures 1J to 1L can be included individually or in any combination in Figure 1I the examples of the devices, features, components, and parts shown.

[0105] Figure 1JA lower perspective view of an example of an HMD 6-200 including a cover or shield 6-204 fixed to a frame 6-230 is shown. In at least one example, sensors 6-203 of a sensor system 6-202 may be disposed around the perimeter of the HDM 6-200 such that the sensors 6-203 are disposed outwardly around the perimeter of a display area or region 6-232 so as not to obstruct viewing of the displayed light. In at least one example, the sensors may be disposed behind the shield 6-204 and aligned with a transparent portion of the shield, thereby allowing the sensors and projectors to allow light to pass back and forth through the shield 6-204. In at least one example, an opaque ink or other opaque material or film / layer may be disposed on the shield 6-204 around the display area 6-232 to hide components of the HMD 6-200 outside of the display area 6-232 rather than a transparent portion defined by an opaque portion through which the sensors and projectors transmit and receive light and electromagnetic signals during operation. In at least one example, the shield 6-204 allows light to pass through from a display (e.g., within the display area 6-232), but does not allow light to pass radially outward from the display area around the perimeter of the display and the shield 6-204.

[0106] In some examples, the shield 6-204 includes a transparent portion 6-205 and an opaque portion 6-207, as described above and elsewhere herein. In at least one example, the opaque portion 6-207 of the shield 6-204 may define one or more transparent regions 6-209 through which sensors 6-203 of the sensor system 6-202 may transmit and receive signals. In the example shown, the sensors 6-203 of the sensor system 6-202 that transmit and receive signals through the shield 6-204, or more specifically through the transparent regions 6-209 of (or defined by) the opaque portion 6-207 of the shield 6-204 may include sensors that are the same as or similar to those shown in the examples of Figure 1I such as depth sensors 6-108 and 6-110, depth projectors 6-112, a first scene camera and a second scene camera 6-106, a first downward camera and a second downward camera 6-114, a first side camera and a second side camera 6-118, and a first infrared illuminator and a second infrared illuminator 6-124. These sensors are also shown in the examples of Figure 1K and Figure 1L Other sensors, sensor types, sensor quantities, and their relative positioning may be included in one or more other examples of the HMD.

[0107] Figure 1J Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination in Figure 1I and Figures 1K to 1Lin any other examples of the devices, features, components, and parts shown and described herein. Similarly, reference Figure 1I and Figures 1K to 1L any one of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination in Figure 1J the examples of the devices, features, components, and parts shown.

[0108] Figure 1K A front view of a portion of an example of an HMD device 6-300 including a display 6-334, brackets 6-336, 6-338, and a frame or housing 6-330 is shown. Figure 1K The example shown does not include a front cover or shield to show the brackets 6-336, 6-338. For example, Figure 1J the shield 6-204 shown includes an opaque portion 6-207 that will visually cover / block the view of anything outside the display / display area 6-334 (e.g., radially / peripherally outside the display / display area), including sensors 6-303 and bracket 6-338.

[0109] In at least one example, various sensors of the sensor system 6-302 are coupled to brackets 6-336, 6-338. In at least one example, the scene cameras 6-306 include tight tolerances on the angles relative to each other. For example, the tolerance on the mounting angle between two scene cameras 6-306 can be 0.5 degrees or less, e.g., 0.3 degrees or less. To achieve and maintain such tight tolerances, in one example, the scene cameras 6-306 may be mounted to bracket 6-338 instead of the shield. The bracket may include a cantilever on which the scene cameras 6-306 and other sensors of the sensor system 6-302 may be mounted to maintain their position and orientation unchanged in the event of a drop event that causes any deformation of other brackets 6-226, housing 6-330, and / or shield by the user.

[0110] Figure 1K any one of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination in Figures 1I to 1J and Figure 1L any other examples of the devices, features, components, and parts shown and described herein. Similarly, reference Figures 1I to 1J and Figure 1L any one of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination in Figure 1K the examples of the devices, features, components, and parts shown.

[0111] Figure 1LShows a bottom view of an example of an HMD 6-400 including a front display / cover assembly 6-404 and a sensor system 6-402. The sensor system 6-402 can be similar to other sensor systems described above and elsewhere in this document (including reference Figures 1I to 1K ). In at least one example, the chin camera 6-416 can face downward to capture images of the user's lower facial features. In one example, the chin camera 6-416 can be directly coupled to the frame or housing 6-430 or one or more internal brackets that are directly coupled to the shown frame or housing 6-430. The frame or housing 6-430 can include one or more holes / openings 6-415 through which the chin camera 6-416 can send and receive signals.

[0112] Figure 1L Any of the features, components, and / or parts shown (including their arrangement and configuration) can be included individually or in any combination in Figures 1I to 1K any other example of the devices, features, components, and parts shown and described herein. Similarly, reference Figures 1I to 1K any of the features, components, and / or parts shown and described (including their arrangement and configuration) can be included individually or in any combination in Figure 1L the example of the devices, features, components, and parts shown.

[0113] Figure 1M Shows a rear perspective view of an interpupillary distance (IPD) adjustment system 11.1.1-102 that includes a first optical module and a second optical module 11.1.1-104a to b that are slidably engaged / coupled to respective guide rods 11.1.1-108a to b and motors 11.1.1-110a to b of a left adjustment subsystem and a right adjustment subsystem 11.1.1-106a to b. The IPD adjustment system 11.1.1-102 can be coupled to a bracket 11.1.1-112 and includes buttons 11.1.1-114 that are in electrical communication with the motors 11.1.1-110a to b. In at least one example, the buttons 11.1.1-114 can be in electrical communication with the first motor and the second motor 11.1.1-110a to b via a processor or other circuit components such that the first motor and the second motor 11.1.1-110a to b are activated and cause the first optical module and the second optical module 11.1.1-104a to b to change their positions relative to each other.

[0114] In at least one example, the first and second optical modules 11.1.1-104a-b may include respective display screens configured to project light toward a user's eyes when wearing the HMD 11.1.1-100. In at least one example, a user may manipulate (e.g., press and / or rotate) buttons 11.1.1-114 to activate position adjustment of the optical modules 11.1.1-104a-b to match the pupil spacing of the user's eyes. The optical modules 11.1.1-104a-b may also include one or more cameras or other sensor / sensor systems for imaging and measuring the user's IPD such that the optical modules 11.1.1-104a-b may be adjusted to match the IPD.

[0115] In one example, a user may manipulate buttons 11.1.1-114 to cause automatic position adjustment of the first and second optical modules 11.1.1-104a-b. In one example, a user may manipulate buttons 11.1.1-114 to cause a manual adjustment such that the optical modules 11.1.1-104a-b move farther or closer (e.g., when the user rotates buttons 11.1.1-114 one way or the other) until the user visually matches her / his own IPD. In one example, the manual adjustment is communicated electronically via one or more circuits and power for moving the optical modules 11.1.1-104a-b via motors 11.1.1-110a-b is provided by a power source. In one example, adjustment and movement of the optical modules 11.1.1-104a-b via manipulation of buttons 11.1.1-114 is mechanically actuated via movement of buttons 11.1.1-114.

[0116] Figure 1M Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included, either alone or in any combination, in any other example of the devices, features, components, and parts shown in any other of the other illustrated figures and described herein. Similarly, any of the features, components, and / or parts shown, described, and herein described in connection with any other of the other illustrated figures (including their arrangement and configuration) may be included, either alone or in any combination, in Figure 1M the examples of the devices, features, components, and parts shown.

[0117] Figure 1N A front perspective view of a portion of the HMD 11.1.2-100 is shown, which includes an outer structural frame 11.1.2-102 and an inner or intermediate structural frame 11.1.2-104 defining a first aperture and a second aperture 11.1.2-106a, 11.1.2-106b. The apertures 11.1.2-106a-b are in Figure 1Nis shown in dashed lines because viewing of the holes 11.1.2-106a to b may be blocked by one or more other components of the HMD 11.1.2-100 that are coupled to the inner frame 11.1.2-104 and / or the outer frame 11.1.2-102, as shown. In at least one example, the HMD 11.1.2-100 may include a first mounting bracket 11.1.2-108 that is coupled to the inner frame 11.1.2-104. In at least one example, the mounting bracket 11.1.2-108 is coupled to the inner frame 11.1.2-104 between the first and second holes 11.1.2-106a to b.

[0118] The mounting bracket 11.1.2-108 may include an intermediate or central portion 11.1.2-109 that is coupled to the inner frame 11.1.2-104. In some examples, the intermediate or central portion 11.1.2-109 may not be the geometric middle or center of the bracket 11.1.2-108. Instead, the intermediate / central portion 11.1.2-109 may be disposed between a first cantilevered extension arm and a second cantilevered extension arm that extend away from the intermediate portion 11.1.2-109. In at least one example, the mounting bracket 108 includes a first cantilever 11.1.2-112 and a second cantilever 11.1.2-114 that extend away from the intermediate portion 11.1.2-109 of the mounting bracket 11.1.2-108 that is coupled to the inner frame 11.1.2-104.

[0119] As Figure 1N shown, the outer frame 11.1.2-102 may define a curved geometry on its lower side to accommodate the user's nose when the user wears the HMD 11.1.2-100. The curved geometry may be referred to as a nose bridge 11.1.2-111 and is shown centered on the lower side of the HMD 11.1.2-100 as shown. In at least one example, the mounting bracket 11.1.2-108 may be connected to the inner frame 11.1.2-104 between the holes 11.1.2-106a to b such that the cantilevers 11.1.2-112, 11.1.2-114 extend downward and laterally outward away from the intermediate portion 11.1.2-109 to be complementary to the geometry of the nose bridge 11.1.2-111 of the outer frame 11.1.2-102. In this way, the mounting bracket 11.1.2-108 is configured to accommodate the user's nose, as described above. The geometry of the nose bridge 11.1.2-111 accommodates the nose because the nose bridge 11.1.2-111 provides a curvature that conforms to the shape of the user's nose, providing a comfortable fit from above, over, and around.

[0120] The first cantilever 11.1.2-112 can extend away from the middle part 11.1.2-109 of the mounting bracket 11.1.2-108 in a first direction, and the second cantilever 11.1.2-114 can extend away from the middle part 11.1.2-109 of the mounting bracket 11.1.2-108 in a second direction opposite to the first direction. The first cantilever 11.1.2-112 and the second cantilever 11.1.2-114 are referred to as "cantilevered" or "cantilever" arms because each arm 11.1.2-112, 11.1.2-114 respectively includes free distal ends 11.1.2-116, 11.1.2-118 that are not attached to the inner frame 11.1.2-102 and the outer frame 11.1.2-104. In this way, the arms 11.1.2-112, 11.1.2-114 overhang from the middle part 11.1.2-109, which can be connected to the inner frame 11.1.2-104, while the distal ends 11.1.2-102, 11.1.2-104 are not attached.

[0121] In at least one example, the HMD 11.1.2-100 can include one or more components coupled to the mounting bracket 11.1.2-108. In one example, the components include a plurality of sensors 11.1.2-110a-f. Each of the plurality of sensors 11.1.2-110a-f can include various types of sensors, including cameras, IR sensors, etc. In some examples, one or more of the sensors 11.1.2-110a-f can be used for object recognition in three-dimensional space, such that it is important to maintain the precise relative positioning of two or more of the plurality of sensors 11.1.2-110a-f. The cantilevered nature of the mounting bracket 11.1.2-108 can protect the sensors 11.1.2-110a-f from damage and misalignment in the event of an accidental drop by the user. Since the sensors 11.1.2-110a-f are cantilevered on the arms 11.1.2-112, 11.1.2-114 of the mounting bracket 11.1.2-108, the stress and deformation of the inner frame and / or the outer frame 11.1.2-104, 11.1.2-102 are not transmitted to the cantilevers 11.1.2-112, 11.1.2-114 and thus do not affect the relative positions of the sensors 11.1.2-110a-f coupled / mounted to the mounting bracket 11.1.2-108.

[0122] Figure 1NAny one of the features, components, and / or parts shown (including their arrangement and configuration) can be included, either individually or in any combination, in any other example of the devices, features, components, and other examples described herein. Similarly, any one of the features, components, and / or parts shown and described herein (including their arrangement and configuration) can be included, either individually or in any combination, in Figure 1N the examples of the devices, features, components, and parts shown.

[0123] Figure 1O An example of an optical module 11.3.2 - 100 for use in an electronic device (such as an HMD, including the HDM devices described herein) is shown. As shown in one or more other examples described herein, the optical module 11.3.2 - 100 can be one of two optical modules within the HMD, where each optical module is aligned to project light towards the user's eyes. In this way, the first optical module can project light towards the user's first eye via a display screen, and the second optical module of the same device can project light towards the user's second eye via another display screen.

[0124] In at least one example, the optical module 11.3.2 - 100 can include an optical frame or housing 11.3.2 - 102, which can also be referred to as a barrel or an optical module barrel. The optical module 11.3.2 - 100 can also include a display 11.3.2 - 104 coupled to the housing 11.3.2 - 102, the display including one or more display screens. The display 11.3.2 - 104 can be coupled to the housing 11.3.2 - 102 such that the display 11.3.2 - 104 is configured to project light towards the user's eyes when wearing the HMD to which the display module 11.3.2 - 100 belongs during use. In at least one example, the housing 11.3.2 - 102 can surround the display 11.3.2 - 104 and provide connection features for coupling other components of the optical module described herein.

[0125] In one example, the optical module 11.3.2-100 may include one or more cameras 11.3.2-106 coupled to a housing 11.3.2-102. The cameras 11.3.2-106 may be positioned relative to a display 11.3.2-104 and the housing 11.3.2-102 such that the cameras 11.3.2-106 are configured to capture one or more images of a user's eyes during use. In at least one example, the optical module 11.3.2-100 may further include a light strip 11.3.2-108 surrounding the display 11.3.2-104. In one example, the light strip 11.3.2-108 is disposed between the display 11.3.2-104 and the cameras 11.3.2-106. The light strip 11.3.2-108 may include a plurality of lights 11.3.2-110. The plurality of lights may include one or more light-emitting diodes (LEDs) or other lights configured to project light toward the user's eyes when wearing the HMD. Each light 11.3.2-110 in the light strip 11.3.2-108 may be spaced apart around the light strip 11.3.2-108 and thus may be spaced apart evenly or unevenly around the display 11.3.2-104 at various positions on the light strip 11.3.2-108 and around the display 11.3.2-104.

[0126] In at least one example, the housing 11.3.2-102 defines a viewing opening 11.3.2-101 through which a user may view the display 11.3.2-104 when wearing the HMD device. In at least one example, the LEDs are configured and arranged to emit light through the viewing opening 11.3.2-101 onto the user's eyes. In one example, the cameras 11.3.2-106 are configured to capture one or more images of the user's eyes through the viewing opening 11.3.2-101.

[0127] As described above, Figure 1O Each of the components and features of the illustrated optical module 11.3.2-100 may be replicated in another (e.g., second) optical module provided with the HMD to interact (e.g., project light and capture images) with the user's other eye.

[0128] Figure 1O Any one of the illustrated features, components, and / or parts (including their arrangement and configuration) may be included individually or in any combination in Figure 1P any other example of the devices, features, components, and parts shown or otherwise described herein. Similarly, reference Figure 1P to any one of the illustrated and described or otherwise described features, components, and / or parts (including their arrangement and configuration) may be included individually or in any combination inFigure 1O in the examples of the devices, features, components, and parts shown.

[0129] Figure 1P A cross-sectional view of an example of an optical module 11.3.2-200 is shown, which includes a housing 11.3.2-202, a display assembly 11.3.2-204 coupled to the housing 11.3.2-202, and a lens 11.3.2-216 coupled to the housing 11.3.2-202. In at least one example, the housing 11.3.2-202 defines a first hole or passage 11.3.2-212 and a second hole or passage 11.3.2-214. The passages 11.3.2-212, 11.3.2-214 may be configured to slidably engage corresponding tracks or guide rods of the HMD device to allow the optical module 11.3.2-200 to adjust its position relative to the user's eyes to match the user's interpupillary distance (IPD). The housing 11.3.2-202 is capable of slidably engaging the guide rods to fix the optical module 11.3.2-200 in place within the HMD.

[0130] In at least one example, the optical module 11.3.2-200 may further include a lens 11.3.2-216 coupled to the housing 11.3.2-202 and disposed between the display assembly 11.3.2-204 and the user's eyes when the HMD is worn. The lens 11.3.2-216 may be configured to direct light from the display assembly 11.3.2-204 to the user's eyes. In at least one example, the lens 11.3.2-216 may be part of a lens assembly that includes a corrective lens removably attached to the optical module 11.3.2-200. In at least one example, the lens 11.3.2-216 is disposed above the light bar 11.3.2-208 and one or more eye tracking cameras 11.3.2-206 such that the cameras 11.3.2-206 are configured to capture images of the user's eyes through the lens 11.3.2-216, and the light bar 11.3.2-208 includes lights configured to project light through the lens 11.3.2-216 onto the user's eyes during use.

[0131] Figure 1P Any of the features, components, and / or parts shown (including their arrangements and configurations) may be included, either alone or in any combination, in any other examples of the devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangements and configurations) may be included, either alone or in any combination, in Figure 1P the examples of the devices, features, components, and parts shown.

[0132] Figure 2FIG. 0 is a block diagram of an example of controller 110 according to some embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and to not obscure more relevant aspects of the embodiments disclosed herein. To that end, as a non-limiting example, in some embodiments, controller 110 includes one or more processing units 202 (e.g., microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., universal serial bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, global system for mobile communications (GSM), code division multiple access (CDMA), time division multiple access (TDMA), global positioning system (GPS), infrared (IR), Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these components and various other components.

[0133] In some embodiments, one or more communication buses 204 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0134] Memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes non-transitory computer readable storage medium. In some embodiments, memory 220 or the non-transitory computer readable storage medium of memory 220 stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and an XR experience module 240.

[0135] The operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate single or multiple XR experiences of one or more users (e.g., single XR experiences of one or more users, or multiple XR experiences of corresponding groups of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.

[0136] In some embodiments, the data acquisition unit 241 is configured to obtain data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 of Figure 1A and optionally from one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.

[0137] In some embodiments, the tracking unit 242 is configured to map the scene 105 and track the positioning / position of at least the display generation component 120 relative to Figure 1A the scene 105, and optionally track the position of one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the positioning / position of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to Figure 1A the scene 105, relative to the display generation component 120, and / or relative to a coordinate system (which is defined relative to the user's hand). The hand tracking unit 244 is described in more detail below with respect to Figure 4 . In some embodiments, the eye tracking unit 243 is configured to track the positioning or movement of the user's gaze (or more generally, the user's eyes, face, or head) relative to the scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hand)) or relative to the XR content displayed via the display generation component 120. The eye tracking unit 243 is described in more detail below with respect to Figure 5 .

[0138] In some embodiments, the coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by the display generation component 120, and optionally by one or more of the output device 155 and / or the peripheral device 195. To this end, in various embodiments, the coordination unit 246 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0139] In some embodiments, the data sending unit 248 is configured to send data (e.g., presentation data, location data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data sending unit 248 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0140] Although the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 are shown as residing on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 may be located in separate computing devices.

[0141] In addition, Figure 2 More serves as a functional description of the various features that may be present in a particular implementation, different from the structural schematic diagrams of the embodiments described herein. As will be recognized by those of ordinary skill in the art, the items shown separately may be combined, and some items may be separated. For example, Figure 2 Some of the functional modules shown separately in may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the specific division of functions, as well as how the features are distributed therein, will vary depending on the specific implementation, and in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for the specific implementation.

[0142] Figure 3FIG. 0 is a block diagram of an example of a display generation component 120 in accordance with some embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and to not obscure more relevant aspects of the embodiments disclosed herein. To that end, as a non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., a microprocessor, an ASIC, an FPGA, a GPU, a CPU, a processing core, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional internal and / or external image sensors 314, a memory 320, and one or more communication buses 304 for interconnecting these components and various other components.

[0143] In some embodiments, one or more communication buses 304 include circuitry for interconnecting and controlling communication between the system components. In some embodiments, one or more I / O devices and sensors 306 include an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, and / or one or more depth sensors (e.g., structured light, time-of-flight, etc.).

[0144] In some embodiments, one or more XR displays 312 are configured to provide a user with an XR experience. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more XR displays 312 correspond to diffractive, reflective, polarization, holographic, and other waveguide displays. For example, the display generation component 120 (e.g., an HMD) includes a single XR display. In another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, one or more XR displays 312 are capable of presenting MR and VR content. In some embodiments, one or more XR displays 312 are capable of presenting MR or VR content.

[0145] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand and optionally the user's arm (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to the scene that the user would see in the absence of the display generation component 120 (e.g., an HMD) (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.

[0146] The memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, the memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. The memory 320 includes non-transitory computer-readable storage medium. In some embodiments, the memory 320 or the non-transitory computer-readable storage medium of the memory 320 stores the following programs, modules, and data structures or subsets thereof, including an optional operating system 330 and an XR rendering module 340.

[0147] The operating system 330 includes procedures for handling various basic system services and for performing hardware-related tasks. In some embodiments, the XR rendering module 340 is configured to present XR content to a user via one or more XR displays 312. To this end, in various embodiments, the XR rendering module 340 includes a data acquisition unit 342, an XR rendering unit 344, an XR mapping generation unit 346, and a data transmission unit 348.

[0148] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, positioning data, etc.) at least from Figure 1A the controller 110. To this end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for instructions and heuristics and metadata for heuristics.

[0149] In some embodiments, the XR rendering unit 344 is configured to present XR content via one or more XR displays 312. To this end, in various embodiments, the XR rendering unit 344 includes instructions and / or logic for instructions and heuristics and metadata for heuristics.

[0150] In some embodiments, the XR mapping generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate extended reality) based on media content data. To this end, in various embodiments, the XR mapping generation unit 346 includes instructions and / or logic for instructions and heuristics and metadata for heuristics.

[0151] In some embodiments, the data sending unit 348 is configured to send data (e.g., presentation data, location data, etc.) to at least the controller 110 and optionally to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data sending unit 348 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.

[0152] Although the data acquisition unit 342, the XR presentation unit 344, the XR mapping generation unit 346, and the data sending unit 348 are shown as residing on a single device (e.g., Figure 1A the display generation component 120 of ), it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR presentation unit 344, the XR mapping generation unit 346, and the data sending unit 348 may be located in separate computing devices.

[0153] In addition, Figure 3 This is more of a functional description of the various features that may be present in a particular embodiment, as opposed to a schematic diagram of the structure of the embodiments described herein. As will be recognized by those of ordinary skill in the art, items shown separately may be combined and some items may be separated. For example, Figure 3 some of the functional modules shown separately in may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the specific functional partitioning and how features are allocated therein will vary depending on the particular implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for the particular implementation.

[0154] Figure 4 is a schematic illustration of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 ( Figure 1A ) is controlled by the hand tracking unit 244 ( Figure 2 ) to track the positioning / position of one or more parts of the user's hand and / or the orientation of one or more parts of the user's hand relative to Figure 1AScene 105 (e.g., movement relative to a portion of the physical environment around the user, relative to the display generation component 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system that is defined relative to the user's hand). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0155] In some embodiments, the hand tracking device 140 includes an image sensor 404 that captures three-dimensional scene information of at least the hand 406 of a human user (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.). The image sensor 404 captures hand images at a sufficient resolution to distinguish the fingers and their corresponding positions. The image sensor 404 typically captures images of other parts of the user's body and may also or possibly capture images of all parts of the body, and may have zoom capabilities or a dedicated sensor with increased magnification to capture images of the hand at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in combination with other image sensors to capture the physical environment of the scene 105, or serves as the image sensor for capturing the physical environment of the scene 105. In some embodiments, the image sensor 404 or a portion of its field of view is positioned relative to the user or the user's environment in a manner that defines an interaction space in which hand movements captured by the image sensor are treated as inputs to the controller 110.

[0156] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and in addition, possibly color image data) to the controller 110, and the controller extracts high-level information from the map data. The high-level information is typically provided to an application running on the controller via an application programming interface (API), and the application accordingly drives the display generation component 120. For example, a user can interact with software running on the controller 110 by moving his hand 406 and changing his hand pose.

[0157] In some embodiments, the image sensor 404 projects a speckle pattern onto a scene that includes the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral offset of the speckles in the pattern. This method is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. The method gives the depth coordinates of points in the scene at a particular distance from the image sensor 404 relative to a pre-determined reference plane. In the present disclosure, it is assumed that the image sensor 404 defines an orthogonal set of x, y, and z axes such that the depth coordinate of a point in the scene corresponds to the z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) may use other 3D mapping methods, such as stereoscopic imaging or time-of-flight measurements, based on a single or multiple cameras or other types of sensors.

[0158] In some embodiments, the hand tracking device 140 captures and processes a time series of depth maps that include the user's hand as the user moves his hand (e.g., the entire hand or one or more fingers). Software running on a processor in the image sensor 404 and / or the controller 110 processes the 3D map data to extract image patch descriptors of the hand in these depth maps. The software may match these descriptors with image patch descriptors stored in the database 408 based on a previous learning process in order to estimate the pose of the hand in each frame. The pose generally includes the 3D positions of the user's hand joints and finger tips.

[0159] The software may also analyze the trajectories of the hand and / or fingers over multiple frames in the sequence to identify gestures. The pose estimation function described herein may alternate with the motion tracking function such that the image patch-based pose estimation is only performed once every two (or more) frames, while tracking is used to find the changes in the pose that occur on the remaining frames. Pose, motion, and gesture information is provided to an application running on the controller 110 via the aforementioned API. The program may, for example, move and modify the image presented on the display generation component 120 in response to the pose and / or gesture information, or perform other functions.

[0160] In some embodiments, the gesture includes an air gesture. An air gesture is detected without the user touching an input element (or independent of an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140)) and is based on the detected movement of a part of the user's body (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs) through the air (including the movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one of the user's hands relative to the other of the user's hands, and / or the movement of the user's finger relative to another finger or part of the user's hand), and / or the absolute movement of a part of the user's body (e.g., a tap gesture including the hand moving a predetermined amount and / or speed in a predetermined pose, or a shake gesture including a predetermined speed or amount of rotation of a part of the user's body)).

[0161] In some embodiments, according to some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures for interacting with an XR environment (e.g., a virtual or mixed reality environment) performed by the movement of the user's fingers relative to other fingers or parts of the user's hand. In some embodiments, an air gesture is detected without the user touching an input element that is part of a device (or independent of an input element that is part of a device) and is based on the detected movement of a part of the user's body through the air (including the movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one of the user's hands relative to the other of the user's hands, and / or the movement of the user's finger relative to another finger or part of the user's hand), and / or the absolute movement of a part of the user's body (e.g., a tap gesture including the hand moving a predetermined amount and / or speed in a predetermined pose, or a shake gesture including a predetermined speed or amount of rotation of a part of the user's body)).

[0162] In some embodiments where the input gesture is an air gesture (e.g., in the absence of physical contact with an input device that provides information to a computer system about which user interface element is the target of a user input, such as contact with a user interface element displayed on a touch screen, or contact with a mouse or touchpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Thus, in embodiments involving air gestures, for example, the input gesture is detected in combination with (e.g., simultaneously) the movement of the user's finger and / or hand towards a user interface element while attention (e.g., gaze) towards the user interface element is detected to perform a pinch and / or tap input, as described below.

[0163] In some embodiments, an input gesture that points to a user interface object is performed with direct or indirect reference to the user interface object. For example, an input gesture is performed directly on the user interface object according to a position corresponding to the positioning of the user's hand in a three-dimensional environment relative to the positioning of the user interface object (e.g., as determined based on the user's current viewpoint). In some embodiments, when attention (e.g., gaze) of the user towards the user interface object is detected, an input gesture is performed indirectly on the user interface object according to the position of the user's hand not being at the position corresponding to the positioning of the user interface object in the three-dimensional environment while the user performs the input gesture. For example, for a direct input gesture, the user can initiate the gesture at or near a position corresponding to the display positioning of the user interface object (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or between 0 and 5 cm measured from the outer edge of the option or the central part of the option) to direct the user's input to the user interface object. For an indirect input gesture, the user can direct the user's input to the user interface object by focusing on the user interface object (e.g., by gazing at the user interface object), and while focusing on the option, the user initiates the input gesture (e.g., at any position detectable by the computer system) (e.g., at a position not corresponding to the display positioning of the user interface object).

[0164] In some embodiments, according to some embodiments, the input gestures (e.g., air gestures) used in the various examples and embodiments described herein include pinch inputs and tap inputs for interacting with virtual or mixed reality environments. For example, the pinch inputs and tap inputs described below are performed as air gestures.

[0165] In some embodiments, a pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double pinch gesture. For example, a pinch gesture as an air gesture includes the movement of two or more fingers of a hand to contact each other, i.e., optionally followed by an immediate (e.g., within 0 seconds to 1 second) interruption of contact with each other. A long pinch gesture as an air gesture includes the movement of two or more fingers of a hand contacting each other for at least a threshold amount of time (e.g., at least 1 second) before detecting an interruption of contact with each other. For example, a long pinch gesture includes a user holding a pinch gesture (e.g., where two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between two or more fingers is detected. In some embodiments, a double pinch gesture as an air gesture includes two (e.g., or more) pinch inputs (e.g., performed by the same hand) detected consecutively and immediately (e.g., within a predefined time period) with each other. For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., interrupts contact between two or more fingers), and performs a second pinch input within a predefined time period (e.g., within 1 second or within 2 seconds) after releasing the first pinch input.

[0166] In some embodiments, a pinch-and-drag gesture as an air gesture includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in combination with (e.g., following) a drag input that changes the positioning of the user's hand from a first positioning (e.g., the starting positioning of the drag) to a second positioning (e.g., the ending positioning of the drag). In some embodiments, a user holds a pinch gesture while performing a drag input and releases the pinch gesture (e.g., opens their two or more fingers) to end the drag gesture (e.g., at the second positioning). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., a user pinches two or more fingers to contact each other and moves the same hand to a second positioning in the air using a drag gesture). In some embodiments, the pinch input is performed by a user's first hand and the drag input is performed by a user's second hand (e.g., while the user continues the pinch input with the user's first hand, the user's second hand moves in the air from a first positioning to a second positioning). In some embodiments, an input gesture as an air gesture includes an input performed using both of the user's hands (e.g., a pinch and / or tap input). For example, an input gesture includes two (e.g., or more) pinch inputs performed in combination with each other (e.g., concurrently or within a predefined time period). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch-and-drag input) is performed using the user's first hand, and a second pinch input is performed using the other hand (e.g., the second of the user's two hands) in combination with performing the pinch input using the first hand.

[0167] In some embodiments, a tap input performed as an air gesture (e.g., pointing to a user interface element) includes a movement of the user's finger towards the user interface element, a movement of the user's hand towards the user interface element (optionally, with the user's finger extended towards the user interface element), a downward movement of the user's finger (e.g., mimicking a mouse click movement or a tap on a touch screen), or other predefined movements of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on movement characteristics of the finger or hand that executes the tap gesture movement, which is the movement of the finger or hand away from the user's viewing point and / or towards an object that is the target of the tap input, followed by an end of the movement. In some embodiments, the end of the movement (e.g., the end of the movement away from the user's viewing point and / or towards an object that is the target of the tap input, a reversal of the direction of movement of the finger or hand, and / or a reversal of the direction of acceleration of the movement of the finger or hand) is detected based on a change in the movement characteristics of the finger or hand that executes the tap gesture.

[0168] In some embodiments, the user's attention being directed to a portion of a three-dimensional environment is determined based on detection of a gaze directed to that portion of the three-dimensional environment (optionally, without requiring additional conditions). In some embodiments, the user's attention being directed to a portion of a three-dimensional environment is determined based on detection of a gaze directed to that portion of the three-dimensional environment using one or more additional conditions, such as requiring the gaze to be directed to that portion of the three-dimensional environment for at least a threshold duration (e.g., a dwell duration) and / or requiring the gaze to be directed to that portion of the three-dimensional environment when the user's viewing point is within a distance threshold of that portion of the three-dimensional environment, such that the device determines that the user's attention is directed to that portion of the three-dimensional environment, where if one of these additional conditions is not met, the device determines that the attention is not directed to that portion of the three-dimensional environment to which the gaze is directed (e.g., until the one or more additional conditions are met).

[0169] In some embodiments, the detection of the ready state configuration of the user or a part of the user is detected by a computer system. The detection of the ready state configuration of the hand is used by the computer system as an indication that the user may be about to use one or more air gesture inputs (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein) performed by the hand to interact with the computer system. For example, based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape where the thumb and one or more fingers are extended and spaced apart to prepare for a pinch or grab gesture, or a pre-tap where one or more fingers are extended and the palm is facing away from the user), based on whether the hand is in a predetermined orientation relative to the user's line of sight (e.g., below the user's head and above the user's waist and extending at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or based on whether the hand has moved in a particular manner (e.g., moved towards an area in front of the user that is above the user's waist and below the user's head or moved away from the user's body or legs) to determine the ready state of the hand. In some embodiments, the ready state is used to determine whether interactive elements of the user interface respond to attention (e.g., gaze) input.

[0170] In a scenario where input is described with reference to an air gesture, it should be understood that a hardware input device attached to one or more of the user's hands or held by one or more of the user's hands can be used to detect a similar gesture, where optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units can be used to track the positioning of the hardware input device in space, and the positioning and / or movement of the hardware input device is used to replace the positioning and / or movement of one or more hands in the corresponding air gesture. In a scenario where input is described with reference to an air pose, it should be understood that a hardware input device attached to one or more of the user's hands or held by one or more of the user's hands can be used to detect a similar pose. User input can be detected using the controls contained in the hardware input device, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, one or more hand or finger overlays that can detect the positioning or change in positioning of parts of the hand and / or fingers relative to each other, relative to the user's body, and / or relative to the user's physical environment, and / or other hardware input device controls, where user input using the controls contained in the hardware input device is used to replace hand and / or finger gestures such as an air tap or an air pinch in the corresponding air gesture. For example, a selection input described as being performed using an air tap or an air pinch input can alternatively be detected using a button press, a tap on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input. As another example, a movement input described as being performed using an air pinch and drag (e.g., an air drag gesture or an air slide gesture) can alternatively be detected based on an interaction with a hardware input control (such as a button press and hold, a touch on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input after the movement of the hardware input device (e.g., along with the hand associated with the hardware input device) through space). Similarly, a two-handed input that includes movement of the hands relative to each other can be performed using one air gesture and one hardware input device that is not performing an air gesture in the hand, two hardware input devices held in different hands, or two air gestures performed using an air gesture and / or an input detected by one or more of the above hardware input devices in different hands in various combinations.

[0171] In some embodiments, the software can be downloaded electronically to the controller 110, for example, via a network, or can alternatively be provided on a tangible non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in the memory associated with the controller 110. Alternatively or in addition, some or all of the described functions of the computer can be implemented in dedicated hardware (such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP)). Although in Figure 4The controller 110 is shown, but by way of example, some or all of the processing functions of the controller, as a unit separate from the image sensor 404, may be performed by a suitable microprocessor and software or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand tracking device) or other devices associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television receiver, a handheld device, or a head-mounted device) or integrated with any other suitable computerized device (such as a game console or a media player). The sensing function of the image sensor 404 may likewise be integrated into a computer or other computerized device that will be controlled by the sensor output.

[0172] Figure 4 Also shown is a schematic diagram of a depth map 410 captured by the image sensor 404 according to some embodiments. As described above, the depth map includes a matrix of pixels having corresponding depth values. The pixels 412 corresponding to the hand 406 have been segmented from the background and the wrist in the figure. The brightness of each pixel within the depth map 410 is inversely proportional to its depth value (i.e., the measured z - distance from the image sensor 404), where the gray shading gets darker as the depth increases. The controller 110 processes these depth values to identify and segment the components of the image having human hand characteristics (i.e., a group of adjacent pixels). These characteristics may include, for example, overall size, shape, and movement from frame to frame in a sequence of depth maps.

[0173] Figure 4 Also schematically shown is a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406. In Figure 4 this, the hand skeleton 414 overlays the hand background 416 that has been segmented from the original depth map. In some embodiments, key feature points on the hand and optionally on the wrist or arm connected to the hand (e.g., points corresponding to knuckles, finger tips, the center of the palm, the end of the hand connected to the wrist, etc.) are identified and located on the hand skeleton 414. In some embodiments, the controller 110 uses the positions and movements of these key feature points over multiple image frames to determine, according to some embodiments, the gesture being performed by the hand or the current state of the hand.

[0174] Figure 5 An example embodiment of an eye tracking device 130 ( Figure 1A ) is shown. In some embodiments, the eye tracking device 130 consists of an eye tracking unit 243 ( Figure 2)Control is used to track the positioning and movement of the user's gaze relative to the scene 105 or relative to the XR content displayed via the display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as, a head-mounted headset, helmet, goggles, or glasses) or a hand-held device placed in a wearable frame, the head-mounted device includes both components for generating XR content for the user to view and components for tracking the user's gaze relative to the XR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a hand-held device or an XR room, the eye tracking device 130 is optionally a device separate from the hand-held device or the XR room. In some embodiments, the eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used in combination with a display generation component that is also head-mounted or a display generation component that is not head-mounted. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally part of a non-head-mounted display generation component.

[0175] In some embodiments, the display generation component 120 uses a display mechanism (e.g., a left near-eye display panel and a right near-eye display panel) to display a frame including a left image and a right image in front of the user's eyes, thereby providing the user with a 3D virtual view. For example, a head-mounted display generation component may include a left optical lens and a right optical lens (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display, and virtual objects are displayed on the transparent or translucent display, through which the user can directly view the physical environment. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects may be projected, for example, onto a physical surface or projected as a hologram such that an individual using the system observes the virtual objects overlapping above the physical environment. In this case, separate display panels and image frames for the left and right eyes may not be required.

[0176] As Figure 5As shown, in some embodiments, the eye tracking device 130 (e.g., a gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) or near-infrared (NIR) camera), and an illumination source (e.g., an IR or NIR light source, such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera can be pointed at the user's eyes to receive the IR or NIR light directly reflected from the eyes by the light source, or alternatively can be pointed at a "hot" mirror located between the user's eyes and the display panel, which reflects the IR or NIR light from the eyes to the eye tracking camera while allowing visible light to pass through. The eye tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60 frames per second - 120 frames per second (fps)), analyzes the images to generate gaze tracking information, and transmits the gaze tracking information to the controller 110. In some embodiments, the user's two eyes are tracked separately by corresponding eye tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by a corresponding eye tracking camera and illumination source.

[0177] In some embodiments, a device-specific calibration process is used to calibrate the eye tracking device 130 to determine the parameters of the eye tracking device for a particular operating environment 100, such as the 3D geometric relationships and parameters of the LEDs, cameras, hot mirrors (if any), eye lenses, and display screens. The device-specific calibration process can be performed at the factory or another facility before the AR / VR equipment is delivered to the end user. The device-specific calibration process can be an automatic calibration process or a manual calibration process. According to some embodiments, the user-specific calibration process can include an estimation of the eye parameters of a particular user, such as pupil position, fovea position, optical axis, visual axis, interpupillary distance, etc. According to some embodiments, once the device-specific parameters and user-specific parameters are determined for the eye tracking device 130, a flash-assisted method can be used to process the images captured by the eye tracking camera to determine the current visual axis and the user's fixation point relative to the display.

[0178] As Figure 5As shown, the eye tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system that includes at least one eye tracking camera 540 (e.g., an infrared (IR) or near-infrared (NIR) camera) positioned on the side of the user's face where eye tracking is to be performed, and an illumination source 530 (e.g., an IR or NIR light source, such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye 592. The eye tracking camera 540 can be directed at a mirror 550 (which reflects IR or NIR light from the eye 592 while allowing visible light to pass through) (e.g., as shown in the top portion of Figure 5 ), or alternatively can be directed at the user's eye 592 to receive reflected IR or NIR light from the eye 592 (e.g., as shown in the bottom portion of Figure 5 ).

[0179] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames for a left display panel and a right display panel) and provides the frames 562 to the display 510. The controller 110 uses the gaze tracking input 542 from the eye tracking camera 540 for various purposes, such as for processing the frames 562 for display. The controller 110 optionally estimates the user's gaze point on the display 510 based on the gaze tracking input 542 obtained from the eye tracking camera 540 using a flash assist method or other suitable method. The gaze point estimated based on the gaze tracking input 542 is optionally used to determine the direction the user is currently looking.

[0180] The following describes several possible use cases of the current gaze direction of a user and is not intended to be limiting. As an example use case, the controller 110 may render virtual content differently based on the determined direction of the user's gaze. For example, the controller 110 may generate virtual content at a higher resolution in the foveal region determined according to the current gaze direction of the user than in the peripheral region. As another example, the controller may at least partially position or move virtual content in the view based on the current gaze direction of the user. As another example, the controller may at least partially display specific virtual content in the view based on the current gaze direction of the user. As another example use case in an AR application, the controller 110 may direct an external camera for capturing the physical environment of the XR experience to focus in the determined direction. Then, the autofocus mechanism of the external camera may focus on an object or a surface in the environment that the user is currently looking at on the display 510. As another example use case, the eye lens 520 may be a focusable lens, and the controller uses the gaze tracking information to adjust the focus of the eye lens 520 such that the virtual object that the user is currently looking at has an appropriate vergence to match the convergence of the user's eyes 592. The controller 110 may utilize the gaze tracking information to direct the eye lens 520 to adjust the focus such that a nearby object that the user is looking at appears at the correct distance.

[0181] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510) mounted in a wearable housing, two eye lenses (e.g., eye lens 520), an eye tracking camera (e.g., eye tracking camera 540), and a light source (e.g., illumination source 530 (e.g., IR or NIR LED)). The light source emits light (e.g., IR or NIR light) towards the user's eyes 592. In some embodiments, the light source may be arranged in a ring or a circle around each of the lenses, as Figure 5 shown. In some embodiments, for example, eight illumination sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer illumination sources 530 may be used, and other arrangements and positions of the illumination sources 530 may be used.

[0182] In some embodiments, the display 510 emits light in the visible light range and does not emit light in the IR or NIR range, and thus does not introduce noise in the gaze tracking system. Note that the position and angle of the eye tracking camera 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, a camera 540 with a wide field of view (FOV) and a camera 540 with a narrow FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.

[0183] As Figure 5 shown, embodiments of the gaze tracking system may be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experience to the user.

[0184] Figure 6 A flash-assisted gaze tracking pipeline according to some embodiments is shown. In some embodiments, the gaze tracking pipeline is implemented by a flash-assisted gaze tracking system (e.g., an eye tracking device 130 as Figure 1A and Figure 5 shown). The flash-assisted gaze tracking system may maintain a tracking state. Initially, the tracking state is off or "no". When in the tracking state, when analyzing the current frame to track the pupil contour and flash in the current frame, the flash-assisted gaze tracking system uses previous information from the previous frame. When not in the tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and flash in the current frame, and if successful, initializes the tracking state to "yes" and continues with the next frame in the tracking state.

[0185] As Figure 6 shown, the gaze tracking camera may capture left and right images of the user's left and right eyes. The captured images are then input into the gaze tracking pipeline for processing to begin at 610. As indicated by the arrow returning to element 600, the gaze tracking system may continue to capture images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images may be input into the pipeline for processing. However, in some embodiments or under some conditions, not all of the captured frames are processed by the pipeline.

[0186] At 610, for the currently captured image, if the tracking state is yes, the method proceeds to element 640. At 610, if the tracking state is no, then as indicated at 620, the image is analyzed to detect the user's pupil and flash in the image. At 630, if the pupil and flash are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eye.

[0187] At 640, if proceeding from element 610, the current frame is analyzed to track the pupil and flash based in part on previous information from a previous frame. At 640, if proceeding from element 630, the tracking state is initialized based on the pupil and flash detected in the current frame. The processing result at element 640 is checked to verify that the result of the tracking or detection can be trusted. For example, the result can be checked to determine whether the pupil and a sufficient number of flashes for performing gaze estimation are successfully tracked or detected in the current frame. At 650, if the result cannot be trusted, then at element 660, the tracking state is set to no, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is trusted, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and flash information is passed to element 680 to estimate the user's point of gaze.

[0188] Figure 6 It is intended to be used as an example of an eye tracking technique that can be used for a particular specific implementation. As would be recognized by one of ordinary skill in the art, according to various embodiments, in the computer system 101 for providing an XR experience to a user, other eye tracking techniques that currently exist or are developed in the future can be used to replace the flash-assisted eye tracking technique described herein or used in combination with the flash-assisted eye tracking technique.

[0189] In some embodiments, a captured portion of the real-world environment 602 is used to provide an XR experience to the user, such as a mixed reality environment in which one or more virtual objects are superimposed over a representation of the real-world environment 602.

[0190] Accordingly, the description herein describes some embodiments of a three-dimensional environment (e.g., an XR environment) that includes a representation of real-world objects and a representation of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table that exists in a physical environment, which is captured and displayed in the three-dimensional environment (e.g., actively displayed via a camera and a display of a computer system or passively displayed via a transparent or semi-transparent display of the computer system). As previously described, the three-dimensional environment is optionally a mixed reality system, where the three-dimensional environment is based on the physical environment captured by one or more sensors of the computer system and displayed via a display generation component. As a mixed reality system, the computer system is optionally capable of selectively displaying portions and / or objects of the physical environment such that the corresponding portions and / or objects of the physical environment appear as if they exist in the three-dimensional environment displayed by the computer system. Similarly, the computer system is optionally capable of displaying virtual objects in the three-dimensional environment at corresponding positions that have corresponding positions in the real world such that the virtual objects appear as if they exist in the real world (e.g., the physical environment). For example, the computer system optionally displays a vase such that the vase appears as if a real vase is placed on top of a table in the physical environment. In some embodiments, the corresponding positions in the three-dimensional environment have corresponding positions in the physical environment. Thus, when the computer system is described as displaying a virtual object at a corresponding position relative to a physical object (e.g., a position at or near the user's hand or a position at or near a physical table), the computer system displays the virtual object at a specific position in the three-dimensional environment such that it appears as if the virtual object is at or near the physical object in the physical environment (e.g., the virtual object is displayed at a position in the three-dimensional environment that corresponds to the position where the virtual object would be displayed in the physical environment if the virtual object were a real object at that specific position).

[0191] In some embodiments, real-world objects that exist in the physical environment and are displayed in the three-dimensional environment (e.g., and / or visible via a display generation component) can interact with virtual objects that only exist in the three-dimensional environment. For example, the three-dimensional environment can include a table and a vase placed on top of the table, where the table is a view (or representation) of a physical table in the physical environment and the vase is a virtual object.

[0192] In a three-dimensional environment (e.g., a real environment, a virtual environment, or an environment that includes a mixture of real and virtual objects), an object is sometimes said to have depth or simulated depth, or an object is said to be visible, displayed, or placed at different depths. In this context, depth refers to a dimension that is different from height or width. In some embodiments, depth is defined relative to a fixed set of coordinates (e.g., where a room or an object has a height, depth, and width defined relative to a fixed set of coordinates). In some embodiments, depth is defined relative to the position or viewpoint of a user, in which case the depth dimension varies based on the position of the user and / or the position and angle of the user's viewpoint. In some embodiments where depth is defined relative to the position of the user relative to a surface of the environment (e.g., the surface of the floor or ground of the environment), an object that is farther from the user along a line extending parallel to the surface is considered to have a greater depth in the environment, and / or the depth of the object is measured along an axis that extends outward from the position of the user and is parallel to the surface of the environment (e.g., depth is defined in a cylindrical or substantially cylindrical coordinate system, where the position of the user is at the center of a cylinder that extends from the user's head toward the user's feet). In some embodiments where depth is defined relative to the user's viewpoint (e.g., the direction relative to a point in space that determines which part of the environment is visible via a head-mounted device or other display), an object that is farther from the user's viewpoint along a line extending parallel to the user's viewpoint is considered to have a greater depth in the environment, and / or the depth of the object is measured along an axis that extends outward from the user's viewpoint and along a line parallel to the direction of the user's viewpoint (e.g., depth is defined in a spherical or substantially spherical coordinate system, where the origin of the viewpoint is at the center of a sphere that extends outward from the user's head). In some embodiments, depth is defined relative to a user interface container (e.g., a window or an application in which an application and / or system content is displayed), where the user interface container has a height and / or width, and depth is a dimension that is orthogonal to the height and / or width of the user interface container. In some embodiments where depth is defined relative to a user interface container, when the container is placed in a three-dimensional environment or is initially displayed (e.g., such that the depth dimension of the container extends outward away from the user or the user's viewpoint), the height and / or width of the container is generally orthogonal or substantially orthogonal to a line that extends from the position of the user (e.g., the user's viewpoint or the position of the user) to the user interface container (e.g., the center of the user interface container or another feature point of the user interface container). In some embodiments where depth is defined relative to a user interface container, the depth of an object relative to the user interface container refers to the positioning of the object along the depth dimension of the user interface container. In some embodiments, multiple different containers may have different depth dimensions (e.g., different depth dimensions that extend in different directions and / or from different starting points away from the user or the user's viewpoint).In some embodiments, when defining depth relative to a user interface container, the direction of the depth dimension remains constant for the user interface container as the position of the user interface container, the user, and / or the user's viewing point changes (e.g., or when multiple different viewers are viewing the same container in a three-dimensional environment, such as during an in-person collaboration session and / or when multiple participants are in a real-time communication session with shared virtual content that includes the container). In some embodiments, for a curved container (e.g., including a container having a curved surface or a curved content area), the depth dimension optionally extends into the surface of the curved container. In some cases, a z-spacing (e.g., the spacing between two objects in the depth dimension), a z-height (e.g., the distance of one object from another object in the depth dimension), a z-position (e.g., the position of one object in the depth dimension), a z-depth (e.g., the position of one object in the depth dimension), or an analog z-dimension (e.g., a depth that serves as a dimension of an object, a dimension of an environment, a direction in space, and / or a direction in an analog space) is used to refer to the concept of depth as described above.

[0193] In some embodiments, a user optionally can interact with virtual objects in a three-dimensional environment using one or more hands as if the virtual objects were real objects in a physical environment. For example, as described above, one or more sensors of a computer system optionally capture one or more hands of the user and display a representation of the user's hands in the three-dimensional environment (e.g., in a manner similar to displaying real-world objects in the three-dimensional environment described above), or in some embodiments, the user's hands are visible via the display generation component, via the ability to see the physical environment through the user interface, due to the transparency / translucency of a portion of the user interface being displayed by the display generation component, or due to the projection of the user interface onto a transparent / translucent surface or onto the user's eyes or into the user's field of view. Thus, in some embodiments, the user's hands are displayed at corresponding positions in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment that can interact with virtual objects in the three-dimensional environment as if those virtual objects were physical objects in a physical environment. In some embodiments, the computer system is able to update the display of the representation of the user's hands in the three-dimensional environment in conjunction with the movement of the user's hands in the physical environment.

[0194] In some of the embodiments described below, the computer system is optionally capable of determining an "effective" distance between a physical object in the physical world and a virtual object in a three-dimensional environment, e.g., for determining whether the physical object is directly interacting with the virtual object (e.g., whether a hand is touching, grasping, holding, etc. the virtual object or is within a threshold distance of the virtual object). For example, a hand directly interacting with a virtual object optionally includes one or more of the following: a finger of the hand pressing a virtual button, a hand of the user grasping a virtual vase, the hands of the user brought together and pinching / holding the user interface of an application, and two fingers performing any other type of interaction described herein. For example, when determining whether a user is interacting with a virtual object and / or how the user is interacting with the virtual object, the computer system optionally determines the distance between the user's hand and the virtual object. In some embodiments, the computer system determines the distance between the user's hand and the virtual object by determining the distance between the position of the hand in the three-dimensional environment and the position of the virtual object of interest in the three-dimensional environment. For example, the one or more hands of the user are located at a particular location in the physical world, and the computer system optionally captures the one or more hands and displays the one or more hands at a particular corresponding location in the three-dimensional environment (e.g., the location where the hand would be displayed in the three-dimensional environment if the hand were a virtual hand rather than a physical hand). Optionally, the location of the hand in the three-dimensional environment is compared with the location of the virtual object of interest in the three-dimensional environment to determine the distance between the one or more hands of the user and the virtual object. In some embodiments, the computer system optionally determines the distance between a physical object and a virtual object by comparing locations in the physical world (e.g., rather than comparing locations in the three-dimensional environment). For example, when determining the distance between one or more hands of the user and a virtual object, the computer system optionally determines the corresponding position of the virtual object in the physical world (e.g., the location where the virtual object would be located in the physical world if the virtual object were a physical object rather than a virtual object), and then determines the distance between the corresponding physical location and the one or more hands of the user. In some embodiments, the same techniques are optionally used to determine the distance between any physical object and any virtual object. Thus, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the computer system optionally performs any of the techniques described above to map the position of the physical object to the three-dimensional environment and / or to map the position of the virtual object to the physical environment.

[0195] In some embodiments, the same or similar techniques are used to determine where and what the user's gaze is directed at, and / or where and what a physical stylus held by the user is directed at. For example, if the user's gaze is directed at a particular location in the physical environment, the computer system optionally determines the corresponding location in the three-dimensional environment (e.g., the virtual location of the gaze), and if a virtual object is located at the corresponding virtual location, the computer system optionally determines that the user's gaze is directed at the virtual object. Similarly, the computer system is optionally able to determine the direction in the physical environment that the stylus is directed based on the orientation of the physical stylus. In some embodiments, based on this determination, the computer system determines the corresponding virtual location in the three-dimensional environment that corresponds to the location in the physical environment that the stylus is directed at, and optionally determines that the stylus is directed at the corresponding virtual location in the three-dimensional environment.

[0196] Similarly, the embodiments described herein may refer to the position of a user (e.g., a user of a computer system) in a three-dimensional environment and / or the position of a computer system in a three-dimensional environment. In some embodiments, the user of the computer system is holding, wearing, or otherwise located at or near the computer system. Thus, in some embodiments, the position of the computer system serves as a proxy for the position of the user. In some embodiments, the position of the computer system and / or the user in the physical environment corresponds to the corresponding position in the three-dimensional environment. For example, the position of the computer system would be the position in the physical environment (and its corresponding position in the three-dimensional environment) where, if the user were standing at that position facing the corresponding portion of the physical environment visible via the display generation component, the user would see in the physical environment those objects that are in the same location, orientation, and / or size (e.g., in an absolute sense and / or relative to each other) as the objects that are displayed in the three-dimensional environment by the display generation component of the computer system or are visible in the three-dimensional environment via the display generation component. Similarly, if the virtual objects displayed in the three-dimensional environment are physical objects in the physical environment (e.g., physical objects placed in the physical environment at the same position as the position of these virtual objects in the three-dimensional environment and having the same size and orientation in the physical environment as they do in the three-dimensional environment), the position of the computer system and / or the user is the position from which the user would see in the physical environment those virtual objects that are in the same location, orientation, and / or size (e.g., in an absolute sense and / or relative to each other and real-world objects) as the virtual objects displayed in the three-dimensional environment by the display generation component of the computer system.

[0197] In the present disclosure, various input methods are described in relation to interaction with a computer system. When one input device or input method is used to provide an example and another input device or input method is used to provide another example, it should be understood that each example may be compatible with and optionally utilize the input device or input method described in relation to the other example. Similarly, various output methods are described in relation to interaction with a computer system. When one output device or output method is used to provide an example and another output device or output method is used to provide another example, it should be understood that each example may be compatible with and optionally utilize the output device or output method described in relation to the other example. Similarly, various methods are described in relation to interaction with a virtual environment or a mixed reality environment via a computer system. When interaction with a virtual environment is used to provide an example and a mixed reality environment is used to provide another example, it should be understood that each example may be compatible with and optionally utilize the methods described in relation to the other example. Accordingly, the present disclosure discloses embodiments that are combinations of features of multiple examples without exhaustively listing all features of the embodiments in the description of each example embodiment.

[0198] User Interface and Associated Processes

[0199] Attention is now turned to embodiments of a user interface (“UI”) and associated processes that may be implemented on a computer system (such as, a portable multifunctional device or a head-mounted device) having a display generation component, one or more input devices, and optionally one or more cameras.

[0200] Figures 7A to 7EE An example of a computer system is shown that changes the visual salience of corresponding virtual objects relative to a three-dimensional environment in response to detecting a threshold amount of overlap between a first virtual object and a second virtual object. In some embodiments, the computer system changes the visual salience of the corresponding virtual objects based on a change in the spatial position of the first virtual object relative to the second virtual object in the three-dimensional environment.

[0201] Figure 7A An example is shown of a computer system (e.g., an electronic device) 101 displaying a three-dimensional environment 702 from the viewpoint of a user (e.g., user 712) of the computer system 101 (e.g., facing the back wall of the physical environment in which the computer system 101 is located) via a display generation component (e.g., the display generation component 120 of FIG. 1). In some embodiments, the computer system 101 includes a display generation component (e.g., a touch screen) and a plurality of image sensors (e.g., Figure 3The image sensor 314). The image sensor optionally includes one or more of the following: a visible light camera; an infrared camera; a depth sensor; or any other sensor that the computer system 101 can use to capture one or more images of the user or a part of the user (e.g., one or more hands of the user) when the user interacts with the computer system 101. In some embodiments, the user interfaces shown and described below may also be implemented on a head-mounted display that includes a display generation component for displaying the user interface or a three-dimensional environment to the user, and sensors for detecting movement of the physical environment and / or the user's hand (e.g., external sensors facing out from the user) and / or sensors for detecting the user's attention (e.g., gaze) (e.g., internal sensors facing in towards the user's face).

[0202] As Figure 7A shown, the computer system 101 displays a first virtual object 704a and a second virtual object 704b in a three-dimensional environment 702. In some embodiments, the first virtual object 704a and the second virtual object 704b have one or more characteristics of the first virtual object, the second virtual object, and / or the corresponding virtual objects described with reference to methods 800 and / or 900. For example, the first virtual object 704a and / or the second virtual object 704b are associated with one or more applications for presenting content in the three-dimensional environment 702 (e.g., the first virtual object 704a is associated with "Application A" and the second virtual object 704b is associated with "Application B"). In some embodiments, the first virtual object 704a and / or the second virtual object 704b present video content (e.g., associated with video media (e.g., from a video streaming application)), website content (e.g., from a web browsing application), phone and / or messaging content (e.g., from a phone, messaging, and / or social media application), or interactive content (e.g., from a video game application).

[0203] In Figure 7A it, one or more objects other than the first virtual object 704a and the second virtual object 704b are visible. Specifically, Figure 7AShows a table 706a, a wall photo 706b, and a door 706c. In some embodiments, the table 706a, the wall photo 706b, and the door 706c are physical objects visible through optical see-through on the display generation component 120 from the user's (e.g., user 712 described below) physical environment. In some embodiments, the table 706a, the wall photo 706b, and the door 706c are virtual representations of physical objects visible through virtual see-through on the display generation component 120 from the user's physical environment. In some embodiments, the three-dimensional environment 702 is an immersive virtual environment (e.g., fully immersive or partially immersive), and one or more objects from the user's physical environment are not visible relative to the user's current viewing point. In some embodiments, the first virtual object 704a and the second virtual object 704b are displayed with a first amount of visual saliency (e.g., including one or more characteristics with a first amount of visual saliency relative to the three-dimensional environment, as described in reference method 800). For example, the first virtual object 704a and the second virtual object 704b are displayed with a certain amount of opacity, brightness, and / or color such that the content associated with the first virtual object 704a and the second virtual object 704b is visible relative to the current viewing point of the user of the computer system 101. In some embodiments, displaying the corresponding virtual object (e.g., the first virtual object 704a or the second virtual object 704b) with a first amount of visual saliency corresponds to the corresponding virtual object being an active virtual object, as described in reference method 800.

[0204] Figures 7A to 7EE Shows a top view 710 of the three-dimensional environment 702. The top view 710 shows the user 712 in the three-dimensional environment 702. In some embodiments, the user 712 is a user of the computer system 101 (e.g., the user 712 is viewing the three-dimensional environment 702 from the current viewing point). In some embodiments, the user 712 in the top view 710 represents the current viewing point of the user 712 relative to the three-dimensional environment 702. In Figure 7A the top view 710 of, the first virtual object 704a and the second virtual object 704b are shown as not overlapping in the three-dimensional environment 702 (e.g., and as Figure 7Ado not overlap with respect to the current viewing point of the user 712). Specifically, the first virtual object 704a and the second virtual object 704b do not spatially conflict in the three-dimensional environment 702 (e.g., at least a portion of the first virtual object 704a and at least a portion of the second virtual object do not appear at the same location in the three-dimensional environment 702). As shown in the top view 710, the first virtual object 704a includes a spatial arrangement that is different from that of the second virtual object 704b with respect to the current viewing point of the user 712. Specifically, with respect to the current viewing point of the user 712 in the three-dimensional environment 702, the first virtual object 704a is located at a first distance in the three-dimensional environment 702, and the second virtual object 704b is located at a second distance greater than the first distance in the three-dimensional environment 702.

[0205] As Figure 7A shown, the user 712 directs an input (e.g., an air pinch input, an air tap input, a pinch input, a tap input, an air pinch and drag input, an air drag input, a drag input, a click and drag input, a gaze input, and / or other input) toward the first virtual object 704a. Specifically, the gaze 708 of the user 712 is directed toward the first virtual object 704a (e.g., represented by the black circle in the three-dimensional environment 702) and the hand 720 of the user 712 is shown. In some embodiments, the user 712 performs an air gesture with the hand 720 (e.g., including one or more of the air gestures described with reference to methods 800 and / or 900), while the attention of the user 712 (e.g., the gaze 708) is concurrently directed toward the first virtual object 704a. In some embodiments, Figure 7A the input shown in corresponds to a request to move (e.g., and / or change the spatial arrangement) the first virtual object 704a in the three-dimensional environment 702 (e.g., and / or change the spatial arrangement of the first virtual object 704a with respect to the current viewing point of the user 712). For example, the input includes hand movements of the hand 720 corresponding to the requested movement of the first virtual object 704a in the three-dimensional environment 702 (e.g., when the attention is directed toward the first virtual object 704a and / or an air gesture is performed). In some embodiments, Figure 7A the input shown in has one or more characteristics of the first input described with reference to methods 800 and / or 900. In some embodiments, an input having Figure 7A one or more characteristics of the input shown in can be directed toward the second virtual object 704b to move the second virtual object 704b in the three-dimensional environment 702 (e.g., to change the spatial arrangement of the second virtual object 704b with respect to the current viewing point of the user 712).

[0206] Figure 7A1 is shown in connection with Figure 7AConcepts that are similar and / or identical to those shown (with many identical reference numerals). It should be understood that, unless otherwise indicated below, Figure 7A1 Elements shown in Figures 7A to 7EE with the same reference numerals as the elements shown in Figure 7A1 include computer system 101, which includes a display generation component 120 (or the same). In some embodiments, computer system 101 and display generation component 120 respectively have Figures 7A to 7EE the computer system 101 shown in Figure 3 and one or more characteristics of the display generation component 120 shown in FIG. 1, and in some embodiments, Figures 7A to 7EE the computer system 101 and display generation component 120 shown in Figure 7A1 have one or more characteristics of the computer system 101 and display generation component 120 shown in

[0207] In Figure 7A1 the display generation component 120 includes one or more internal image sensors 314a oriented towards the user's face (e.g., the eye tracking camera 540 described with reference to Figure 5 ). In some embodiments, the internal image sensor 314a is used for eye tracking (e.g., detecting the user's gaze). The internal image sensor 314a is optionally arranged on the left and right portions of the display generation component 120 to enable eye tracking of the user's left and right eyes. The display generation component 120 also includes external image sensors 314b and 314c facing outwards from the user to detect and / or capture the physical environment and / or the movement of the user's hand. In some embodiments, the image sensors 314a, 314b, and 314c have one or more characteristics of the image sensor 314 described with reference to Figures 7A to 7EE ).

[0208] In Figure 7A1 the display generation component 120 is shown displaying content that optionally corresponds to the content described with reference to Figures 7A to 7EE as being displayed and / or visible via the display generation component 120. In some embodiments, the content is displayed by a single display included in the display generation component 120 (e.g., the display 510 of Figure 5 ). In some embodiments, the display generation component 120 includes two or more displays with display outputs that are combined (e.g., by the user's brain) to create the view of the content shown in Figure 7A1 (e.g., a left display panel and a right display panel for the user's left and right eyes respectively, as described with reference to Figure 5 ).

[0209] The display generation component 120 has a field of view corresponding to Figure 7A1 the content shown in (e.g., the field of view captured by the external image sensors 314b and 314c and / or visible to the user via the display generation component 120). Since the display generation component 120 is optionally a head-mounted device, the field of view of the display generation component 120 is optionally the same as or similar to the user's field of view.

[0210] In Figure 7A1 , the user is depicted as performing an air pinch gesture (e.g., with hand 720) to provide input to the computer system 101, thereby providing user input for the content displayed by the computer system 101. This description is intended to be exemplary and not restrictive; the user optionally uses different air gestures and / or other forms of input to provide user input, as referenced in Figures 7A to 7EE described.

[0211] In some embodiments, the computer system 101 responds to user input, as referenced in Figures 7A to 7EE described.

[0212] In Figure 7A1 's example, since the user's hand is within the field of view of the display generation component 120, it is visible within the three-dimensional environment. That is, the user can optionally see any part of their own body within the field of view of the display generation component 120 in the three-dimensional environment. It should be understood that one or more or all aspects of the present disclosure, as shown in Figures 7A to 7EE or referenced in Figures 7A to 7EE described and / or referenced in the corresponding method, are optionally implemented on the computer system 101 and the display generation unit 120 in a manner similar or analogous to that shown in Figure 7A1 .

[0213] Figure 7B shows the movement of the first virtual object 704a in the three-dimensional environment 702 (e.g., relative to the current viewpoint of the user 712) in response to input provided by the user 712 in Figure 7A . As in Figure 7BAs shown, the movement of the first virtual object 704a in the three-dimensional environment 702 causes the first virtual object 704a to at least partially overlap with the second virtual object 704b (e.g., at least a portion of the first virtual object 704a spatially conflicts with (e.g., visually occludes) the second virtual object 704b relative to the current viewing point of the user 712) (e.g., from the user's viewing point, the first virtual object overlaps with the second virtual object, and optionally, the first virtual object is within a threshold distance of the second virtual object in the depth dimension). Specifically, the first virtual object 704a is displayed in the three-dimensional environment 702 at a distance closer to the user 712 (e.g., relative to the current viewing point of the user 712) than the second virtual object 704b, such that the portion of the first virtual object 704a that overlaps with the second virtual object 704b visually occludes a portion of the second virtual object 704b relative to the current viewing point of the user 712.

[0214] In some embodiments, based on a threshold amount of overlap detected by computer system 101 between a portion of a first virtual object 704a and a second virtual object 704b, a corresponding virtual object (e.g., the first virtual object 704a or the second virtual object 704b) displayed in a three-dimensional environment 702 is displayed with a different visual prominence (e.g., computer system 101 reduces the visual prominence of at least a portion of the corresponding virtual object). Thus, the top view 710 shows a schematic illustration of an area (e.g., an area) of an overlap threshold 714a and an angle (e.g., an angular distance) of an overlap threshold 714b corresponding to a threshold amount of overlap (e.g., or optionally one or more threshold amounts of overlap) to be detected by computer system 101 to change the visual prominence of a corresponding virtual object displayed in the three-dimensional environment 702. In some embodiments, based on computer system 101 detecting that the overlap between the first virtual object 704a and the second virtual object 704b exceeds an overlap area threshold 714a and / or an overlap angle threshold 716b, at least a portion of the first virtual object 704a or the second virtual object 704b is displayed with a different (e.g., reduced) visual prominence. For example, based on the user 712's attention being directed to the first virtual object 704a (e.g., by a gaze 708 while concurrently performing an air gesture (e.g., an air pinch) with a hand 720), the second virtual object 704b is displayed with a different visual prominence (e.g., the first virtual object 704a is the active virtual object). For example, based on the user 712's attention being directed to the second virtual object 704b (e.g., by a gaze 708 while concurrently performing an air gesture (e.g., an air pinch) with a hand 720), the first virtual object 704a is displayed with a different visual prominence (e.g., the second virtual object 704b is the active virtual object). In some embodiments, the threshold amount of overlap (e.g., the overlap area threshold 714a and / or the overlap angle threshold 714b) has one or more characteristics of a threshold amount of overlap between at least a portion of the first virtual object and the second virtual object, as described in reference method 800.

[0215] As Figure 7B shown, the overlap between the first virtual object 704a and the second virtual object 704b does not exceed the overlap area threshold 714a or the overlap angle threshold 714b. Based on the overlap between the first virtual object 704a and the second virtual object 704b not exceeding the threshold amount of overlap, computer system 101 maintains displaying the first virtual object 704a and the second virtual object 704b with a first visual prominence relative to the three-dimensional environment 702.

[0216] In Figure 7BIn this case, the user 712 directs an input (e.g., an air pinch input, an air tap input, a pinch input, a tap input, an air pinch and drag input, an air drag input, a drag input, a click and drag input, a gaze input, and / or other input) corresponding to a request to move a first virtual object 704a in a three-dimensional environment 702 (e.g., corresponding to a gaze 708 directed at the first virtual object 704a and an air gesture and / or hand movement performed by the hand 720) at the first virtual object 704a. In some embodiments, Figure 7B shows that the computer system 101 continues to receive the input initiated by the user 712 in Figure 7A . For example, Figure 7B the input shown in Figure 7A is a continuation of the input shown in Figure 7A (e.g., the user 712 continues to move the first virtual object 704a in the three-dimensional environment 702 by continuing to direct the gaze 708 at the first virtual object 704a while continuing to perform the air gesture and / or hand movement initiated in

[0217] Figure 7C shows the movement of the first virtual object 704a in the three-dimensional environment 702 (e.g., relative to the current viewpoint of the user 712) based on the input provided by the user 712 in Figures 7A to 7B . Due to the movement of the first virtual object 704a in the three-dimensional environment 702 (e.g., relative to the current viewpoint of the user 712), the first virtual object 704a overlaps with the second virtual object 704b by more than a threshold amount of overlap (e.g., more than an overlap region threshold 714 and / or an overlap angle threshold 714b, as shown in the top view 710) relative to the current viewpoint of the user 712. In response to the movement of the first virtual object 704a causing the overlap with the second virtual object 704b to exceed the threshold amount, the second virtual object 704b (e.g., or optionally a portion of the second virtual object 704b) is displayed with a second amount of visual prominence (e.g., including one or more characteristics of the second visual prominence, as described in reference method 800). In some embodiments, displaying the second virtual object 704b with the second amount of visual prominence includes displaying it with a first amount of visual prominence (e.g., in Figures 7A to 7BThe visual prominence of the amount of the second virtual object 704b is shown) As compared to showing the second virtual object 704b, the second virtual object 704b is shown with a reduced amount of brightness, color, saturation, and / or opacity (e.g., or optionally a portion of the second virtual object 704b). In some embodiments, showing the second virtual object 704b with a second amount of visual prominence includes stopping showing a portion of the second virtual object 704b that is overlapped by the first virtual object 704a with respect to the current viewpoint of the user 712 in the three-dimensional environment (e.g., this portion of the second virtual object 704b spatially conflicts with the first virtual object 704a with respect to the current viewpoint of the user 712 (e.g., is visually occluded by the first virtual object 704a)) (e.g., from the user's viewpoint, the second virtual object overlaps the first virtual object, and optionally, the second virtual object is within a threshold distance of the first virtual object in the depth dimension). In some embodiments, the second virtual object 704b is shown with a second amount of visual prominence because the attention of the user 712 (e.g., via the gaze 708 and an air gesture and / or hand movement performed by the hand 720) is directed to the first virtual object 704a when performing Figures 7A to 7B the input shown in (e.g., the first virtual object 704a is the active virtual object).

[0218] As Figure 7C shown, the user 712 stops directing the input to the first virtual object 704a (e.g., stops moving the hand for more than a threshold amount of time, releases the user's finger for an air pinch input, closes the user's eyes, or other input indicating the end of the input), and directs the input (e.g., an air pinch input, an air tap input, a pinch input, a tap input, an air pinch and drag input, an air drag input, a drag input, a click and drag input, a gaze input, and / or other input) to the second virtual object 704b. Specifically, the gaze 708 is directed to the second virtual object 704b. In some embodiments, when the user 712 directs the gaze 708 to the second virtual object 704b, the user 712 performs an air gesture (e.g., an air pinch) with the hand 720. In some embodiments, Figure 7C the input shown in corresponds to a request to interact with the second virtual object 704b (e.g., and a request to show the second virtual object 704b with a first amount of visual prominence and the first virtual object 704a with a second amount of visual prominence). For example, Figure 7C the input shown in corresponds to a request to make the second virtual object 704b the active virtual object.

[0219] Fig.7D shows in response to being caused by Figure 7CThe second virtual object 704b displayed with a first amount of visual prominence and the first virtual object 704a displayed with a second amount of visual prominence based on the input provided by the user 712 in it. In some embodiments, compared to as shown in FIG. 7A to FIG. 7C , the first virtual object 704a is displayed with a reduced amount of brightness, color, saturation, and / or opacity in Fig.7D . In some embodiments, the computer system 101 stops displaying the portion of the first virtual object 704a that is overlapped by the second virtual object 704b in the three-dimensional environment 702 (e.g., this portion of the first virtual object 704a has one or more characteristics of the corresponding portion of the corresponding virtual object as described in reference method 800 and / or one or more characteristics of the first portion of at least a part of the second virtual object as described in reference method 900). For example, this portion of the first virtual object 704a has a size relative to the three-dimensional environment that corresponds to the size of this portion of the second virtual object 704b that overlaps the first virtual object 704a relative to the current viewpoint of the user 712.

[0220] As 7A to 7D shown (e.g., in the top view 710), the second virtual object 704b is displayed at a greater distance from the current viewpoint of the user 712 compared to the first virtual object 704a. In some embodiments, based on the second virtual object 704b being displayed at a greater distance from the current viewpoint of the user 712 compared to the first virtual object 704a, the portion 718a of the first virtual object 704a is displayed with a greater amount of transparency compared to the portion of the first virtual object 704a displayed with the first amount of visual prominence (e.g., the portion 718a of the first virtual object 704a has one or more characteristics of the second portion of the corresponding portion of the corresponding virtual object as described in reference method 800 and / or one or more characteristics of the second portion of at least a part of the second virtual object as described in reference method 900). As Fig.7D shown, the portion 718a of the first virtual object 704a surrounds the portion of the second virtual object 704b that overlaps the first virtual object 704a relative to the current viewpoint of the user 712 (e.g., the portion 718a of the first virtual object 704a surrounds the portion of the first virtual object 704a that the computer system 101 stops displaying in the three-dimensional environment 702). In some embodiments, in Fig.7DIn [context], although there is a spatial conflict (e.g., overlap) between the first virtual object 704a and the second virtual object 704b and the first virtual object 704a is displayed at a closer distance relative to the current viewing point of the user 712 (e.g., because the computer system 101 stops displaying the portion of the first virtual object 704a that visually occludes the second virtual object 704b and displays the portion 718a of the first virtual object 704a surrounding the second virtual object 704b in a transparent manner), the second virtual object 704b is visible (e.g., not visually occluded by the first virtual object 704a).

[0221] In Fig.7D [context], the user 712 directs an input (e.g., an air pinch input, an air tap input, a pinch input, a tap input, an air pinch and drag input, an air drag input, a drag input, a click and drag input, a gaze input, and / or other input) to a blank space in the three-dimensional environment 702 (e.g., a region of the three-dimensional environment that does not include one or more virtual objects (e.g., the first virtual object 704a or the second virtual object 704b)). In some embodiments, the blank space in the three-dimensional environment 702 has one or more characteristics of the blank space in the three-dimensional environment as described in the reference method 800. As Fig.7D shown, the input directed to the blank space in the three-dimensional environment 702 includes a gaze 708 directed to the blank space while the user 712 performs an air gesture (e.g., an air pinch) with the hand 720. In some embodiments, Fig.7D the input shown in [context] corresponds to a request to change the respective virtual object (e.g., the first virtual object 704a or the second virtual object 704b) displayed with a first amount of visual prominence (e.g., which respective virtual object is displayed as the active virtual object). For example, Fig.7D the input shown in [context] corresponds to a request to display the respective virtual object (e.g., the first virtual object 704a) that is displayed closest to the current viewing point of the user 712 with a first amount of visual prominence (e.g., and one or more virtual objects different from the respective virtual object (e.g., the second virtual object 704b) displayed with a second amount of visual prominence in the three-dimensional environment 702).

[0222] Fig. 7E shows the first virtual object 704a displayed with a first amount of visual prominence and the second virtual object 704b displayed with a second amount of visual prominence in response to an input provided by the user 712 in Fig.7D [context]. In some embodiments, displaying the first virtual object 704a with a first amount of visual prominence and the second virtual object 704b with a second amount of visual prominence includes referring to Figure 7CThe illustrated and described display one or more characteristics of the first virtual object 704a with a first amount of visual prominence and the second virtual object 704b with a second amount of visual prominence.

[0223] In some embodiments, the computer system 101 changes the visual prominence of the second virtual object 704b based on the change in the spatial position of the first virtual object 704a relative to the second virtual object 704b (e.g., including changing one or more characteristics of the visual prominence of at least a portion of the second virtual object relative to the three-dimensional environment based on the change in the spatial position of the first virtual object relative to the second virtual object during movement of the first virtual object in the three-dimensional environment, as described with reference to method 900). Fig. 7E , the top view 710 includes a schematic diagram of spatial position thresholds 716a and 716b. In some embodiments, the spatial position thresholds 716a and 716b correspond to distance thresholds relative to the second virtual object 704b. For example, the distance threshold corresponds to the distance from the second virtual object 704b in the three-dimensional environment 702 in a first dimension (e.g., in the depth direction relative to the current viewpoint of the user 712). In some embodiments, the spatial position thresholds 716a and 716b correspond to distance thresholds relative to the current viewpoint of the user 712. For example, the distance threshold is associated with the distance from the current viewpoint of the user 712 in the first dimension in the three-dimensional environment 702, which distance differs from the distance of the second virtual object 704b from the current viewpoint of the user 712 by more than a threshold amount.

[0224] like Fig. 7E As shown, the input (e.g., air pinch input, air tap input, pinch input, tap input, air pinch and drag input, air drag input, drag input, click and drag input, gaze input, and / or other input) is directed to the first virtual object 704a. In some embodiments, the input corresponds to a request to move the first virtual object 704a in a first dimension (e.g., in a depth direction relative to a current viewpoint of the user 712) in the three-dimensional environment. Fig. 7E The input shown in 708 includes attention (e.g., gaze 708) of the user 712 directed toward the first virtual object 704a. In some embodiments, when the gaze is directed toward the first virtual object 704a, the user 712 performs an air gesture (e.g., air pinch) and / or hand movement relative to the three-dimensional environment 702 (e.g., hand movement in the depth direction in the three-dimensional environment 702 relative to the current viewpoint of the user 712).

[0225] Figure 7F shows the response to Fig. 7E The first virtual object 704a moves in the three-dimensional environment 702 based on the input provided by the user 712 in the three-dimensional environment 702. Fig. 7EThe input provided in , the first virtual object 704a moves in the three-dimensional environment 702 (e.g., in the first dimension) to a greater distance relative to the current viewpoint of the user 712. As shown in the top view 710, the movement of the first virtual object 704a in the first dimension in the three-dimensional environment 702 positions the first virtual object 704a at a spatial position within the spatial position thresholds 716a and 716b relative to the second virtual object 704b. In some embodiments, since the first virtual object 704a is at a spatial position within the spatial position thresholds 716a and 716b relative to the second virtual object 704b, the computer system 101 changes the visual salience of a portion 718b of the second virtual object 704b. In some embodiments, changing the visual salience of the portion 718b of the second virtual object 704b includes changing one or more characteristics of the visual salience of the portion 718a of the first virtual object 704a as described above. In some embodiments, changing the visual salience of the portion 718b of the second virtual object 704b includes changing one or more characteristics of the visual salience of at least a portion of the second virtual object relative to the three-dimensional environment based on a change in the spatial position of the first virtual object relative to the second virtual object during the movement of the first virtual object in the three-dimensional environment, as described in reference method 900. For example, compared to displaying the portion 718b with a first amount of visual salience, the portion 718b of the second virtual object 704b is displayed with a greater amount of transparency. In some embodiments, based on the spatial position of the first virtual object 704a relative to the second virtual object 704b during the movement of the first virtual object 704a in the three-dimensional environment 702 (e.g., and within the spatial position thresholds 716a and 716b), the computer system 101 reduces the visual salience of the second virtual object 704b by different amounts. In some embodiments, reducing the visual salience by different amounts includes changing the size of the portion 718b displayed with a greater amount of transparency based on the spatial position of the first virtual object 704a relative to the second virtual object 704b. For example, in Figure 7F In , the portion 718b of the second virtual object 704b has a first size relative to the three-dimensional environment 702. In some embodiments, the size of the portion 718b increases as the first virtual object 704a moves closer to the second virtual object 704b in the three-dimensional environment 702 (e.g., relative to the first dimension) (e.g., as the difference between the distance of the first virtual object 704a from the current viewpoint of the user 712 and the distance of the second virtual object 704b from the current viewpoint of the user 712 becomes smaller, the size of the portion 718b increases relative to the three-dimensional environment 702). In Figure 7FIn this case, since part 718b of the second virtual object 704b is displayed with a relatively large amount of transparency, the parts of the second virtual object 704b that are different from part 718b (e.g., the remaining parts of the second virtual object 704b outside of part 718b) continue to be displayed with a second amount of visual prominence (e.g., with the amount of visual prominence as shown in Fig. 7E ). In Figure 7F , the parts of the second virtual object 704b that spatially conflict with the first virtual object 704a (e.g., are visually occluded by the first virtual object) relative to the current viewing point of the user 712 (e.g., from the user's viewpoint, this part of the second virtual object is overlapped by the first virtual object, and optionally, the second virtual object is within a threshold distance of the first virtual object in the depth dimension) stop being displayed in the three-dimensional environment 702 (e.g., as shown and described with reference to Figure 7C ).

[0226] As shown in Figure 7F , an input corresponding to a request to move the first virtual object 704a in the three-dimensional environment 702 (e.g., an air pinch input, an air tap input, a pinch input, a tap input, an air pinch and drag input, an air drag input, a drag input, a click and drag input, a gaze input, and / or other inputs) is directed at the first virtual object 704a (e.g., the input shown in Figure 7F has one or more characteristics of the input shown and described with reference to Fig. 7E ). In some embodiments, Figure 7F shows that the computer system 101 continues to receive an input initiated by the user 712 in Fig. 7E . For example, the input shown in Figure 7F is a continuation of the input shown in Fig. 7E (e.g., the user 712 continues to move the first virtual object 704a in the three-dimensional environment 702 (e.g., in the first dimension) by continuing to direct the gaze 708 at the first virtual object 704a while performing the air gesture and / or hand movement initiated in Fig. 7E ).

[0227] Figure 7G shows a response to an input by Figure 7Fthe input provided by user 712 in, the movement of the first virtual object 704a in the three-dimensional environment 702. As shown in the top view 710, the first virtual object 704a spatially conflicts with the second virtual object 704b in the three-dimensional environment 702 (e.g., a portion of the first virtual object 704a is at the same position in the three-dimensional environment 702 as a portion of the second virtual object 704b) (e.g., from the user's viewpoint, the first virtual object overlaps the second virtual object, and optionally, the first virtual object is within a threshold distance of the second virtual object in the depth dimension). In some embodiments, in Figure 7G , the first virtual object 704a and the second virtual object 704b are at the same distance from the current viewpoint of the user 712 in the three-dimensional environment 702.

[0228] Due to the change in the spatial position of the first virtual object 704a relative to the second virtual object 704b (e.g., the first virtual object 704a has moved closer to the second virtual object 704b in the three-dimensional environment 702 compared to that shown and described previously in Figure 7F ), the visual prominence of the second virtual object 704b decreases by a greater amount in Figure 7G . For example, relative to the three-dimensional environment 702, part 718b has a second size that is greater than the first size of part 718b (e.g., as shown and described with reference to Figure 7F ). For example, part 718b is displayed with a greater amount of transparency compared to part 718b shown in Figure 7F . In some embodiments, the size of part 718b is relative to the maximum size of the three-dimensional environment 702 (e.g., because the first virtual object 704a and the second virtual object 704b are at the same distance from the current viewpoint of the user 712 in the three-dimensional environment 702). In some embodiments, part 718b is displayed with the maximum amount of transparency (e.g., because the first virtual object 704a and the second virtual object 704b are at the same distance from the current viewpoint of the user 712 in the three-dimensional environment 702). In Figure 7G , since part 718b of the second virtual object 704b is displayed with a greater amount of transparency, the portion of the second virtual object 704b that is different from part 718b (e.g., the remaining portion of the second virtual object 704b outside of part 718b) continues to be displayed with a second amount of visual prominence. In Figure 7G , the portion of the second virtual object 704b that spatially conflicts with the first virtual object 704a (e.g., is visually occluded by the first virtual object) relative to the current viewpoint of the user 712 stops being displayed in the three-dimensional environment 702 (e.g., as shown with reference to Figure 7Cas shown and described (e.g., from the user's perspective, this part of the second virtual object is overlapped by the first virtual object, and optionally, the second virtual object is within a threshold distance of the first virtual object in the depth dimension).

[0229] As Figure 7G shown, the input corresponding to the request to move the first virtual object 704a in the three-dimensional environment 702 points to the first virtual object 704a (e.g., Figure 7G the input shown in has a reference Fig. 7E one or more characteristics of the input shown and described. In some embodiments, Figure 7G shows that the computer system 101 continues to receive the input initiated by the user 712 in Fig. 7E e.g., Figure 7G the input shown in is FIG. 7E to FIG. 7F a continuation of the input shown in (e.g., the user 712 continues to move the first virtual object 704a in the three-dimensional environment 702 by continuing to point the gaze 708 at the first virtual object 704a while concurrently performing the air gesture and / or hand movement initiated in Fig. 7E ). In some embodiments, Figure 7G the input shown in corresponds to a request to move the first virtual object 704a in a second (e.g., and / or third) dimension different from the first dimension (e.g., the input corresponds to a request to move the first virtual object 704a laterally and / or vertically (e.g., rather than in the depth direction) relative to the current perspective of the user 712).

[0230] Figure 7H shows in response to that by Figure 7GThe input provided by user 712 in [description], the movement of the first virtual object 704a in the three-dimensional environment 702. Specifically, the first virtual object 704a moves vertically and laterally in the three-dimensional environment 702 relative to the current viewing point of the user 712. Due to the movement of the first virtual object 704a in the three-dimensional environment 702 (e.g., relative to the current viewing point of the user 712), the spatial conflict (e.g., the amount of overlap) between the first virtual object 704a and the second virtual object 704b changes (e.g., the first virtual object 704a overlaps with the second virtual object 704b by a greater amount (e.g., the first virtual object 704a and the second virtual object 704b overlap in a larger area relative to the current viewing point of the user 712)). In response to the movement of the first virtual object 704a, the computer system 101 changes the display of the portion 718b of the second virtual object 704b that is displayed with a greater amount of transparency, and changes the size of the portion of the second virtual object 704b that ceases to be displayed in the three-dimensional environment 702 (e.g., changing the display of the portion 718b of the second virtual object and changing the size of the portion of the second virtual object that ceases to be displayed in the three-dimensional environment 702 includes changing one or more characteristics of redisplaying at least a first portion of at least a part of the second virtual object and ceasing to display a third portion different from the first portion of the second virtual object based on the change in the spatial conflict of the second virtual object relative to the first virtual object during the movement of the first virtual object in the three-dimensional environment, as described in the reference method 900). As Figure 7H shown, the portion 718b of the second virtual object 704b corresponds to different parts of the second virtual object 704b (e.g., because compared to as Figure 7G shown, different parts of the second virtual object 704b spatially conflict with the first virtual object 704a relative to the current viewing point of the user 712 (e.g., from the user's viewing point, this part of the second virtual object overlaps with the first virtual object, and optionally, the second virtual object is within a threshold distance of the first virtual object in the depth dimension)). In Figure 7H it, different parts of the second virtual object 704b (e.g., parts that are larger in size compared to as Figure 7G shown) cease to be displayed in the three-dimensional environment 702 (e.g., because compared to as Figure 7G shown, a larger part of the first virtual object 704a spatially conflicts with the second virtual object 704b relative to the current viewing point of the user 712 (e.g., from the user's viewing point, this part of the first virtual object overlaps with the second virtual object, and optionally, the first virtual object is within a threshold distance of the second virtual object in the depth dimension)). In Figure 7HIn [description], when a portion 718b of the second virtual object 704b is displayed with a greater amount of transparency, a portion of the second virtual object 704b that is different from the portion 718b (e.g., the remaining portion of the second virtual object 704b outside the portion 718b (e.g., due to a change in the spatial conflict between the first virtual object 704a and the second virtual object 704b, optionally having different dimensions relative to the three-dimensional environment 702 compared to as Figure 7G shown) continues to be displayed with a second amount of visual prominence.

[0231] As Figure 7H shown, an input corresponding to a request to move the first virtual object 704a in the three-dimensional environment 702 (e.g., an air pinch input, an air tap input, a pinch input, a tap input, an air pinch and drag input, an air drag input, a drag input, a click and drag input, a gaze input, and / or other inputs) points to the first virtual object 704a (e.g., Figure 7H the input shown in [description] has one or more characteristics of the input shown and described with reference to Fig. 7E shown). In some embodiments, Figure 7F shows that the computer system 101 continues to receive an input initiated by the user 712 in Fig. 7E [description]. For example, Figure 7F the input shown in [description] is a continuation of the input shown in FIG. 7E to FIG. 7G [description] (e.g., the user 712 continues to move the first virtual object 704a in the three-dimensional environment 702 by continuing to point the gaze 708 at the first virtual object 704a while concurrently performing the air gesture and / or hand movement initiated in Fig. 7E [description]). In some embodiments, Figure 7H the input shown in [description] corresponds to a request to move the first virtual object 704a in a first dimension (e.g., in the depth direction) relative to the current viewpoint of the user 712.

[0232] Fig.7I shows in response to that by Figure 7HThe input provided by user 712 in [context], the movement of the first virtual object 704a in the three-dimensional environment 702. As shown in the top view 710, compared with the distance of the second virtual object 704b relative to the current viewpoint of user 712, the first virtual object 704a moves to a position with a greater distance relative to the current viewpoint of user 712 in the three-dimensional environment 702. Additionally, as shown in the top view 710, the first virtual object 704a is displayed at a position within the spatial position thresholds 716a and 716b in the three-dimensional environment 702. Due to the movement of the first virtual object 704a in the three-dimensional environment 702 (e.g., the change in the spatial arrangement of the first virtual object 704a relative to the current viewpoint of user 712), the spatial position of the first virtual object 704a relative to the second virtual object 704b changes (e.g., compared with as Figure 7H shown (e.g., the first virtual object 704a is no longer at the same distance from the current viewpoint of user 712 as the second virtual object 704b in the three-dimensional environment 702)). Based on the change in the spatial position of the first virtual object 704a relative to the second virtual object 704b, the computer system 101 changes the visual prominence of the display of the second virtual object 704b. In some embodiments, due to the difference between the distance of the first virtual object 704a relative to the current viewpoint of user 712 and the distance of the second virtual object 704b relative to the current viewpoint of user 712 being greater compared with as Figure 7H shown (e.g., and since the first virtual object 704a is displayed within the spatial position thresholds 716a and 716b), the computer system 101 displays the second virtual object 704b with a greater amount of visual prominence compared with as Figure 7H shown. For example, as Fig.7I shown, compared with as Figure 7H shown, the portion 718b is displayed with a reduced size relative to the three-dimensional environment 702. In some embodiments, compared with as Figure 7H shown, the portion 718b is displayed with a reduced amount of transparency. In Fig.7I , when the portion 718b of the second virtual object 704b is displayed with a greater amount of transparency, the portion of the second virtual object 704b that is different from the portion 718b (e.g., the remaining portion of the second virtual object 704b outside the portion 718b (e.g., optionally having a different size relative to the three-dimensional environment 702 compared with as Figure 7H shown due to the change in the size of the portion 718b)) continues to be displayed with a second amount of visual prominence. In Fig.7I , the portion of the second virtual object 704b that spatially conflicts with the first virtual object 704a relative to the current viewpoint of user 712 (e.g., is visually occluded by the first virtual object) stops being displayed in the three-dimensional environment 702 (e.g., as referenced Figure 7Has shown and described (e.g., from the user's perspective, this portion of the second virtual object is overlapped by the first virtual object, and optionally, the second virtual object is within a threshold distance of the first virtual object in the depth dimension).

[0233] As Fig.7I shown, an input corresponding to a request to move a first virtual object 704a in a three-dimensional environment 702 (e.g., an air pinch input, an air tap input, a pinch input, a tap input, an air pinch and drag input, an air drag input, a drag input, a click and drag input, a gaze input, and / or other inputs) is directed at the first virtual object 704a (e.g., Fig.7I the input shown in has reference Fig. 7E to one or more characteristics of the input shown and described). In some embodiments, Fig.7I shows that the computer system 101 continues to receive an input initiated by the user 712 in Fig. 7E . For example, Fig.7I the input shown in is FIG. 7E to FIG. 7H a continuation of the input shown in (e.g., the user 712 continues to move the first virtual object 704a in the three-dimensional environment 702 by continuing to direct the gaze 708 at the first virtual object 704a while concurrently continuing to perform the air gesture and / or hand movement initiated in Fig. 7E ). In some embodiments, Fig.7I the input shown in corresponds to a request to move the first virtual object 704a further in a first dimension (e.g., in the depth direction) relative to the current viewpoint of the user 712.

[0234] Figure 7J shows the movement of the first virtual object 704a in the three-dimensional environment 702 in response to an input provided by the user 712 in Fig.7I . As shown in the top view 710, compared to the distance of the first virtual object 704a relative to the current viewpoint of the user 712 shown in Fig.7I , the first virtual object 704a moves to a position at a greater distance relative to the current viewpoint of the user 712 in the three-dimensional environment 702. Due to the movement of the first virtual object 704a, the first virtual object 704a is not shown at a spatial position within the spatial position thresholds 716a and 716b relative to the second virtual object 704b. Since the first virtual object 704a moves to a position in the three-dimensional environment 702 that is not within the spatial position thresholds 716a and 716b, the computer system 101 changes the amount of visual prominence of the second virtual object 704b. Specifically, as Figure 7J shown, the second virtual object 704b visually occludes the first virtual object 704a relative to the current viewpoint of the user 712 (e.g., as compared to Fig.7IAs compared to that shown, a greater portion of the first virtual object 704a is not visible from the current viewpoint of the user 712). In some embodiments, the computer system 101 displays the first virtual object 704a at a greater distance relative to the current viewpoint of the user 712 (e.g., as compared to the second virtual object 704b) and when moving in the three-dimensional environment 702, does not fall within the spatial position thresholds 716b and 716b, and displays a portion of the second virtual object 704b (e.g., different from the portion 718b) with transparency. For example, as Figure 7J shown, the portion 718c of the second virtual object 704b is displayed with a greater amount of transparency (e.g., in some embodiments, the portion of the first virtual object 704a corresponding to the size of the portion 718c is visible relative to the current viewpoint of the user 712 (e.g., because the portion 718c is displayed as transparent)). Optionally, the computer system 101 stops displaying the portion 718c of the second virtual object 704b in the three-dimensional environment 702 (e.g., the portion 718c corresponds to a smaller size of the portion of the second virtual object 704b that the computer system 101 stops displaying when the first virtual object 704a moves within the spatial position thresholds 716a and 716b (e.g., as FIG. 7F to FIG. 7I shown)). In some embodiments, when the first virtual object 704a moves in the three-dimensional environment 702 beyond the spatial position thresholds 716b and 716a and is located at a position corresponding to a greater distance from the current viewpoint of the user 712 as compared to the second virtual object 704b, the second virtual object 704b visually occludes the entire portion of the first virtual object 704a that overlaps with the second virtual object 704b (e.g., the second virtual object 704b is not displayed together with the transparent portion 718c, and the portion of the first virtual object 704a that overlaps with the second virtual object 704b is not visible relative to the current viewpoint of the user 712).

[0235] As Figure 7J shown, the input corresponding to the request to move the first virtual object 704a in the three-dimensional environment 702 (e.g., an air pinch input, an air tap input, a pinch input, a tap input, an air pinch and drag input, an air drag input, a drag input, a click and drag input, a gaze input, and / or other inputs) points to the first virtual object 704a (e.g., the input shown in Figure 7J has one or more characteristics of the input shown and described with reference to Fig. 7E ). In some embodiments, Figure 7J shows that the computer system 101 continues to receive the input initiated by the user 712 in Fig. 7E . For example, the input shown in Fig.7I is FIG. 7E to FIG. 7IA continuation of the input shown (e.g., User 712 continues to move the first virtual object 704a in the three-dimensional environment 702 by continuing to direct the gaze 708 at the first virtual object 704a while concurrently performing the air gestures and / or hand movements initiated in Fig. 7E ). In some embodiments, Fig.7I the input shown corresponds to a request to move the first virtual object 704a further in a first dimension (e.g., in the depth direction) relative to the current viewpoint of the user 712. In some embodiments, based on the first virtual object 704a moving to a greater distance relative to the current viewpoint of the user 712 in the three-dimensional environment 702 in response to the input shown in Figure 7J , the computer system 101 continues to change the visual salience of the second virtual object 704b. For example, the portion 718c continues to change in size relative to the size of the three-dimensional environment 702 (e.g., when the first virtual object 704a moves further from the current viewpoint of the user 712 in the three-dimensional environment 702, the size of the portion 718c (e.g., and the amount of the first virtual object 704a visible from the current viewpoint of the user 712) decreases relative to the three-dimensional environment 702). In some embodiments, based on the first virtual object 704a moving to a position within the spatial position thresholds 716a and 716b in the three-dimensional environment 702, the computer system 101 changes the visual salience of the second virtual object 704b such that the first virtual object 704a is fully visible from the current viewpoint of the user 712 (e.g., because the computer system 101 stops displaying the portion of the second virtual object 704b that spatially conflicts with the first virtual object 704a and displays the portion 718b with a greater amount of transparency, as shown and described with reference to FIG. 7F to FIG. 7I ).

[0236] Figure 7K Shows the second virtual object 704b displayed with a second amount of visual salience and the first virtual object 704a displayed with a first amount of visual salience based on the user 712 stopping providing the input shown and described with reference to FIG. 7E to FIG. 7J (e.g., the movement of the first virtual object 704a in the three-dimensional environment 702 corresponds to the continued movement of the first virtual object 704a provided by the user 712 in accordance with the input in FIG. 7E to FIG. 7J ). In some embodiments, the user 712 stops providing air gestures and / or hand movements relative to the three-dimensional environment 702 (e.g., with the hand 720, as shown in FIG. 7E to FIG. 7J ). In some embodiments, displaying the second virtual object 704b with a second amount of visual salience includes, as shown with reference to Figure 7GDisplaying one or more characteristics of the second virtual object 704b with a second amount of visual prominence (e.g., the portion of the second virtual object 704b that overlaps with the first virtual object 704a stops being displayed in the three-dimensional environment 702, and the portion 718b is displayed with a greater amount of transparency (e.g., compared to displaying the second virtual object 704b with a first amount of visual prominence)). In some embodiments, based on the user 712 stopping providing a reference FIG. 7E to FIG. 7K Displaying the second virtual object 704b with a second amount of visual prominence and the first virtual object 704a with a first amount of visual prominence as shown and described includes reducing the visual prominence of at least a portion of the second virtual object to a visual prominence less than a third visual prominence with respect to the three-dimensional environment in response to detecting the termination of a first input, as described in reference method 900. In some embodiments, in Figure 7K which the first virtual object 704a is displayed with a first amount of visual prominence and the second virtual object 704b is displayed with a second amount of visual prominence because the user 712 previously pointed an input to the first virtual object 704a (e.g., and has not subsequently pointed an input to the second virtual object (e.g., the first virtual object 704a is the active virtual object)). In some embodiments, in Figure 7K displaying the first virtual object 704a with a first amount of visual prominence and the second virtual object 704b with a second amount of visual prominence includes one or more characteristics of displaying the first virtual object with a first visual prominence based on determining that the first virtual object is the active virtual object regardless of whether the first virtual object overlaps with other virtual objects, as described in reference method 800. In some embodiments, in response to an input provided by the user 712 pointing to the second virtual object 704b (e.g., as shown and described Figure 7C or optionally to a blank space in the three-dimensional environment 702 (e.g., as shown and described Fig.7D ), the computer system 101 displays the second virtual object 704b with a first amount of visual prominence and the first virtual object 704a with a second amount of visual prominence (e.g., in response to this input making the second virtual object the active virtual object, and the portion 718a of the first virtual object 704a that does not display including a greater amount of transparency because the first virtual object 704a is located at a greater distance from the current viewpoint of the user 712 in the three-dimensional environment 702 than the second virtual object 704b).

[0237] Figure 7L Shows a first virtual object 704c and a second virtual object 704d displayed in the three-dimensional environment 702. In some embodiments, the first virtual object 704c has a reference FIG. 7A to FIG. 7KOne or more characteristics of the first virtual object 704a as shown and described. In some embodiments, the second virtual object 704d has reference FIG. 7A to FIG. 7K One or more characteristics of the second virtual object 704b as shown and described. As Figure 7L Shown in the top view 710 in, the difference between the distance of the first virtual object 704c from the current viewing point of the user 712 and the distance of the second virtual object 704d from the current viewing point of the user 712 is greater than as 7A to 7E Shown, the difference between the distance of the first virtual object 704a from the current viewing point of the user 712 and the distance of the second virtual object 704b from the current viewing point of the user 712 (e.g., Figure 7L In, the distance of the first virtual object 704c relative to the second virtual object 704d is greater than 7A to 7E Shown in, the distance of the first virtual object 704a relative to the second virtual object 704b). According to Figure 7L In, the difference between the distance of the first virtual object 704c from the current viewing point of the user 712 and the distance of the second virtual object 704d from the current viewing point of the user 712 is different from 7A to 7E In, the difference between the distance of the first virtual object 704a from the current viewing point of the user 712 and the distance of the second virtual object 704b from the current viewing point of the user 712, Figure 7L Shown in, the threshold amount of overlap between the first virtual object 704c and the second virtual object 704d (e.g., for displaying the corresponding virtual objects with a second amount of visual prominence) is different from FIG. 7B to FIG. 7D Shown in, the threshold amount of overlap between the first virtual object 704a and the second virtual object 704b.

[0238] As Figure 7L Shown in the top view 710 in, as compared to FIG. 7B to FIG. 7DCompared with that shown, the overlap region threshold 714a and the overlap angle threshold 714b are decreased (e.g., because the difference between the distance of the first virtual object 704c from the current view point of the user 712 and the distance of the second virtual object 704d from the current view point of the user 712 is greater than the difference between the distance of the first virtual object 704a from the current view point of the user 712 and the distance of the second virtual object 704b from the current view point of the user 712). In some embodiments, the threshold amount of overlap (e.g., the overlap region threshold 714a and / or the overlap angle threshold 714b) is increased according to the fact that the difference between the distance of the first virtual object 704c from the current view point of the user 712 and the distance of the second virtual object 704d from the current view point of the user 712 is relatively large (e.g., as compared with the first virtual object 704a and the second virtual object 704b). In some embodiments, changing the threshold amount of overlap based on the difference between the distances of the first corresponding virtual object (e.g., the first virtual object 704c) and the second corresponding virtual object (e.g., the second virtual object 704d) from the current view point of the user 712 includes one or more characteristics of the threshold amount being the first threshold amount and / or the second threshold amount according to whether the difference between the distance between the first virtual object and the current view point of the user and the distance between the second virtual object and the current view point of the user is the first distance or the second distance, as described by the reference method 800.

[0239] Figure 7M The second virtual object 704d displayed with a second amount of visual prominence and the first virtual object 704c displayed with a first amount of visual prominence are shown after the current view point of the user 712 has changed with respect to the three-dimensional environment 702. As shown in the top view 710, the current view point of the user 712 has changed the spatial arrangement (e.g., position and orientation) with respect to the three-dimensional environment 702 (e.g., as compared with FIG. 7A to FIG. 7L(as compared to that shown). In some embodiments, the movement of the current viewpoint of user 712 has one or more characteristics of the movement of the current viewpoint of the user from a first viewpoint relative to the three-dimensional environment to a second viewpoint relative to the three-dimensional environment, as described in reference method 800. As shown in the top view 710, the movement of the current viewpoint of user 712 causes the first virtual object 704c to overlap the second virtual object 704d by more than a threshold amount of overlap (e.g., more than a threshold overlap angle 714b relative to the current viewpoint of user 712). Based on the movement of the current viewpoint of user 712 causing the first virtual object 704c to overlap the second virtual object 704d by more than the threshold amount, the computer system 101 changes the visual salience of the second virtual object 704d (e.g., because prior to or during the movement of the current viewpoint of user 712, the input previously pointed to the first virtual object 704c (e.g., the first virtual object 704c is the active virtual object)). In some embodiments, according to the input previously pointing to the second virtual object 704d prior to or during the movement of the current viewpoint of user 712 (e.g., the second virtual object 704d is the active virtual object), the computer system 101 displays the first virtual object 704c with a second amount of visual salience and the second virtual object 704d with a first amount of visual salience (e.g., the computer system 101 stops displaying a first portion of the first virtual object 704c that spatially conflicts with the second virtual object 704d relative to the current viewpoint of user 712 (e.g., from the user's viewpoint, the first virtual object overlaps the second virtual object, and optionally, the first virtual object is within a threshold distance of the second virtual object in the depth dimension), and displays a second portion of the first virtual object 704c surrounding the first portion with a greater amount of transparency (e.g., including reference Fig.7D one or more characteristics of that shown and described and portion 718a).

[0240] Figure 7NShows a first virtual object 704e displayed with a first amount of visual prominence, a second virtual object 704f displayed with a second amount of visual prominence, and a third virtual object 704g displayed with the second amount of visual prominence in a three-dimensional environment 702. In some embodiments, the first virtual object 704e, the second virtual object 704f, and the third virtual object 704g have one or more characteristics of the above-mentioned first virtual object 704a and / or second virtual object 704b. As shown in the top view 710, the first virtual object 704e is displayed at a first distance relative to the current viewing point of the user 712, the second virtual object 704f is displayed at a second distance different from the first distance relative to the current viewing point of the user 712, and the third virtual object 704g is displayed at a third distance different from the first distance and the second distance relative to the current viewing point of the user 712. As shown in the top view 710, based on the difference between the distance of the first virtual object 704e from the current viewing point of the user 712 and the distance of the second virtual object 704f from the current viewing point of the user 712 being the first distance, the threshold amount of overlap between the first virtual object 704e and the second virtual object 704f corresponds to the first overlap region threshold amount 714a-1 and the first overlap angle threshold amount 714b-1. As shown in the top view 710, based on the difference between the distance of the first virtual object 704e from the current viewing point of the user 712 and the distance of the third virtual object 704g from the current viewing point of the user 712 being a second distance different from the first distance, the threshold amount of overlap between the first virtual object 704e and the third virtual object 704g corresponds to a second overlap region threshold amount 714a-2 different from the first overlap region threshold amount 714a-1, and a second overlap angle threshold amount 714b-2 different from the first overlap angle threshold amount 714b-1. In the top view 710, the first overlap region threshold amount 714a-1 is less than the second overlap region threshold amount 714a-2. In some embodiments, according to the first distance being less than the second distance, the first overlap region threshold amount 714a-1 is greater than the second overlap region threshold amount 714a-2. In the top view 710, the first overlap angle threshold amount 714b-1 is less than the second overlap angle threshold amount 714b-2. In some embodiments, according to the first distance being less than the second distance, the first overlap angle threshold amount 714b-1 is greater than the second overlap angle threshold amount 714b-2. As Figure 7NAs shown (e.g., in the top view 710), the first virtual object 704e overlaps with the second virtual object 704f (e.g., has a spatial conflict) by more than a first threshold amount (e.g., the first overlap region threshold amount 714a-1 and / or the first overlap angle threshold amount 714b-1) and overlaps with the third virtual object 704g by more than a second threshold amount (e.g., the second overlap region threshold amount 714a-2 and / or the second overlap angle threshold amount 714b-2). Based on the first virtual object 704e overlapping with the second virtual object 704f and the third virtual object 704g by more than the respective overlap threshold amounts, the computer system 101 displays the second virtual object 704f and the third virtual object 704g with a second amount of visual prominence (e.g., because the attention of the user 712 is directed to the first virtual object 704e).

[0241] As Figure 7N shown, the input points to the first virtual object 704e. In some embodiments, the input corresponds to a request to move the first virtual object 704a in a first dimension (e.g., in the depth direction relative to the current viewpoint of the user 712) in a three-dimensional environment. In some embodiments, Figure 7N the input shown in Fig. 7E has one or more characteristics of the input shown and described above.

[0242] Fig.7O shows in response to that by Figure 7NThe input provided by user 712 in [description], the movement of the first virtual object 704e in the three-dimensional environment 702. As shown in the top view 710, compared with the second virtual object 704f and the third virtual object 704g, the first virtual object 704e moves to a greater distance relative to the current viewing point of the user 712 in the three-dimensional environment 702. In some embodiments, the computer system 101 changes the visual salience of the second virtual object 704f and the third virtual object 704g during the movement (e.g., change in spatial arrangement) of the first virtual object 704e relative to the current viewing point of the user 704 based on the spatial positions of the first virtual object 704e relative to the second virtual object 704f and the first virtual object 704e relative to the third virtual object 712g. In some embodiments, the computer system 101 changes the visual salience of the second virtual object 704f independently of (e.g., not based on) the spatial position of the first virtual object 704e relative to the third virtual object 704g. In some embodiments, the computer system 101 changes the visual salience of the third virtual object 704g independently of (e.g., not based on) the spatial position of the first virtual object 704e relative to the second virtual object 704f. As shown in the top view 710, the first spatial position thresholds 716a-1 and 716b-1 are shown relative to the position of the second virtual object 704f in the three-dimensional environment 702, and the second spatial position thresholds 716a-2 and 716b-2 are shown relative to the position of the third virtual object 704g in the three-dimensional environment. In some embodiments, the spatial position thresholds 716a-1, 716a-2, 716b-1, 716b-2 have one or more characteristics of the spatial position thresholds 716a and 716b as shown and described. FIG. 7E to FIG. 7J shown and described.

[0243] In some embodiments, the computer system 101 reduces the visual salience of the second virtual object 704f by a first amount based on the spatial position of the first virtual object 704e relative to the second virtual object 704f. For example, as Fig.7O shown, reducing the visual salience of the second virtual object 704f by a first amount includes stopping displaying the portion of the second virtual object 704f that spatially conflicts with the first virtual object 704e (e.g., this portion of the second virtual object 704f has dimensions corresponding to the dimensions of the portion of the first virtual object 704e that overlaps with the second virtual object 704f) (e.g., from the user's viewing point, this portion of the second virtual object overlaps with the first virtual object, and optionally, the second virtual object is within a threshold distance of the first virtual object in the depth dimension). For example, as Fig.7OAs shown, reducing the visual prominence of the second virtual object 704f by a first amount includes displaying a portion 724a including a first size (e.g., including one or more characteristics of the portions 718a and / or 718b described above) with a greater amount of transparency relative to the three-dimensional environment 702 than a portion 724a is displayed with a visual prominence of the first amount. In some embodiments, the computer system 101 reduces the visual prominence of the third virtual object 704g by a second amount that is less than the first amount based on the spatial position of the first virtual object 704e relative to the third virtual object 704g (e.g., the second amount is less than the first amount because the difference between the distance of the first virtual object 704e from the current viewing point of the user 712 and the distance of the second virtual object 704f from the current viewing point of the user 712 is less than the difference between the distance of the first virtual object 704e from the current viewing point of the user 712 and the distance of the third virtual object 704g from the current viewing point of the user 712). For example, as Fig.7O As shown, reducing the visual prominence of the third virtual object 704g by a second amount includes ceasing to display a portion of the third virtual object 704g that spatially conflicts with the first virtual object 704e (e.g., this portion of the third virtual object 704g has a size corresponding to the size of the portion of the first virtual object 704e that overlaps the third virtual object 704g) (e.g., from the viewpoint of the user, this portion of the third virtual object overlaps the first virtual object, and optionally, the third virtual object is within a threshold distance of the first virtual object in the depth dimension). For example, as Fig.7O As shown, reducing the visual prominence of the third virtual object 704g by a second amount includes displaying a portion 724b including a second size that is less than the first size (e.g., including one or more characteristics of the portions 718a and / or 718b described above) with a greater amount of transparency relative to the three-dimensional environment 702 than a portion 724b is displayed with a visual prominence of the first amount (e.g., the second size is less than the first size because the difference between the distance of the first virtual object 704e from the current viewing point of the user 712 and the distance of the second virtual object 704f from the current viewing point of the user 712 is less than the difference between the distance of the first virtual object 704e from the current viewing point of the user 712 and the distance of the third virtual object 704g from the current viewing point of the user 712).

[0244] As Fig.7O shown, an input corresponding to a request to move the first virtual object 704e in the three-dimensional environment 702 points to the first virtual object 704e (e.g., Fig.7O the input shown in has a reference Fig. 7E(one or more characteristics of the input shown and described). In some embodiments, as the first virtual object 704e moves to different spatial positions in the three-dimensional environment 702 relative to the second virtual object 704f and / or the third virtual object 704g, the computer system 101 changes the visual salience of the second virtual object 704f and / or the third virtual object 704g during the movement of the first virtual object 704e. For example, based on the movement of the first virtual object 704e that is displayed at a position within the first spatial position thresholds 716a-1 and 716b-1 and not within the second spatial position thresholds 716a-2 and 716b-2 in the three-dimensional environment 702, the third virtual object 704g visually occludes the first virtual object 704e relative to the current viewing point of the user 712, and the second virtual object 704f does not visually occlude the first virtual object 704e relative to the current viewing point of the user 712 (e.g., the computer system 101 stops displaying the portion of the second virtual object 704f that corresponds to the first portion of the first virtual object 704e that overlaps with the second virtual object 704f, and does not stop displaying the portion of the third virtual object 704f that corresponds to the second portion of the first virtual object 704e that overlaps with the second virtual object 704 relative to the current viewing point of the user 712). For example, based on the movement of the first virtual object 704e, the first virtual object 704e is displayed at a position within the second spatial position thresholds 716a-2 and 716b-2 and not within the first spatial position thresholds 716a-1 and 716b-1 in the three-dimensional environment 702, the second virtual object 704f does not display the transparent portion 724a (e.g., because the first virtual object 704e is displayed at a position in the three-dimensional environment that corresponds to a distance closer to the current viewing point of the user 712 compared to the second virtual object 704f and not within the spatial position thresholds 716a-1 and 716b-1) and the third virtual object 704g displays the transparent portion 724b (e.g., because the first virtual object 704e is located at a position within the second spatial position thresholds 716a-2 and 716b-2).

[0245] Figure 7P An input corresponding to the user 712's attention directed to the second virtual object 704f is shown. As Figure 7P shown, the input corresponds to a gaze 708 (e.g., represented by the eyes in Figure 7P ) directed at the virtual object 704f, while the user 712 concurrently performs an air gesture with the hand 720 (e.g., an air pinch as Figure 7P shown) (e.g., for a threshold time period (e.g., 0.1, 0.2, 0.5, 1, 2, 5, or 10 seconds)). In response to the input corresponding to the attention directed to the second virtual object 704f, the computer system 101 increases the visual salience of the second virtual object 704f (e.g., compared to as Fig.7O shown (e.g., the visual prominence of the first amount) and decreases the visual prominence of the first virtual object 704e (e.g., as compared to as Fig.7O shown (e.g., the visual prominence of the second amount)). For example, in response to an input corresponding to the attention directed to the second virtual object 704f, the computer system 101 increases the opacity, brightness, color, saturation, and / or sharpness of the second virtual object 704f and decreases the opacity, brightness, color, saturation, and / or sharpness of the first virtual object 704e. Additionally, as Figure 7P shown, in response to an input corresponding to the attention directed to the virtual object 704f, the computer system 101 maintains the display of the third virtual object 704g with the same amount (e.g., the second amount and / or the decreased amount) of visual prominence (e.g., as compared to Fig.7O the amount of visual prominence of the third virtual object 704g shown in Figure 7P . In some embodiments, based on the first virtual object 704e continuing to overlap the third virtual object 704g by more than a threshold amount, the computer system 101 maintains the display of the third virtual object 704g with the second amount of visual prominence. In Fig.7O , a portion 724b of the third virtual object 704g is shown with a greater magnitude of transparency (e.g., the portion 724b is shown with an increased amount of transparency and / or a larger size) (e.g., as compared to as Fig.7O shown) (e.g., because the user 712 terminates

[0246] as Figure 7P (e.g., and Figure 7Q to Figure 7X) As shown, the first virtual object 704e, the second virtual object 704f, and the third virtual object 704g are respectively displayed together with virtual elements 740a, 740b, and 740c. In some embodiments, the virtual elements 740a - 740c can be selected by the user 712 to move the virtual objects 704e - 704g in the three - dimensional environment 702. For example, to move the virtual object 704f in the three - dimensional environment 702, the user 712 provides an input corresponding to the attention (e.g., gaze) directed at the virtual element 704a, while concurrently performing an air gesture (e.g., an air pinch such as Figure 7P shown in) that includes the movement of the user 712's hand (e.g., hand 720) relative to the three - dimensional environment 702. As Figure 7P shown, the virtual elements 740a - 740c are displayed together with virtual affordance representations (e.g., to the right of each respective virtual element 740a - 740c). In some embodiments, these virtual affordance representations can be selected by the user 712 (e.g., via an input corresponding to the attention directed at the virtual affordance representation while performing an air gesture) to stop displaying the corresponding virtual object in the three - dimensional environment 702. For example, in response to a user input corresponding to the selection of the virtual element 740a associated with the virtual affordance representation, the computer system 101 stops displaying the first virtual object 704e in the three - dimensional environment.

[0247] Figure 7Q shows the user 712 performing an input corresponding to the attention directed at the third virtual object 704g. As Figure 7Q shown, this input includes a gaze 708 directed at the virtual object 704g while the user 712 performs an air gesture (e.g., an air pinch such as Figure 7Q shown) with the hand 720 (e.g., for a threshold time period such as 0.1, 0.2, 0.5, 1, 2, 5, or 10 seconds). In response to detecting the Figure 7Q input shown, the computer system 101 displays the third virtual object 704g with an increased amount of visual salience (e.g., a first amount of visual salience) compared to that shown in Figure 7P . For example, the third virtual object 704g is displayed with a greater amount of opacity, brightness, color, saturation, and / or sharpness in Figure 7P compared to that shown in Figure 7Q . Additionally, in response to detecting the Figure 7Q input shown, the computer system 101 maintains the display of the second virtual object 704f with the same amount of visual salience (e.g., a first amount of visual salience) as shown in Figure 7P . For example, the computer system 101 does not respond to Figure 7Qreduces the visual prominence of the second virtual object 704f due to the input shown in Figure 7Q as shown, in response to detecting Figure 7Q the input shown in Figure 7P the computer system 101 maintains displaying the first virtual object 704e with the same amount of visual prominence (e.g., a second amount of visual prominence) as shown in Figure 7Q For example, the computer system 101 maintains displaying the first virtual object 704e with a reduced amount of visual prominence because the first virtual object 704e is overlapped by the third virtual object 704g (e.g., which is displayed with an increased amount of visual prominence) by more than a threshold amount. Additionally, for example, the computer system 101 maintains displaying the first virtual object 704e with a reduced amount of visual prominence because the second virtual object 704f, which was previously displayed with an increased amount of visual prominence, continues to overlap the first virtual object 704e by more than a threshold amount when detecting

[0248] Figure 7R shows Figure 7P an alternative implementation, which includes that when the second virtual object 704f overlaps the first virtual object 704e by no more than a threshold amount, the user 712 performs an input corresponding to the attention directed to the second virtual object 704f. As Figure 7R shown, in response to detecting the input corresponding to the attention directed to the second virtual object 704f, the computer system 101 displays the second virtual object 704f with an increased amount of visual prominence (e.g., a first amount of visual prominence) relative to the three-dimensional environment 702. Additionally, as Figure 7R shown, in response to detecting the input corresponding to the attention directed to the second virtual object 704f, the computer system 101 maintains displaying the first virtual object 704e with the first amount of visual prominence and the third virtual object 704g with the second amount of visual prominence. In some implementations, the computer system 101 maintains displaying the first virtual object 704e with the first amount of visual prominence because Figure 7R the second virtual object 704f pointed to by the input shown in Figure 7R overlaps the first virtual object 704e by no more than a threshold amount. In some implementations, the computer system 101 maintains displaying the third virtual object 704g with the second amount of visual prominence because when detecting Figure 7RWhen the input shown in FIG. is received, the first virtual object 704e is finally displayed with a first amount of visual prominence (e.g., since the first virtual object 704e is displayed with a first amount of visual prominence and there is more than a threshold amount of overlap between the first virtual object 704e and the third virtual object 704g, the computer system 101 displays the virtual object 704g with a second amount of visual prominence). In some embodiments, the portion 724b is displayed with increased transparency and / or an increased transparency at a maximum value (e.g., corresponding to an increased amount of transparency and / or an increased size relative to the three-dimensional environment 702), because the input corresponding to the request to move the first virtual object 704e in the three-dimensional environment 702 as shown in Fig.7O terminates (e.g., compared to the decreased transparency and / or an increased transparency at a minimum value of the portion 724b as shown during the movement of the first virtual object 704e relative to the third virtual object 704g in the three-dimensional environment 702 Fig.7O ).

[0249] Figure 7S FIG. shows a plurality of virtual elements displayed within the second virtual object 704f. Specifically, the virtual elements 730a - 730d are included within the second virtual object 704f in the three-dimensional environment 702. In some embodiments, the virtual elements 730a - 730d have one or more characteristics of virtual elements that move in the three-dimensional environment in response to the detection of a second input, as described with reference to method 800. For example, the virtual elements 730a - 730d are content such as images, files, documents, and / or text. In some embodiments, the virtual elements 730a - 730d are content associated with the respective application programs associated with the second virtual object 704f (e.g., a file (e.g., an image) storage application program). In some embodiments, the virtual elements 730a - 730d are displayed in one or more positions in the three-dimensional environment 702 that are not associated with the respective virtual objects (e.g., the virtual elements 730a - 730d are not included within the virtual objects 704e - 704g in the three-dimensional environment 702). It should be understood that although four virtual elements are shown within the second virtual object 704f, more or fewer virtual elements may also be displayed. In some embodiments, the second virtual object 704f includes a user interface that can be scrolled (e.g., by user input) by the user 712 to display one or more additional virtual elements that were not previously displayed within the second virtual object 704f.

[0250] As Figure 7SAs shown, user 712 performs an input directed at virtual element 730a. The input includes a gaze 708 directed at virtual element 730a while performing an air gesture (e.g., an air pinch) with hand 720 (e.g., the air gesture is performed for a threshold period of time, such as 0.1, 0.2, 0.5, 1, 2, 5, or 10 seconds). In some embodiments, Figure 7S the input shown in corresponds to the selection of virtual element 730a. In some embodiments, after selecting virtual element 730a, user 712 may move virtual element 730a relative to the three-dimensional environment 702 by maintaining the air gesture (e.g., the air pinch as shown in Figure 7S ) and moving hand 720 relative to the three-dimensional environment 702.

[0251] Figure 7T Shows user 712 performing an input corresponding to a request to move virtual element 730a in three-dimensional environment 702 towards a third virtual object 704g. In some embodiments, Figure 7T the input shown in is Figure 7S a continuation of the input initiated in (e.g., user 712 maintains the air gesture performed by hand 720 while moving hand 720 relative to the three-dimensional environment 702). In some embodiments, the movement of virtual element 730a-2 in the three-dimensional environment 702 corresponds to the movement of hand 720 relative to the three-dimensional environment 702 (e.g., user 712 moves hand 720 towards the position in the three-dimensional environment 702 corresponding to the third virtual object 704g). As shown in Figure 7T , when computer system 101 detects an input corresponding to a request to move virtual element 730a in three-dimensional environment 702 towards a third virtual object 704g, computer system 101 maintains the display of the first virtual object 704e with an increased amount of visual saliency (e.g., a first amount of visual saliency), the display of the second virtual object 704f with an increased amount of visual saliency, and the display of the third virtual object 704g with a decreased amount of visual saliency (e.g., a second amount of visual saliency).

[0252] In some embodiments, when moving virtual element 730a in the three-dimensional environment 702 according to the input shown in Figure 7T , computer system 101 changes the visual appearance of virtual element 730a. In some embodiments, in Figure 7S , virtual element 730a is displayed with a first visual appearance (e.g., the first visual appearance of virtual element 730a is labeled 730a-1 in Figure 7S ). For example, Figure 7S the virtual element 730a-1 shown in includes a first size, shape, and / or amount of opacity, brightness, color, saturation, and / or clarity. In some embodiments, in Figure 7TIn [context], the virtual element 730a is displayed with a second visual appearance that is different from the first visual appearance (e.g., the second visual appearance of the virtual element 730a is labeled as 730a-2 in Figure 7T . For example, Figure 7T the virtual element 730a-2 shown in [context] includes a second dimension, shape, and / or amounts of opacity, brightness, color, saturation, and / or clarity (e.g., compared to the virtual element 730a-1 shown in Figure 7S , the virtual element 730a-2 shown in Figure 7T is displayed with a smaller size, different shape, and / or more or less opacity, brightness, color, saturation, and / or clarity).

[0253] Figure 7U Shows the movement of the virtual element 730a in the three-dimensional environment 702 towards the third virtual object 704g. As Figure 7U shown, the user 712 continues to provide input corresponding to the request to move the virtual element 730a towards the third virtual object 704g as shown in Figure 7T (e.g., and initiated in Figure 7S ). In some embodiments, based on the virtual element 730a being within a threshold distance (e.g., 0.01m, 0.05m, 0.1m, 0.2m, 0.5m, or 1m) of the third virtual object 730g during the movement of the virtual element 730a in the three-dimensional environment 702, the computer system 101 moves the virtual element 730a to the third virtual object 704g (e.g., as described with reference to method 800). As Figure 7U shown, the virtual element 730a is displayed at a position in the three-dimensional environment 702 corresponding to the third virtual object 704g. For example, based on the virtual element 730a being within a threshold distance of the third virtual object 704g during the movement of the virtual element 730a in the three-dimensional environment 702, the computer system 101 moves the virtual element 730a to a position in the three-dimensional environment 702 corresponding to the third virtual object 704g. As Figure 7U shown, based on the computer system 101 moving the virtual element 730a to a position in the three-dimensional environment 702 corresponding to the third virtual object 704g, the computer system 101 maintains displaying the third virtual object 704g with a reduced amount of visual prominence. Additionally, as Figure 7U shown, the computer system 101 maintains displaying the first virtual object 704e and the second virtual object 704f with an increased amount of visual prominence.

[0254] As Figure 7UAs shown, the virtual element 730a is displayed together with the visual item 732. In some embodiments, the visual item 732 corresponds to visual feedback displayed in the three-dimensional environment 702 according to the movement of the virtual element 730a to a position corresponding to the third virtual object 704g in the three-dimensional environment 702. In some embodiments, when the visual item 732 is displayed in the three-dimensional environment 702, in accordance with the input corresponding to the user 712 terminating the request to move the virtual element 730a towards the third virtual object 704g, the computer system 101 adds the virtual element 730a to the third virtual object 704g (e.g., as described with reference to Figure 7V . The display of the visual item 732 in the three-dimensional environment 702 notifies the user 712 that if the user 712 stops providing the Figure 7U input shown (e.g., the user 712 stops performing an air pinch with the hand 720), the computer system 101 will add the virtual element 730a to the third virtual object 704g (e.g., and provide the user 712 with an opportunity to move the virtual element 730a to a different position outside a threshold distance from the third virtual object 704g in the three-dimensional environment 702 before terminating the input (e.g., to avoid the virtual element 730a being added to the third virtual object 704g)).

[0255] Figure 7V shows the addition of the virtual element 730a to the third virtual object 704g after the user 712 terminates the input corresponding to the request to move the virtual element 730a towards the third virtual object 704g. In some embodiments, adding the virtual element 730a to the third virtual object 704g includes adding the virtual element to one or more characteristics of the corresponding virtual object in the three-dimensional environment, as described with reference to the method 800. For example, as Figure 7V shown, the virtual element 730a is displayed within the third virtual object 704g. Additionally, in Figure 7V , when the virtual element 730a is added to the third virtual object 704g, the computer system 101 maintains the display of the third virtual object 704g with a second amount of visual prominence. Further, as Figure 7VAs shown, computer system 101 maintains the display of first virtual object 704e and second virtual object 704f with an increasing amount of visual prominence. In some embodiments, after virtual element 730a is added to third virtual object 704g, computer system 101 displays third virtual object 704g with an increasing amount of visual prominence (e.g., a first amount of visual prominence or a third visual prominence greater than a second visual prominence, as described in reference method 800). For example, before virtual element 730a is added to third virtual object 704g or when virtual element 730a is added to third virtual object 704g, computer system 101 does not display third virtual object 704g with an increasing amount of visual prominence (e.g., computer system 101 maintains the display of third virtual object 704g with a second amount of visual prominence). In some embodiments, in accordance with the display of third virtual object 704g with an increasing amount of visual prominence, computer system 101 displays first virtual object 704e with a decreasing amount of visual prominence (e.g., a second amount of visual prominence).

[0256] In some embodiments, adding virtual element 730a to third virtual object 704g includes changing the visual appearance of virtual element 730a. For example, virtual element 730a is displayed with a third visual appearance (e.g., the third visual appearance of virtual element 730a is labeled 730-3 in Figure 7V . Optionally, displaying virtual element 730a with a third visual appearance is different from displaying virtual element 730a with a first visual appearance and / or a second visual appearance. In some embodiments, the third visual appearance of virtual element 730a includes displaying virtual element 730a with less opacity, color, brightness, saturation, and / or sharpness compared to displaying virtual element 730a with a first visual appearance (e.g., because in Figure 7V , virtual element 730a is included in a corresponding virtual object that is displayed with less opacity, color, brightness, saturation, and / or sharpness compared to the corresponding virtual object that includes virtual element 730a when virtual element 730a is displayed with a first visual appearance). In some embodiments, virtual element 730a-3 includes a different size and / or shape compared to virtual element 730-2 (e.g., as shown in Figures 7T to 7U ).

[0257] Figure 7W is shown Figure 7UAlternative embodiments thereof include that, based on one or more criteria being met during the movement of the virtual element 730a in the three-dimensional environment 702, the computer system 101 displays the third virtual object 704g with an increased amount of visual prominence (e.g., a first amount of visual prominence or a third visual prominence greater than a second visual prominence, as described in reference method 800). In some embodiments, one or more of the criteria have one or more characteristics of one or more of the first criteria described in reference method 800. In some embodiments, based on the virtual element 730a being within a threshold distance of the third virtual object 704g (e.g., for a threshold period of time (e.g., 0.1, 0.2, 0.5, 1, 2, 5, or 10 seconds)) during the movement of the virtual element 730a in the three-dimensional environment 702, the computer system 101 displays the third virtual object 704g with an increased amount of visual prominence. In some embodiments, based on the movement of the virtual element 730a being less than a threshold amount of movement (e.g., less than 0.01 m, 0.05 m, 0.1 m, 0.2 m, 0.5 m, or 1 m relative to the three-dimensional environment 702 within 0.1, 0.2, 0.5, 1, 2, 5, or 10 seconds, or having an average speed less than 0.01 m / s, 0.02 m / s, 0.05 m / s, 0.1 m / s, 0.2 m / s, 0.5 m / s, or 1 m / s within 0.1, 0.2, 0.5, 1, 2, 5, or 10 seconds), the computer system 101 displays the third virtual object 704g with an increased amount of visual prominence. For example, when executing an input corresponding to a request to move the virtual element 730a towards the virtual object 704g, the virtual element 730a is displayed at a position corresponding to the third virtual object 704g in the three-dimensional environment 702 for a period of time exceeding a threshold (e.g., 0.1, 0.2, 0.5, 1, 2, 5, or 10 seconds). Based on the virtual element 730a being displayed at a position corresponding to the third virtual object 704g in the three-dimensional environment 702 for a period of time exceeding a threshold, the computer system 101 displays the third virtual object 704g with an increased amount of visual prominence. In some embodiments, based on a threshold amount of the third virtual object 704g being visible in the three-dimensional environment 702, the computer system 101 displays the third virtual object 704g with an increased amount of visual prominence (e.g., as described in reference to one or more first criteria, the one or more first criteria including criteria met based on a first portion of a corresponding virtual object being visible in the three-dimensional environment in method 800). As Figure 7W shown, based on the computer system 101 displaying the third virtual object 704g with an increased amount of visual prominence, the computer system 101 displays the first virtual object 704e with a decreased amount of visual prominence (e.g., a second amount of visual prominence) (e.g., which continues to be overlapped by the third virtual object 704g by more than a threshold amount). Additionally, as Figure 7WAs shown, the computer system 101 maintains the display of the second virtual object 704f with an increasing amount of visual prominence (e.g., because the second virtual object 704f is displayed with an increasing amount of visual prominence before the visual prominence of the first virtual object 704e and the third virtual object 704g changes, and the second virtual object 704f overlaps with the first virtual object 704e or the third virtual object 704g by no more than a threshold amount).

[0258] Figure 7X An input corresponding to a request to move the virtual element 730a in the three-dimensional environment 702 away from the third virtual object 704g by the user 712 is shown. In some embodiments, Figure 7X the input shown in Figure 7S is initiated and Figure 7T and Figure 7W is a continuation of the input shown in Figure 7X (e.g., the user 712 maintains an air gesture (e.g., an air pinch) and performs a movement with the hand 720 relative to the three-dimensional environment 702). As Figure 7W shown, the virtual element 730a is moved to a position in the three-dimensional environment 702 that does not correspond to the third virtual object 704g (e.g., after the computer system 101 moves the virtual element 730a to the third virtual object 704g based on the virtual element 730a being within a threshold distance of the third virtual object 704g, the user 712 moves the virtual element 730a away from the third virtual object 704g). In some embodiments, after one or more criteria are met, the user 712 moves the virtual element 730a away from the third virtual object 704g (e.g., as described with reference to Figure 7XAs shown, as the virtual element 730a is moved away from the location in the three-dimensional environment 702 corresponding to the third virtual object 704g, the computer system 101 maintains displaying the third virtual object 704g with an increased amount of visual prominence. In some embodiments, in accordance with the computer system 101 detecting the termination of an input corresponding to a request to move the virtual element 730a in the three-dimensional environment 702 when the virtual element 730a is displayed at a location away from the third virtual object 704g, the computer system 101 maintains displaying the third virtual object 704g with an increased amount of visual prominence (e.g., and displaying the first virtual object 704e with a decreased amount of visual prominence and displaying the second virtual object 704f with an increased amount of visual prominence). In some embodiments, in accordance with the computer system 101 detecting the termination of an input corresponding to a request to move the virtual element 730a in the three-dimensional environment 702 when the virtual element 730a is displayed at a location away from the third virtual object 704g, the computer system 101 abandons adding the virtual element 730a to the third virtual object 704g (e.g., because the virtual element 730a is not within a threshold distance of the third virtual object 704g and / or is not displayed at the location in the three-dimensional environment 702 corresponding to the third virtual object 704g). For example, after detecting the termination of the input, the computer system 101 maintains displaying the virtual element 730a at a location away from the third virtual object 704g. For example, after detecting the termination of the input, the computer system 101 adds (e.g., returns) the virtual element 730a to the second virtual object 704f.

[0259] Figure 7Y Illustrated are a first virtual object and a second virtual object displayed in a three-dimensional environment 702 having an input interface. In some embodiments, the first virtual object 704h and the second virtual object 704i are associated with an application to which the user 712 can provide input. For example, the first virtual object 704h is associated with a word processing application, while the second virtual object 704i is associated with a web browsing application or a search engine application. In Figure 7Y which, the first virtual object 704h and the second virtual object 704i are displayed together with virtual elements 740d and 740e, respectively. In some embodiments, the virtual elements 740d and 740e are selectable (e.g., by an input corresponding to a gaze directed at the virtual element 740d or the virtual element 740e and an air gesture) to move the first virtual object 704h or the second virtual object 704i relative to the three-dimensional environment 702 (e.g., selection of the virtual element 740d corresponds to initiating movement of the first virtual object 704h relative to the three-dimensional environment 702). As Figure 7Y shown, the virtual elements 740d and 740e are displayed with a virtual affordance representation having one or more characteristics of the virtual affordance representation described above.

[0260] As Figure 7Y shown, the input interface 736 is a virtual keyboard (e.g., the input interface 736 has one or more characteristics of the input elements described in reference method 800). In some embodiments, the input interface 736 is associated with the first virtual object 704h (e.g., the input provided through the input interface 736 corresponds to the input provided to the corresponding application associated with the first virtual object 704h). In some embodiments, based on the input interface 736 being associated with the first virtual object 704h, the user 712 can provide input through the input interface 736 to add and / or edit text in the user interface of the first virtual object 704h. Specifically, referring to Figure 7Y , the input provided by the user 712 through the input interface 736 corresponds to adding and / or editing text to the text input user interface 742a associated with the first virtual object 704h. For example, the text input user interface 742a is associated with a document. As Figure 7Y shown, the cursor 734a is shown to indicate the position in the text input user interface 742a where text will be added in response to the input provided through the input interface 736.

[0261] In some embodiments, based on the input interface 736 being associated with the first virtual object 704h, the input interface 736 is displayed at a position in the three-dimensional environment 702 based on the position of the first virtual object 704h in the three-dimensional environment 702. For example, as Figure 7Y shown, the input interface 736 is shown to be aligned with the first virtual object 704h (e.g., from the current viewpoint of the user 712 (e.g., from the current viewpoint of the user 712, the input interface 736 is centered on the first virtual object 704h)). In some embodiments, the input interface 736 is displayed at a position in the three-dimensional environment 702 independent of the position of the corresponding virtual object associated with the input interface 736. For example, in some embodiments, the input interface 736 is displayed at a position based on the current viewpoint of the user 712 (e.g., at a position aligned with the center of the current viewpoint of the user 712). As shown in the top view 710 in Figure 7Y , the input interface 736 is displayed at a position in the three-dimensional environment 702 that is closer to the current viewpoint of the user 712 than the first virtual object 704h. In some embodiments, the input interface 736 is displayed at a position in the three-dimensional environment 702 such that the user 712 can successfully interact with the input interface 736. For example, based on the input interface 736 being a virtual keyboard (e.g., as Figure 7YAs shown, the input interface 736 is displayed at a certain distance from the current viewpoint of the user 712, such that the user 712 can read the keys of the virtual keyboard. For example, the input interface 736 is displayed at a certain distance from the current viewpoint of the user 712, such that the input interface 736 is near one or more parts of the user 712 (e.g., near the hand 720), so that the user 712 can move the hand 720 to a position in the three-dimensional environment 702 corresponding to the input interface 736 (e.g., and / or one or more keys of the virtual keyboard). As Figure 7Y shown, the user 712 provides an input pointing to the input interface 736 (e.g., an air gesture corresponding to pointing to a key of the virtual keyboard (e.g., an air tap)). The input includes an air gesture (e.g., an air tap) performed by the hand 720 on a part of the input interface 736 corresponding to a key of the virtual keyboard. In some embodiments, the input includes an attention (e.g., a gaze) directed to a part of the input interface 736 corresponding to a key of the virtual keyboard when the user 712 performs an air gesture.

[0262] Figure 7Z shows the text typed in the text input user interface 742a associated with the first virtual object 704h due to the input pointing to Figure 7Y the input interface 736 therein. As Figure 7Z shown, the letter "D" is typed in the text input user interface 742a due to the input (e.g., the letter "D" corresponds to Figure 7Y the key of the virtual keyboard to which the input is directed). In Figure 7Z it, due to the addition of text in the text input user interface 742a, the position of the cursor 734a is updated within the text input user interface 742a (e.g., the updated position of the cursor 734a corresponds to the position in the text input user interface 742a where additional text will be inserted due to additional input provided through the input interface 736). In some embodiments, in response to the addition, modification, and / or removal of text that occurs in the text input user interface 742a in response to the input provided through the input interface 742, the position of the cursor 734a is further updated (e.g., in response to an input corresponding to a request to move the cursor 734a within the text input user interface 736a provided through the input interface 736, the position of the cursor 734a is further updated).

[0263] Figure 7AA shows that the user 712 provides an input corresponding to a request to move the second virtual object 704i in the three-dimensional environment 702. As Figure 7AA shown, the gaze 708 points to the virtual element 740e while the user 712 concurrently performs an air gesture (e.g., an air pinch) with the hand 720. In some embodiments, Figure 7AAThe input shown corresponds to the selection of the second virtual object 704i. When the second virtual object 704i is selected, in response to the user 712 holding an air gesture (e.g., an air pinch) with the hand 720 while performing a movement of the hand 720 relative to the three-dimensional environment 702, the second virtual object 704i can move within the three-dimensional environment 702. In some embodiments, the movement of the second virtual object 704i within the three-dimensional environment 702 is based on the movement of the hand 720 associated with the Figure 7AA input shown.

[0264] Figure 7BB An input interface 736 is shown that is displayed with a reduced amount of visual prominence in response to the movement of the second virtual object 704i that causes an overlap between the first virtual object 704h and the second virtual object 704i to exceed a threshold amount. As Figure 7BB shown, the movement of the second virtual object 704i caused by the Figure 7AA input initiated therein causes the second virtual object 704i to overlap the first virtual object 704h by more than a threshold amount. Based on the second virtual object 704i overlapping the first virtual object 704h by more than a threshold amount, the computer system 101 displays the first virtual object 704h with a second amount of visual prominence. In some embodiments, as Figure 7BB shown, since the input interface 736 is associated with the first virtual object 704h and the first virtual object 704h is displayed with a second amount of visual prominence, the computer system 101 displays the input interface 736 with a reduced amount of visual prominence. For example, compared to the Figure 7Y to Figure 7AACompared with the visual prominence display of the quantity shown in [the figure], the input interface 736 with a reduced quantity of visual prominence display includes displaying the input interface 736 with less opacity, brightness, color, saturation, and / or clarity. In some embodiments, in response to the computer system 101 detecting an input provided by the user 712 that is directed to the input interface 736 while the input interface 736 is displayed with a reduced quantity of visual prominence (e.g., and while the second virtual object 704i overlaps the first virtual object 704h by more than a threshold amount), the computer system 101 forgoes updating (e.g., by adding and / or modifying text) the text input user interface 742a based on that input. In some embodiments, the computer system 101 displays the input interface 736 with a reduced quantity of visual prominence based on the first virtual object 704h being displayed with a reduced quantity of visual prominence (e.g., in response to the first virtual object 704h being displayed with an increased quantity of visual prominence, the computer system 101 displays the input interface 736 with an increased quantity of visual prominence). In some embodiments, the computer system 101 displays the input interface 736 with a reduced visual prominence that is independent of the amount of overlap between the second virtual object 704i and the input interface 736 (e.g., the input interface 736 is displayed with a reduced quantity of visual prominence because the first virtual object 704h is displayed with a reduced quantity of visual prominence, rather than because the input interface 736 overlaps the second virtual object 704i by more than a threshold amount (e.g., as Figure 7BB shown, from the current viewpoint of the user 712, the second virtual object 704i does not overlap the input interface 736 in the three-dimensional environment 702)).

[0265] Figure 7CC Illustrated is the user 712 providing input to the text input user interface of the second virtual object 704i. In some embodiments, the text input user interface 742b of the second virtual object 704i is a text field associated with a search engine. As Figure 7CC shown, the input includes a gaze 708 directed at the text input user interface 742b while the user 712 concurrently performs an air gesture with the hand 720. In some embodiments, Figure 7CC the input shown in [the figure] corresponds to a request to associate the input interface 736 with the second virtual object 704i (e.g., the user 712 requests to use the input interface 736 to type text in the text field associated with the second virtual object 704i).

[0266] Figure 7DD Illustrated is the display of the input interface 736 associated with the second virtual object 704i in the three-dimensional environment 702 due to the input provided by the user 712 in Figure 7CC [the figure]. As Figure 7DDAs shown, the input interface 736 is displayed in the three-dimensional environment 702 with an increased amount of visual prominence (e.g., corres...

Claims

1. A method, comprising: at a computer system in communication with one or more input devices and a display generation component: displaying, via the display generation component, a plurality of virtual objects including a first virtual object and a second virtual object in a three-dimensional environment in a first spatial relationship relative to a current viewpoint of a user of the computer system, wherein displaying the first virtual object and the second virtual object in the first spatial relationship includes displaying the first virtual object and the second virtual object in a manner such that there is no overlapping portion relative to the current viewpoint of the user, and displaying the first virtual object and the second virtual object with a first visual prominence relative to the three-dimensional environment; detecting, via the one or more input devices, a first input corresponding to a request to change a spatial relationship between the first virtual object and the second virtual object from the first spatial relationship to a second spatial relationship different from the first spatial relationship relative to the current viewpoint of the user; in response to detecting the first input: based on determining that at least a portion of the first virtual object overlaps with the second virtual object by more than a threshold amount from the current viewpoint of the user, displaying, via the display generation component, a corresponding portion of a corresponding virtual object among the plurality of virtual objects with a second visual prominence less than the first visual prominence relative to the three-dimensional environment; and based on determining that the first virtual object and the second virtual object overlap by no more than the threshold amount from the current viewpoint of the user, displaying, via the display generation component, the corresponding portion of the corresponding virtual object with the first visual prominence relative to the three-dimensional environment.

2. The method according to claim 1, wherein based on determining that the first input includes an attention directed to the first virtual object, the corresponding virtual object among the plurality of virtual objects is the second virtual object, and the method further comprises: after detecting the first input, detecting a second input corresponding to an attention directed to the second virtual object; and in response to detecting the second input, based on determining that at least a portion of the first virtual object overlaps with the second virtual object by more than the threshold amount from the current viewpoint of the user: displaying the corresponding portion of the second virtual object with the first visual prominence relative to the three-dimensional environment; and displaying a corresponding portion of the first virtual object with the second visual prominence relative to the three-dimensional environment.

3. The method according to any one of claims 1 to 2, further comprising: after detecting the first input and while displaying the first virtual object with the first visual prominence, detecting a second input corresponding to an attention directed to the second virtual object; and in response to detecting the second input, based on determining that at least a portion of the first virtual object overlaps with the second virtual object by more than the threshold amount from the current viewpoint of the user: displaying the corresponding portion of the second virtual object with the first visual prominence relative to the three-dimensional environment; and Display the corresponding portion of the first virtual object with the second visual prominence relative to the three-dimensional environment.

4. The method according to claim 3, further comprising: After detecting the second input and while displaying the corresponding portion of the second virtual object with the first visual prominence, detecting a third input corresponding to an attention directed to a third virtual object among the plurality of virtual objects in the three-dimensional environment; And In response to detecting the third input, based on determining that at least a portion of the third virtual object overlaps with the second virtual object by more than the threshold amount from the user's current viewpoint: Display the corresponding portion of the second virtual object with the second visual prominence relative to the three-dimensional environment; and Maintain displaying the corresponding portion of the first virtual object with the second visual prominence relative to the three-dimensional environment.

5. The method according to claim 4, further comprising: In response to detecting the third input, based on determining that the third virtual object does not overlap with the second virtual object by more than the threshold amount from the user's current viewpoint: Maintain displaying the corresponding portion of the second virtual object with the first visual prominence relative to the three-dimensional environment; and Maintain displaying the corresponding portion of the first virtual object with the second visual prominence relative to the three-dimensional environment.

6. The method according to claim 3, further comprising: In response to detecting the second input, based on determining that at least a portion of the first virtual object overlaps with the second virtual object by more than the threshold amount from the user's current viewpoint and at least a portion of the second virtual object overlaps with a third virtual object among the plurality of virtual objects in the three-dimensional environment by more than the threshold amount from the user's current viewpoint: Display the corresponding portion of the second virtual object with the first visual prominence relative to the three-dimensional environment; Display the corresponding portion of the first virtual object with the second visual prominence relative to the three-dimensional environment; And Display the corresponding portion of the third virtual object with the second visual prominence relative to the three-dimensional environment.

7. The method according to any one of claims 1 to 6, further comprising: While displaying the plurality of virtual objects in the three-dimensional environment, displaying input elements associated with the corresponding virtual objects in the three-dimensional environment; And In response to detecting the first input: Based on determining that at least a portion of the first virtual object overlaps with the second virtual object by more than the threshold amount from the user's current viewpoint, display the input element with a third visual prominence less than the first visual prominence relative to the three-dimensional environment; And Based on determining that the first virtual object does not overlap with the second virtual object by more than the threshold amount from the user's current viewpoint, display the input element with a fourth visual prominence greater than the second visual prominence relative to the three-dimensional environment.

8. The method according to claim 7, further comprising: After detecting the first input, detect a second input corresponding to a request to display an input element associated with a third virtual object among the plurality of virtual objects in the three-dimensional environment; and In response to detecting the second input: Stop displaying the input element associated with the corresponding virtual object in the three-dimensional environment; and Display the input element associated with the third virtual object in the three-dimensional environment.

9. The method according to any one of claims 1 to 8, wherein the corresponding part of the corresponding virtual object among the plurality of virtual objects is the corresponding part of the second virtual object, and the method further includes: After detecting the first input, detect a second input corresponding to an attention directed to a position in the three-dimensional environment, the position corresponding to a blank space in the three-dimensional environment; And In response to detecting the second input, according to determining that at least a part of the first virtual object overlaps with the second virtual object by more than the threshold amount from the current view point of the user: Display the corresponding part of the second virtual object with a first visual prominence relative to the three-dimensional environment; And Display the corresponding part of the first virtual object with a second visual prominence relative to the three-dimensional environment.

10. The method according to any one of claims 1 to 9, further includes: In response to detecting the first input, move the corresponding virtual object from a first position in the three-dimensional environment to a second position in the three-dimensional environment, wherein the movement of the corresponding virtual object causes at least a part of the first virtual object to overlap with the second virtual object.

11. The method according to any one of claims 1 to 9, wherein detecting the first input includes detecting a movement of the current view point of the user from a first view point relative to the three-dimensional environment to a second view point relative to the three-dimensional environment, wherein the movement of the current view point of the user relative to the three-dimensional environment causes at least a part of the first virtual object to overlap with the second virtual object from the current view point of the user.

12. The method according to any one of claims 1 to 11, wherein: According to determining that the difference between the distance between the first virtual object and the current view point of the user and the distance between the second virtual object and the current view point of the user is a first distance, the threshold amount is a first threshold amount; And According to determining that the difference between the distance between the first virtual object and the current view point of the user and the distance between the second virtual object and the current view point of the user is a second distance different from the first distance, the threshold amount is a second threshold amount different from the first threshold amount.

13. The method according to claim 12, wherein: According to the first distance being greater than the second distance, the first threshold amount is greater than the second threshold amount; and According to the second distance being greater than the first distance, the second threshold amount is greater than the first threshold amount.

14. The method according to any one of claims 1 to 13, wherein: displaying the corresponding part of the corresponding virtual object among the plurality of virtual objects with the first visual saliency relative to the three-dimensional environment includes displaying the corresponding part of the corresponding virtual object with a first value of a first visual characteristic; and displaying the corresponding part of the corresponding virtual object among the plurality of virtual objects with the second visual saliency relative to the three-dimensional environment includes displaying the corresponding part of the corresponding virtual object with a second value less than the first value of the first visual characteristic.

15. The method according to any one of claims 1 to 14, wherein displaying the corresponding part of the corresponding virtual object with the second visual saliency relative to the three-dimensional environment includes stopping displaying a first part of the corresponding part of the corresponding virtual object in the three-dimensional environment, wherein the first part of the corresponding part of the corresponding virtual object has a relative size corresponding to the relative size of the at least part of the first virtual object that overlaps with the second virtual object.

16. The method according to claim 15, wherein displaying the corresponding part of the corresponding virtual object with the second visual saliency relative to the three-dimensional environment includes displaying the second part of the corresponding part of the corresponding virtual object with a greater amount of transparency compared to displaying the second part of the corresponding part of the corresponding virtual object with the first visual saliency, wherein the second part of the corresponding part of the corresponding virtual object surrounds the first part of the corresponding part of the corresponding virtual object.

17. The method according to any one of claims 1 to 15, wherein displaying the corresponding part of the corresponding virtual object with the second visual saliency relative to the three-dimensional environment includes, when the first virtual object is an active virtual object that overlaps with the second virtual object: stopping displaying the corresponding part of the second virtual object in the three-dimensional environment according to determining that the first virtual object is farther from the user's viewing point than the second virtual object; and maintaining displaying the corresponding part of the second virtual object in the three-dimensional environment according to determining that the first virtual object is closer to the user's viewing point than the second virtual object.

18. The method according to any one of claims 1 to 17, further comprising: in response to detecting the first input, according to determining that a first part of a third virtual object among the plurality of virtual objects overlaps with the first virtual object by more than the threshold amount and a second part of the third virtual object overlaps with the second virtual object by more than the threshold amount from the user's current viewing point: displaying a first corresponding part of a first corresponding virtual object among the plurality of virtual objects with the second visual saliency; and displaying a second corresponding part of a second corresponding virtual object among the plurality of virtual objects with the second visual saliency.

19. The method according to any one of claims 1 to 18, wherein displaying the plurality of virtual objects includes: displaying the first virtual object with the first visual prominence according to determining that the first virtual object is an active virtual object, regardless of whether the first virtual object overlaps with other virtual objects; and displaying the second virtual object with the first visual prominence according to determining that the second virtual object is an active virtual object, regardless of whether the first virtual object overlaps with other virtual objects.

20. The method according to any one of claims 1 to 19, further comprising: while displaying the corresponding virtual object with the second visual prominence, detecting a second input corresponding to a request to move a virtual element in the three-dimensional environment towards a position associated with the corresponding virtual object in the three-dimensional environment; and when the second input is detected, moving the virtual element in the three-dimensional environment according to the movement associated with the second input when displaying the corresponding virtual object with the second visual prominence.

21. The method according to claim 20, further comprising: after moving the virtual element to the position associated with the corresponding virtual object, detecting the termination of the second input via the one or more input devices; and in response to detecting the termination of the second input, adding the virtual element to the corresponding virtual object in the three-dimensional environment while maintaining the corresponding portion of the corresponding virtual object being displayed with the second visual prominence.

22. The method according to claim 20, further comprising: when the second input is detected: displaying the corresponding portion of the corresponding virtual object with a third visual prominence greater than the second visual prominence according to determining that the movement of the virtual element in the three-dimensional environment meets one or more first criteria; and maintaining the corresponding portion of the corresponding virtual object being displayed with the second visual prominence according to determining that the movement of the virtual element in the three-dimensional environment does not meet the one or more first criteria.

23. The method according to claim 22, wherein the one or more first criteria include criteria that are met when the virtual element is within a threshold distance of the corresponding virtual object.

24. The method according to any one of claims 22 to 23, wherein the one or more first criteria include criteria that are met when the movement of the virtual element is less than a threshold movement amount.

25. The method according to any one of claims 22 to 24, wherein the one or more first criteria include criteria that are met when the virtual element is within a threshold distance of the corresponding virtual object for a threshold period of time.

26. The method according to any one of claims 22 to 25, wherein the one or more first criteria include criteria that are met when a first portion of the corresponding virtual object is visible from the user's current viewpoint in the three-dimensional environment.

27. The method according to any one of claims 22 to 26, further comprising: Upon detecting the second input, move the virtual element within a threshold distance of the corresponding virtual object according to the movement associated with the second input; and Based on determining that the movement of the virtual element in the three-dimensional environment satisfies the one or more first criteria, move the virtual element to the corresponding virtual object in the three-dimensional environment before displaying the corresponding portion of the corresponding virtual object with the third visual prominence.

28. The method according to any one of claims 22 to 27, further comprising: While displaying the corresponding portion of the corresponding virtual object with the third visual prominence based on determining that the movement of the virtual element in the three-dimensional environment satisfies the one or more first criteria, detect the termination of the second input via the one or more input devices; and In response to detecting the termination of the second input, maintain displaying the corresponding portion of the corresponding virtual object with the third visual prominence based on the virtual element being at a position in the three-dimensional environment away from the corresponding virtual object.

29. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: Displaying, via the display generation component, a plurality of virtual objects including a first virtual object and a second virtual object in a three-dimensional environment in a first spatial relationship relative to a current viewpoint of a user of the computer system, wherein displaying the first virtual object and the second virtual object in the first spatial relationship includes displaying the first virtual object and the second virtual object in a manner that there is no overlapping portion relative to the current viewpoint of the user, and displaying the first virtual object and the second virtual object with a first visual prominence relative to the three-dimensional environment; Detecting, via the one or more input devices, a first input corresponding to a request to change a spatial relationship between the first virtual object and the second virtual object from the first spatial relationship to a second spatial relationship different from the first spatial relationship relative to the current viewpoint of the user; In response to detecting the first input: Based on determining that at least a portion of the first virtual object overlaps with the second virtual object by more than a threshold amount from the current viewpoint of the user, displaying, via the display generation component, a corresponding portion of a corresponding virtual object among the plurality of virtual objects with a second visual prominence less than the first visual prominence relative to the three-dimensional environment; and Based on determining that the first virtual object and the second virtual object overlap by no more than the threshold amount from the current viewpoint of the user, displaying, via the display generation component, the corresponding portion of the corresponding virtual object with the first visual prominence relative to the three-dimensional environment.

30. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: Display, via the display generation component, a plurality of virtual objects including a first virtual object and a second virtual object in a three-dimensional environment in a first spatial relationship with respect to a current viewpoint of a user of the computer system, wherein displaying the first virtual object and the second virtual object in the first spatial relationship includes displaying the first virtual object and the second virtual object in a manner that has no overlapping portions with respect to the current viewpoint of the user, and displaying the first virtual object and the second virtual object with a first visual prominence with respect to the three-dimensional environment; Detect, via the one or more input devices, a first input corresponding to a request to change a spatial relationship between the first virtual object and the second virtual object from the first spatial relationship to a second spatial relationship different from the first spatial relationship with respect to the current viewpoint of the user; In response to detecting the first input: Based on determining that at least a portion of the first virtual object overlaps with the second virtual object by more than a threshold amount from the current viewpoint of the user, display, via the display generation component, a corresponding portion of a corresponding virtual object among the plurality of virtual objects with a second visual prominence less than the first visual prominence with respect to the three-dimensional environment; And Based on determining that the first virtual object and the second virtual object overlap by no more than the threshold amount from the current viewpoint of the user, display, via the display generation component, the corresponding portion of the corresponding virtual object with the first visual prominence with respect to the three-dimensional environment.

31. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; Means for: displaying, via the display generation component, a plurality of virtual objects including a first virtual object and a second virtual object in a three-dimensional environment in a first spatial relationship with respect to a current viewpoint of a user of the computer system, wherein displaying the first virtual object and the second virtual object in the first spatial relationship includes displaying the first virtual object and the second virtual object in a manner that has no overlapping portions with respect to the current viewpoint of the user, and displaying the first virtual object and the second virtual object with a first visual prominence with respect to the three-dimensional environment; Means for: detecting, via the one or more input devices, a first input corresponding to a request to change a spatial relationship between the first virtual object and the second virtual object from the first spatial relationship to a second spatial relationship different from the first spatial relationship with respect to the current viewpoint of the user; Means for: in response to detecting the first input: Based on determining that at least a portion of the first virtual object overlaps the second virtual object by more than a threshold amount from the user's current viewpoint, displaying, via the display generation component, a corresponding portion of the corresponding virtual object among the plurality of virtual objects with a second visual prominence that is less than the first visual prominence with respect to the three-dimensional environment; and Based on determining that the first virtual object overlaps the second virtual object by no more than the threshold amount from the user's current viewpoint, displaying, via the display generation component, the corresponding portion of the corresponding virtual object with the first visual prominence with respect to the three-dimensional environment.

32. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods according to claims 1 to 28.

33. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform any of the methods according to claims 1 to 28.

34. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and means for performing any of the methods according to claims 1 to 28.

35. A method, comprising: at a computer system in communication with one or more input devices and a display generation component: displaying, via the display generation component, a first virtual object and a second virtual object in a three-dimensional environment, wherein the three-dimensional environment is visible from the current viewpoint of a user of the computer system, the second virtual object has a first visual prominence with respect to the three-dimensional environment, and the second virtual object does not spatially conflict with the first virtual object; while displaying the first virtual object and the second virtual object in the three-dimensional environment, detecting, via the one or more input devices, a first input corresponding to a request to change the position of the first virtual object in the three-dimensional environment from a first position to a second position; and in response to receiving the first input, moving the first virtual object from the first position to the second position in the three-dimensional environment, wherein moving the first virtual object from the first position to the second position includes: when at least a portion of the second virtual object spatially conflicts with at least a portion of the first virtual object with respect to the user's current viewpoint: reducing the visual prominence of at least a portion of the second virtual object with respect to the three-dimensional environment from the first visual prominence to a second visual prominence that is less than the first visual prominence; and Changing at least a portion of the second virtual object's visual salience relative to the three-dimensional environment based on a change in the spatial position of the first virtual object relative to the second virtual object during the movement of the first virtual object in the three-dimensional environment.

36. The method according to claim 35, wherein changing the visual salience of at least a portion of the second virtual object based on the spatial position of the first virtual object relative to the second virtual object includes changing the visual salience of at least a portion of the second virtual object based on a change in the depth of the first virtual object relative to the current viewpoint of the user.

37. The method according to any one of claims 35 to 36, wherein changing the visual salience of at least a portion of the second virtual object includes changing the magnitude of the second visual salience of at least a portion of the second virtual object based on the change in the spatial position of the first virtual object relative to the second virtual object during the movement of the first virtual object in the three-dimensional environment.

38. The method according to any one of claims 35 to 37, wherein changing the visual salience of at least a portion of the second virtual object includes changing the size of at least a portion of the second virtual object that is displayed with a reduced visual salience relative to the three-dimensional environment.

39. The method according to any one of claims 35 to 38, the method further comprising: Detecting an end of the first input when at least a portion of the second virtual object is displayed with a third visual salience that is less than the first visual salience relative to the three-dimensional environment while receiving the first input; And In response to detecting the end of the first input, reducing the visual salience of at least a portion of the second virtual object relative to the three-dimensional environment to a visual salience that is less than the third visual salience.

40. The method according to any one of claims 35 to 39, wherein changing the visual salience of at least a portion of the second virtual object includes reducing the visual salience of at least a portion of the second virtual object as the distance between the first virtual object and the current viewpoint of the user in the three-dimensional environment increases during the movement of the first virtual object.

41. The method according to any one of claims 35 to 40, wherein changing the visual salience of at least a portion of the second virtual object includes increasing the visual salience of at least a portion of the second virtual object as the distance between the first virtual object and the current viewpoint of the user in the three-dimensional environment increases during the movement of the first virtual object.

42. The method according to any one of claims 35 to 41, wherein changing the visual salience of at least a portion of the second virtual object includes: During the first part of the movement of the first virtual object, reducing the visual prominence of at least a portion of the second virtual object as the distance between the first virtual object and the user's current viewing point in the three-dimensional environment increases; And After the first part of the movement of the first virtual object and after reducing the visual prominence of at least a portion of the second virtual object, during the second part of the movement of the first virtual object, increasing the visual prominence of at least a portion of the second virtual object as the distance between the first virtual object and the user's current viewing point in the three-dimensional environment increases.

43. The method according to any one of claims 35 to 42, further comprising: While displaying the first virtual object and the second virtual object in the three-dimensional environment, displaying a third virtual object in the three-dimensional environment, wherein the third virtual object does not spatially conflict with the first virtual object and the second virtual object; While displaying the first virtual object, the second virtual object, and the third virtual object in the three-dimensional environment, detecting a second input corresponding to a request to change the position of the first virtual object in the three-dimensional environment from the second position to a third position; And In response to receiving the second input, and when at least a first portion of the second virtual object spatially conflicts with the first virtual object relative to the user's current viewing point and at least a second portion of the third virtual object spatially conflicts with the first virtual object: Reducing the visual prominence of at least a portion of the second virtual object relative to the three-dimensional environment from the first visual prominence to a third visual prominence lower than the first visual prominence; Reducing the visual prominence of at least a portion of the third virtual object relative to the three-dimensional environment from the first visual prominence to a fourth visual prominence lower than the first visual prominence; Changing the visual prominence of at least a portion of the second virtual object relative to the three-dimensional environment based on a change in the spatial position of the first virtual object relative to the second virtual object during the movement of the first virtual object in the three-dimensional environment; And Changing the visual prominence of at least a portion of the third virtual object relative to the three-dimensional environment based on a change in the spatial position of the first virtual object relative to the third virtual object during the movement of the first virtual object in the three-dimensional environment.

44. The method according to any one of claims 35 to 43, wherein reducing the visual prominence of at least a portion of the second virtual object relative to the three-dimensional environment to the second visual prominence includes: Stop displaying a first portion of at least a portion of the second virtual object in the three-dimensional environment, wherein the first portion of at least a portion of the second virtual object has a first size corresponding to a relative size of at least a portion of the first virtual object; And Display at least a portion of the second virtual object with a greater amount of transparency as compared to displaying a second portion of at least a portion of the second virtual object with a first visual prominence relative to the three-dimensional environment, wherein the second portion of at least a portion of the second virtual object at least partially surrounds a perimeter of the first portion of at least a portion of the second virtual object.

45. The method according to claim 44, wherein changing the visual prominence of at least a portion of the second virtual object relative to the three-dimensional environment based on the change in the spatial position of the first virtual object relative to the second virtual object comprises: Based on a change in a spatial conflict between the second virtual object and the first virtual object during the movement of the first virtual object in the three-dimensional environment, redisplay a first portion of at least a portion of the second virtual object in the three-dimensional environment, and stop displaying a third portion of at least a portion of the second virtual object that is different from the first portion in the three-dimensional environment; And Display at least a portion of the second virtual object with a greater amount of transparency as compared to displaying a fourth portion of at least a portion of the second virtual object that is different from the third portion with a first visual prominence relative to the three-dimensional environment, wherein the fourth portion of at least a portion of the second virtual object at least partially surrounds a perimeter of the third portion of at least a portion of the second virtual object.

46. The method according to any one of claims 35 to 45, wherein at least a portion of the second virtual object at least partially surrounds a perimeter of at least a portion of the first virtual object relative to the user's current viewing point.

47. The method according to any one of claims 35 to 46, further comprising: While reducing the visual prominence of at least a portion of the second virtual object, display the first virtual object at a first distance from the user's current viewing point in the three-dimensional environment, and display the second virtual object at a second distance greater than the first distance from the user's current viewing point in the three-dimensional environment.

48. The method according to any one of claims 35 to 47, further comprising: After receiving the first input, detect a second input pointing to the second virtual object; And In response to detecting the second input: Display at least a portion of the second virtual object with a first visual prominence relative to the three-dimensional environment; And Display at least a portion of the first virtual object with a third visual prominence less than the first visual prominence relative to the three-dimensional environment.

49. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: Displaying a first virtual object and a second virtual object in a three-dimensional environment via the display generation component, wherein the three-dimensional environment is visible from a current viewpoint of a user of the computer system, the second virtual object has a first visual prominence relative to the three-dimensional environment, and the second virtual object does not spatially conflict with the first virtual object; While displaying the first virtual object and the second virtual object in the three-dimensional environment, detecting, via the one or more input devices, a first input corresponding to a request to change a position of the first virtual object in the three-dimensional environment from a first position to a second position; And In response to receiving the first input, moving the first virtual object from the first position in the three-dimensional environment to the second position, wherein moving the first virtual object from the first position to the second position includes: When at least a portion of the second virtual object spatially conflicts with at least a portion of the first virtual object relative to the current viewpoint of the user: Reducing the visual prominence of at least a portion of the second virtual object relative to the three-dimensional environment from the first visual prominence to a second visual prominence less than the first visual prominence; And Changing the visual prominence of at least a portion of the second virtual object relative to the three-dimensional environment based on a change in a spatial position of the first virtual object relative to the second virtual object during the movement of the first virtual object in the three-dimensional environment.

50. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, cause the computer system to perform a method including: Displaying a first virtual object and a second virtual object in a three-dimensional environment via the display generation component, wherein the three-dimensional environment is visible from a current viewpoint of a user of the computer system, the second virtual object has a first visual prominence relative to the three-dimensional environment, and the second virtual object does not spatially conflict with the first virtual object; While displaying the first virtual object and the second virtual object in the three-dimensional environment, detecting, via the one or more input devices, a first input corresponding to a request to change a position of the first virtual object in the three-dimensional environment from a first position to a second position; And In response to receiving the first input, move the first virtual object from the first position in the three-dimensional environment to the second position, where moving the first virtual object from the first position to the second position includes: When at least a part of the first virtual object spatially conflicts with the second virtual object relative to the current viewpoint of the user: Reduce the visual prominence of at least a part of the second virtual object from the first visual prominence to a second visual prominence less than the first visual prominence relative to the three-dimensional environment; And Change the visual prominence of at least a part of the second virtual object relative to the three-dimensional environment based on a change in the spatial position of the first virtual object relative to the second virtual object during the movement of the first virtual object in the three-dimensional environment.

51. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; Means for: displaying a first virtual object and a second virtual object in a three-dimensional environment via the display generation component, where the three-dimensional environment is visible from the current viewpoint of a user of the computer system, the second virtual object has a first visual prominence relative to the three-dimensional environment, and the second virtual object does not spatially conflict with the first virtual object; Means for: while displaying the first virtual object and the second virtual object in the three-dimensional environment, detecting a first input corresponding to a request to change the position of the first virtual object in the three-dimensional environment from a first position to a second position via the one or more input devices; And Means for: in response to receiving the first input, move the first virtual object from the first position in the three-dimensional environment to the second position, where moving the first virtual object from the first position to the second position includes: When at least a part of the first virtual object spatially conflicts with the second virtual object relative to the current viewpoint of the user: Reduce the visual prominence of at least a part of the second virtual object from the first visual prominence to a second visual prominence less than the first visual prominence relative to the three-dimensional environment; And Change the visual prominence of at least a part of the second virtual object relative to the three-dimensional environment based on a change in the spatial position of the first virtual object relative to the second virtual object during the movement of the first virtual object in the three-dimensional environment.

52. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include instructions for performing any one of the methods according to claims 35 to 48.

53. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 35 to 48.

54. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and means for performing any one of the methods according to claims 35 to 48.

55. A method, comprising: at a computer system in communication with one or more input devices and a display generation component: while displaying virtual content via the display generation component, wherein at least a portion of the virtual content obscures the visibility of at least a portion of the physical environment of a user of the computer system, detecting a passthrough visibility event via the one or more input devices; and in response to detecting the passthrough visibility event, replacing the display of at least a portion of the virtual content with a representation of a real-world object in the physical environment of the user via the display generation component, wherein presenting the representation of the real-world object includes: presenting the representation of the real-world object with a first visual effect applied to the representation of the real-world object according to determining that the state of the virtual content is a first state; and not presenting the representation of the real-world object with the first visual effect applied to the representation of the real-world object according to determining that the state of the virtual content is not the first state.

56. The method according to claim 55, wherein detecting the passthrough visibility event includes detecting, via the one or more input devices, that a portion of the user has moved into at least a portion of the physical environment, and presenting the representation of the real-world object includes presenting a representation of the portion of the user.

57. The method according to claim 55, wherein detecting the passthrough visibility event includes detecting, via the one or more input devices, that at least a portion of the virtual content has a spatial conflict with at least a portion of the real-world object, and presenting the representation of the real-world object includes presenting at least a portion of the real-world object.

58. The method according to claim 55, wherein detecting the passthrough visibility event includes detecting, via the one or more input devices, that the real-world object has moved within a threshold distance of the position of the user in the physical environment.

59. The method according to claim 55, wherein detecting the passthrough visibility event includes detecting, via the one or more input devices, that the user's viewing point points to a boundary of the virtual content, wherein the real-world object is covered by at least a portion of the virtual content, and wherein at least a portion of the virtual content is adjacent to the boundary of the virtual content.

60. The method according to claim 55, wherein detecting the passthrough visibility event includes detecting, via the one or more input devices, that the user's viewing point has moved more than a threshold distance from the position of the user's viewing point when the virtual content was first displayed.

61. The method according to claim 55, wherein detecting the passthrough visibility event includes detecting, via the one or more input devices, a user input corresponding to a request to stop displaying the application associated with the virtual content.

62. The method according to any one of claims 55 to 61, wherein in response to detecting the passthrough visibility event and based on determining that the state of the virtual content is a second state, wherein in the second state the virtual content includes an application window, based on the state of the virtual content being the second state, the representation of the real-world object is not presented with the visual effects applied to the representation of the real-world object.

63. The method according to any one of claims 55 to 61, wherein when the virtual content includes a user interface for inputting information associated with an application, the virtual content is in a second state, the user interface is displayed simultaneously with the application window associated with the application, and in response to detecting the passthrough visibility event and based on determining that the virtual content is in the second state, the representation of the physical object is presented with a second visual effect different from the first visual effect.

64. The method according to any one of claims 55 to 61, wherein the virtual content is in the first state at least partially based on determining that the user's attention is directed to the virtual content.

65. The method according to any one of claims 55 to 61 and 64, wherein applying the first visual effect includes reducing the visual prominence of the representation of the real-world object.

66. The method according to any one of claims 55 to 61, wherein the first visual effect includes a coloring effect applied to the representation of the real-world object.

67. The method according to claim 66, wherein the virtual content includes virtual media content, and the coloring effect is associated with one or more colors included in the virtual media content.

68. The method according to any one of claims 66 to 67, wherein the virtual content is associated with an application, and the coloring effect is selected based on the application associated with the virtual content.

69. The method according to any one of claims 55 to 68, wherein the first visual effect includes a change in the saturation of the representation of the real-world object.

70. The method according to any one of claims 55 to 69, wherein the virtual content includes an application window and a virtual environment, and the first visual effect is at least partially based on the application window and the virtual environment.

71. The method according to any one of claims 55 to 70, wherein the virtual content includes a virtual environment, and presenting the representation of the real-world object with the first visual effect applied to the representation of the real-world object includes: presenting the representation of the real-world object with the first visual effect including a first shading effect associated with the first virtual environment according to determining that the virtual environment is the first virtual environment, and presenting the representation of the real-world object with the first visual effect including a second shading effect associated with the second virtual environment different from the first virtual environment according to determining that the virtual environment is a second virtual environment different from the first virtual environment, the second shading effect being different from the first shading effect.

72. The method according to claim 71, further comprising: before presenting the representation of the real-world object with the first visual effect applied to the representation of the real-world object, not presenting the representation of the real-world object with the first visual effect applied to the representation of the real-world object; detecting a request to display the virtual environment when not presenting the representation of the real-world object with the first visual effect applied to the representation of the real-world object; and in response to detecting the request to display the virtual environment, displaying the virtual environment, wherein presenting the representation of the real-world object with the first visual effect applied to the representation of the real-world object is based on displaying the virtual environment.

73. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: while displaying virtual content via the display generation component, wherein at least a portion of the virtual content obscures the visibility of at least a portion of the physical environment of a user of the computer system, detecting a see-through visibility event via the one or more input devices; and in response to detecting the see-through visibility event, replacing the display of at least a portion of the virtual content with a representation of a real-world object in the physical environment of the user via the display generation component, wherein presenting the representation of the real-world object includes: presenting the representation of the real-world object with a first visual effect applied to the representation of the real-world object according to determining that the state of the virtual content is a first state; and not presenting the representation of the real-world object with the first visual effect applied to the representation of the real-world object according to determining that the state of the virtual content is not the first state.

74. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: While displaying virtual content via the display generation component, where at least a portion of the virtual content obscures the visibility of at least a portion of the physical environment of a user of the computer system, detect a passthrough visibility event via the one or more input devices; And In response to detecting the passthrough visibility event, replacing at least a portion of the display of the virtual content with a representation of a real-world object in the user's physical environment via the display generation component, wherein presenting the representation of the real-world object includes: Presenting the representation of the real-world object with a first visual effect applied to the representation of the real-world object according to determining that the state of the virtual content is a first state; And Not presenting the representation of the real-world object with the first visual effect applied to the representation of the real-world object according to determining that the state of the virtual content is not the first state.

75. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; Means for: while displaying virtual content via the display generation component, wherein at least a portion of the virtual content obscures at least a portion of the visibility of the user's physical environment of the computer system, detecting a passthrough visibility event via the one or more input devices; And Means for: in response to detecting the passthrough visibility event, replacing at least a portion of the display of the virtual content with a representation of a real-world object in the user's physical environment via the display generation component, wherein presenting the representation of the real-world object includes: Presenting the representation of the real-world object with a first visual effect applied to the representation of the real-world object according to determining that the state of the virtual content is a first state; And Not presenting the representation of the real-world object with the first visual effect applied to the representation of the real-world object according to determining that the state of the virtual content is not the first state.

76. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one of the methods according to claims 55 to 72.

77. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 55 to 72.

78. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And Means for performing any of the methods according to claims 55 to 72.

79. A method, comprising: At a computer system that communicates with a display generation component: While a second part of a three-dimensional environment with a background behind virtual content is visible, when displaying the virtual content in a first part of the three-dimensional environment via the display generation component: Detecting an event corresponding to the virtual content; And In response to detecting the event corresponding to the virtual content: Presenting the background with a first visual effect applied to the background according to determining that the state of the background is a first state; and Not presenting the background with the first visual effect according to determining that the state of the background is not the first state.

80. The method according to claim 79, wherein the background includes a representation of the physical environment of a user of the computer system, and the background includes a representation of a virtual environment.

81. The method according to any one of claims 79 to 80, wherein the background includes a virtual environment.

82. The method according to any one of claims 79 to 81, wherein the virtual content includes visual media content, and wherein detecting the event includes detecting that the state of the visual media content is a first state.

83. The method according to any one of claims 79 to 81, wherein detecting the event includes detecting that a user's attention is directed to the virtual content.

84. The method according to any one of claims 79 to 83, wherein the background includes a virtual environment, and the first state of the background corresponds to a first time-of-day setting of the virtual environment, and a second state of the virtual environment corresponds to a second time-of-day setting different from the first time-of-day setting.

85. The method according to any one of claims 79 to 84, wherein the first state corresponds to a daytime time-of-day setting, the second state corresponds to a nighttime time-of-day setting and the background is in the second state, and wherein not presenting the background with the first visual effect according to determining that the state is not the first state includes not presenting the background with any visual effect based on the background being in the second state.

86. The method according to claim 85, wherein detecting the event includes detecting that the state of the background has changed to the first state while the virtual content is being displayed.

87. The method according to any one of claims 84 to 85, wherein detecting the event includes detecting that the state of the background has changed to the second state while the virtual content is being displayed, and wherein not presenting the background with the first visual effect includes not presenting the background with a visual effect corresponding to the second state, regardless of whether the virtual content is associated with the first visual effect.

88. The method according to any one of claims 79 to 87, wherein the virtual content includes media content and the background is in the first state, the method further comprising: When displaying the media content in the three-dimensional environment including the background and when the media content is not being played, detecting a first input corresponding to a request to play the media content via one or more input devices; In response to detecting the first input, playing the media content in the three-dimensional environment including the background; When playing the media content and displaying the media content in the three-dimensional environment including the background, detecting a second input corresponding to a request to display the media content at a corresponding position of the media content in the background via the one or more input devices; And In response to detecting the second input, displaying the media content at the corresponding position of the media content in the background and changing the state of the background to the second state.

89. The method according to claim 88, wherein detecting the event includes detecting the first input.

90. The method according to any one of claims 79 to 89, wherein presenting the background with the first visual effect includes: According to determining that the background includes a first virtual environment, presenting the background with a first corresponding visual effect corresponding to the first virtual environment; And According to determining that the background includes a second virtual environment different from the first virtual environment, presenting the background with a second corresponding visual effect different from the first corresponding visual effect corresponding to the second virtual environment.

91. The method according to claim 90, wherein not presenting the background with the first visual effect according to determining that the background is not in the first state includes presenting the background with a third corresponding visual effect different from the first corresponding visual effect and the second corresponding visual effect, and the third corresponding visual effect is independent of whether the background includes the first virtual environment or the second virtual environment.

92. The method according to any one of claims 79 to 91, wherein presenting the background with the first visual effect includes dimming the background.

93. The method according to any one of claims 79 to 92, wherein presenting the background with the first visual effect includes applying color shading to the background.

94. The method according to any one of claims 79 to 93, wherein the background includes a representation of the physical environment of the user of the computer system, the method further comprising: Displaying second virtual content in the three-dimensional environment, and wherein presenting the background with the first visual effect includes presenting the representation of the physical environment with a combination of the first visual effect and the second visual effect.

95. The method according to claim 94, wherein the background includes a first virtual environment, and wherein presenting the background with the first visual effect includes presenting the first virtual environment with the combination of the first visual effect and the second visual effect.

96. The method according to any one of claims 94 to 95, wherein presenting the background with the first visual effect includes: presenting the background with a first corresponding visual effect according to determining that the state of the virtual content is a first state of the virtual content; and abandoning presenting the background with the first corresponding visual effect according to determining that the state of the virtual content is a second state of the virtual content.

97. The method according to any one of claims 79 to 96, wherein the background includes a first virtual environment, and presenting the background with the first visual effect includes presenting the background with a first amount of the first visual effect applied to the background, regardless of the immersion level of the first virtual environment.

98. The method according to any one of claims 79 to 97, wherein the background includes a virtual environment, and the method further includes: while displaying the virtual content and while presenting the background with the first visual effect, detecting that the orientation of the user's viewpoint of the computer system has changed from a first orientation relative to the virtual environment to a second orientation relative to the virtual environment; and in response to detecting that the orientation of the user's viewpoint has changed relative to the virtual environment and according to determining that the second orientation is greater than a threshold orientation away from the virtual environment, reducing the first visual effect applied to the background.

99. The method according to claim 98, wherein: according to determining that the immersion level of the virtual environment is a first immersion level, the threshold orientation is a first threshold orientation; and according to determining that the immersion level of the virtual environment is a second immersion level greater than the first immersion level, the threshold orientation is a second threshold orientation greater than the first threshold orientation.

100. The method according to any one of claims 79 to 99, wherein the background includes a virtual environment and a representation of the physical environment of the computer system, and wherein presenting the background with the first visual effect includes presenting a part of the three-dimensional environment with the first visual effect, and the part of the three-dimensional environment includes a transition region between the virtual environment and the representation of the physical environment.

101. The method according to any one of claims 79 to 100, the method further includes: while presenting the background with the first virtual effect according to determining that the background is in the first state, detecting a passthrough visibility event associated with a real-world object in the physical environment of the computer system; and in response to detecting the passthrough visibility event, presenting a representation of the real-world object with the first visual effect applied to the representation of the real-world object.

102. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and One or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for the following operations: When displaying the virtual content in the first part of the three-dimensional environment via the display generating component while the background is visible in the second part of the three-dimensional environment behind the virtual content: Detect an event corresponding to the virtual content; And In response to detecting the event corresponding to the virtual content: Present the background with a first visual effect applied to the background according to determining that the state of the background is a first state; and Do not present the background with the first visual effect according to determining that the state of the background is not the first state.

103. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generating component and one or more input devices, cause the computer system to perform a method including the following operations: When displaying the virtual content in the first part of the three-dimensional environment via the display generating component while the background is visible in the second part of the three-dimensional environment behind the virtual content: Detect an event corresponding to the virtual content; and In response to detecting the event corresponding to the virtual content: Present the background with a first visual effect applied to the background according to determining that the state of the background is a first state; and Do not present the background with the first visual effect according to determining that the state of the background is not the first state.

104. A computer system in communication with a display generating component and one or more input devices, the computer system including: One or more processors; Memory; Means for the following operations: When displaying the virtual content in the first part of the three-dimensional environment via the display generating component while the background is visible in the second part of the three-dimensional environment behind the virtual content: Means for the following operation: Detect an event corresponding to the virtual content; And Means for the following operation: In response to detecting the event corresponding to the virtual content: Present the background with a first visual effect applied to the background according to determining that the state of the background is a first state; and Do not present the background with the first visual effect according to determining that the state of the background is not the first state.

105. A computer system in communication with a display generating component and one or more input devices, the computer system including: One or more processors; Memory; And One or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any one of the methods according to claims 79 to 101.

106. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform any of the methods according to claims 79 to 101.

107. A computer system in communication with a display generation component and one or more input devices, the computer system including: one or more processors; a memory; and means for performing any of the methods according to claims 79 to 101.

108. A method, comprising: at a computer system in communication with one or more input devices and a display generation component, the computer system being associated with a user: while displaying a plurality of virtual objects in a three-dimensional environment via the display generation component, detecting, via the one or more input devices, a transfer of the user's attention to a first virtual object among the plurality of virtual objects, wherein the first virtual object is associated with a first visual effect; and in response to detecting the transfer of the user's attention to the first virtual object: displaying the first visual effect applied to the three-dimensional environment based on determining that the first virtual object is active and meets one or more first criteria; and abandoning display of the first visual effect applied to the three-dimensional environment based on determining that the first virtual object is not active.

109. The method according to claim 108, further comprising: prior to detecting the transfer of the user's attention to the first virtual object, presenting the three-dimensional environment without any visual effect applied to the three-dimensional environment.

110. The method according to any one of claims 108 to 109, further comprising: prior to detecting the transfer of the user's attention to the first virtual object via the one or more input devices, displaying a second visual effect applied to the three-dimensional environment based on determining that a second criterion is met, the second visual effect being different from the first visual effect.

111. The method according to claim 110, wherein the second visual effect is a visual effect selected based on an application program that was active prior to detecting the transfer of the user's attention to the first virtual object.

112. The method according to claim 110, wherein the second visual effect is a visual effect selected based on a system visual effect that is part of the enhanced three-dimensional environment in which the first virtual object is displayed.

113. The method according to any one of claims 108 to 112, wherein the first virtual object is not active prior to detecting the transfer of the user's attention to the first virtual object, the method further comprising: After detecting the transfer of the user's attention to the first virtual object and while the user's attention is directed to the first virtual object, detect user input via the one or more input devices; And In response to detecting the user input while the user's attention is directed to the first virtual object, change the state of the first virtual object to the active state, Wherein the display of the first visual effect applied to the three-dimensional environment is based on the state of the first virtual object being in the active state.

114. The method according to claim 113, wherein before detecting the user input, the first virtual object is at least partially displayed behind the second virtual object and occluded by the second virtual object relative to the user's current viewing point.

115. The method according to claim 113, wherein after detecting the transfer of the user's attention to the first virtual object and before detecting the user input, the first virtual object is not in the active state.

116. The method according to claim 115, wherein displaying the first virtual object includes: Displaying the first virtual object with a first visual appearance according to determining that the first virtual object is in the active state; And Displaying the first virtual object with a second visual appearance different from the first visual appearance according to determining that the first virtual object is not in the active state.

117. The method according to claim 113, further comprising: While displaying the first virtual object and while the first virtual object is in the active state, display a second virtual object among the plurality of virtual objects, wherein the second virtual object is in the active state.

118. The method according to any one of claims 108 to 117, further comprising: While displaying the first visual effect applied to the three-dimensional environment according to determining that the first virtual object is in the active state, detect an event corresponding to stopping the display of the first virtual object via the one or more input devices; And In response to detecting the event, stop displaying the first virtual object and stop displaying the first visual effect applied to the three-dimensional environment.

119. The method according to any one of claims 108 to 118, wherein displaying the first visual effect applied to the three-dimensional environment includes gradually changing the visual salience of the first visual effect to a final visual salience through a plurality of intermediate states over a period of time.

120. The method according to claim 119, wherein the visual salience of the first visual effect changes in a manner simulating a critically damped spring over the period of time.

121. The method according to any one of claims 119 to 120, wherein the duration of the change in the visual prominence of the first visual effect occurs after a time delay after detecting the transfer of the user's attention to the first virtual object and determining that the first virtual object is in the active state.

122. The method according to any one of claims 119 to 121, wherein changing the visual prominence of the first visual effect to the final visual prominence during the duration includes: changing the visual prominence of the first visual effect to the final visual prominence within a first duration according to determining that the transfer of the user's attention is from a portion of the three-dimensional environment associated with a second visual effect different from the first visual effect; and changing the visual prominence of the first visual effect to the final visual prominence within a second duration different from the first duration according to determining that the transfer of the user's attention is from a portion of the three-dimensional environment not associated with a visual effect.

123. The method according to any one of claims 108 to 122, further comprising: after detecting the transfer of the user's attention to the first virtual object, detecting the transfer of the user's attention to a second virtual object among the plurality of virtual objects; and in response to detecting the transfer of the user's attention to the second virtual object: displaying the second visual effect applied to the three-dimensional environment according to determining that the second virtual object is associated with a second visual effect, and abandoning the display of the second visual effect applied to the three-dimensional environment according to determining that the second virtual object is not associated with the second visual effect.

124. The method according to any one of claims 108 to 123, wherein the first visual effect is displayed according to determining that the first virtual object is in the active state, regardless of whether the three-dimensional environment is associated with a second visual effect different from the first visual effect.

125. The method according to claim 124, further comprising: while displaying the first visual effect applied to the three-dimensional environment, detecting the transfer of the user's attention away from the first virtual object; and in response to detecting the transfer of the user's attention away from the first virtual object, displaying the second visual effect applied to the three-dimensional environment.

126. The method according to any one of claims 108 to 125, wherein the first criterion includes a criterion satisfied when the three-dimensional environment does not include a virtual environment associated with a second visual effect different from the first visual effect, and the method further comprises: In response to detecting the transfer of the user's attention to the first virtual object and based on the determination that the first criterion is not satisfied because the three-dimensional environment includes the virtual environment associated with the second visual effect, display the second visual effect applied to the three-dimensional environment, regardless of whether the state of the first virtual object is active.

127. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for the following operations: While displaying a plurality of virtual objects in a three-dimensional environment via the display generation component, detect, via the one or more input devices, a transfer of the user's attention of the computer system to a first virtual object among the plurality of virtual objects, wherein the first virtual object is associated with a first visual effect; And In response to detecting the transfer of the user's attention to the first virtual object: Based on the determination that the first virtual object is in an active state and meets one or more first criteria, display the first visual effect applied to the three-dimensional environment; And Based on the determination that the first virtual object is not in the active state, refrain from displaying the first visual effect applied to the three-dimensional environment.

128. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, cause the computer system to perform a method including the following operations: While displaying a plurality of virtual objects in a three-dimensional environment via the display generation component, detecting, via the one or more input devices, a shift of the user's attention to a first virtual object among the plurality of virtual objects, wherein the first virtual object is associated with a first visual effect; And In response to detecting the transfer of the user's attention to the first virtual object: Based on the determination that the first virtual object is in an active state and meets one or more first criteria, display the first visual effect applied to the three-dimensional environment; And Based on the determination that the first virtual object is not in the active state, refrain from displaying the first visual effect applied to the three-dimensional environment.

129. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; Means for: while displaying a plurality of virtual objects in a three-dimensional environment via the display generation component, detecting, via the one or more input devices, a transfer of the user's attention to a first virtual object among the plurality of virtual objects, wherein the first virtual object is associated with a first visual effect; And Means for: in response to detecting the transfer of the user's attention to the first virtual object: Based on the determination that the first virtual object is in an active state and meets one or more first criteria, display the first visual effect applied to the three-dimensional environment; And In accordance with the determination that the first virtual object is not in the active state, abandon the display of the first visual effect applied to the three-dimensional environment.

130. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods according to claims 108 to 126.

131. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform any of the methods according to claims 108 to 126.

132. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; a memory; and means for performing any of the methods according to claims 108 to 126.

133. A method, comprising: at a computer system in communication with a display generation component and one or more input devices: while displaying a first user interface element in the environment from a current viewpoint of a user of the computer system via the display generation component, detecting via the one or more input devices that a first event has occurred; and in response to detecting that the first event has occurred, displaying a second user interface element different from the first user interface element in the environment, wherein: from the current viewpoint of the user of the computer system, the second user interface element at least partially overlaps the first user interface element; in accordance with the determination that the second user interface element is a first type of user interface element that overlaps the first user interface element, the first user interface element is visually weakened relative to the environment in a first manner; and in accordance with the determination that the second user interface element is a second type of user interface element that is different from the first type of user interface element and overlaps the first user interface element, the first user interface element is not visually weakened relative to the environment in the first manner.

134. The method according to claim 133, wherein: the first user interface element includes first content; and visually weakening the first user interface element relative to the environment in the first manner includes dimming the first content of the first user interface element relative to the environment.

135. The method according to any one of claims 133 to 134, wherein visually weakening the first user interface element relative to the environment in the first manner includes increasing the transparency of the first user interface element relative to the environment, such that a first portion of the environment occluded by the first user interface is visible relative to the user's viewing point.

136. The method according to any one of claims 133 to 135, wherein, based on determining that the second user interface element is a second type of user interface element that overlaps with the first user interface element, the first user interface element is visually weakened relative to the environment in a second manner different from the first manner.

137. The method according to any one of claims 133 to 136, wherein visually weakening the first user interface element relative to the environment in the first manner includes visually weakening the first user interface element relative to a three-dimensional environment in which the first user interface element is displayed via the display generating component.

138. The method according to any one of claims 133 to 137, wherein: In response to detecting that the first event has occurred, based on determining that the second user interface element is a third type of user interface element different from the first type of user interface element and the second type of user interface element that overlaps with the first user interface element, the first user interface element is visually weakened relative to the environment in a third manner different from the first manner.

139. The method according to any one of claims 133 to 138, further comprising: While displaying the first user interface element in the environment, detecting via the one or more input devices that a second event has occurred; And In response to detecting that the second event has occurred, displaying in the environment a third user interface element different from the first user interface element and the second user interface element, wherein: Based on determining that the third user interface element is a third type of user interface element different from the first type of user interface element and the second type of user interface element, the first user interface element is visually weakened relative to the environment in a third manner, regardless of whether the third user interface element overlaps with the first user interface element.

140. The method according to claim 139, wherein: In response to detecting that the second event has occurred, the second user interface element is visually weakened relative to the environment in the third manner.

141. The method according to any one of claims 139 to 140, wherein: In response to detecting that the second event has occurred, based on determining that the third user interface element is a fourth type of user interface element different from the first type of user interface element, the second type of user interface element, and the third type of user interface element, the first user interface element is visually attenuated relative to the environment in a fourth manner different from the third manner, regardless of whether the third user interface element overlaps with the first user interface element.

142. The method according to any one of claims 133 to 141, wherein: In response to detecting that the first event has occurred, based on determining that the second user interface element is a fourth type of user interface element different from the first type of user interface element and the second type of user interface element that overlaps with the first user interface element, the first user interface element is not visually attenuated relative to the environment.

143. The method according to claim 142, wherein the fourth type of user interface includes a virtual keyboard.

144. The method according to claim 143, further comprising: While displaying the second user interface element based on determining that the second user interface element is the fourth type of user interface element without visually attenuating the first user interface element relative to the environment, detecting via the one or more input devices that the second event has occurred; And In response to detecting that the second event has occurred, displaying in the environment a third user interface element different from the first user interface element and the second user interface element, wherein: From the current viewpoint of the user of the computer system, the third user interface element at least partially overlaps with the first user interface element; Based on determining that the third user interface element is the first type of user interface element that overlaps with the first user interface element, the first user interface element and the second user interface element are visually attenuated relative to the environment in the first manner; and Based on determining that the third user interface element is the second type of user interface element that overlaps with the first user interface element, the first user interface element and the second user interface element are not visually attenuated relative to the environment in the first manner.

145. The method according to any one of claims 133 to 144, comprising: In response to detecting that the first event has occurred, based on determining that when the second user interface element is displayed, the second user interface element does not at least partially overlap with the first user interface element from the current viewpoint of the user, the first user interface element is not visually attenuated relative to the environment.

146. The method according to any one of claims 133 to 145, wherein visually attenuating the first user interface element relative to the environment in the first manner includes: Based on determining that the second user interface element overlaps a first portion of the first user interface element, visually weakening the first portion of the first user interface element relative to the environment, without visually weakening a second portion of the first user interface element that is different from the first portion and that is not overlapped by the second user interface element relative to the environment.

147. The method according to any one of claims 133 to 146, wherein visually weakening the first user interface element relative to the environment in the first manner includes: Based on determining that the second user interface element overlaps a first portion of the first user interface element, visually weakening the first portion of the first user interface element relative to the environment, and visually weakening a second portion of the first user interface element that is different from the first portion and that is not overlapped by the second user interface element relative to the environment.

148. The method according to any one of claims 133 to 147, wherein detecting that the first event has occurred includes detecting a first warning event at the computer system.

149. The method according to any one of claims 133 to 148, wherein detecting that the first event has occurred includes detecting a first input corresponding to a request to display the second user interface element in the environment via the one or more input devices.

150. The method according to claim 149, wherein detecting the first input includes detecting the user's gaze directed at a predetermined portion of the display generation component.

151. The method according to claim 150, further comprising: While displaying the second user interface element in the environment in response to detecting the first input, detecting a second input directed at the second user interface element via the one or more input devices; And In response to detecting the second input: Stop displaying the second user interface element via the display generation component; And Display a third user interface element different from the first user interface element and the second user interface element in the environment via the display generation component, wherein the third user interface element is associated with the second user interface element.

152. The method according to any one of claims 149 to 151, wherein detecting the first input includes detecting a selection of a hardware button of the computer system via the one or more input devices.

153. The method according to any one of claims 133 to 152, wherein the first user interface element corresponds to a virtual application window associated with a corresponding application program running on the computer system.

154. The method according to any one of claims 133 to 153, wherein the first user interface element corresponds to an immersive virtual object.

155. The method according to claim 154, further comprising: While displaying the second user interface element in the environment in response to detecting that the first event has occurred, detecting, via the one or more input devices, a corresponding input directed to the first user interface element in the environment; and In response to detecting the corresponding input: Abandon performing an operation associated with the corresponding input directed to the first user interface element.

156. The method according to any one of claims 154 to 155, wherein displaying the first user interface element in the environment includes applying a visual effect to the first portion of the user based on determining that the first portion of the user is positioned within the environment relative to the viewpoint of the user, the visual effect causing the first portion of the user to be displayed as a corresponding virtual representation in the environment relative to the viewpoint of the user, the method further including: In response to detecting that the first event has occurred: Based on determining that the first portion of the user is positioned within the environment relative to the viewpoint of the user when detecting that the first event has occurred, stop applying the visual effect to the first portion of the user, such that the corresponding virtual representation is no longer displayed in the environment relative to the viewpoint of the user.

157. A computer system in communication with a display generation component and one or more input devices, the computer system including: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: While displaying a first user interface element in the environment from a current viewpoint of a user of the computer system via the display generation component, detecting that a first event has occurred via the one or more input devices; And In response to detecting that the first event has occurred, displaying a second user interface element in the environment that is different from the first user interface element, wherein: Viewed from the current viewpoint of the user of the computer system, the second user interface element at least partially overlaps the first user interface element; Based on determining that the second user interface element is a first type of user interface element that overlaps the first user interface element, the first user interface element is visually attenuated relative to the environment in a first manner; and Based on determining that the second user interface element is a second type of user interface element that is different from the first type of user interface element and overlaps the first user interface element, the first user interface element is not visually attenuated relative to the environment in the first manner.

158. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method including: While displaying a first user interface element in the environment from a current viewpoint of a user of the computer system via the display generation component, detect that a first event has occurred via the one or more input devices; And In response to detecting that the first event has occurred, a second user interface element different from the first user interface element is displayed in the environment, where: From the current viewpoint of the user of the computer system, the second user interface element at least partially overlaps with the first user interface element; Based on determining that the second user interface element is a first type of user interface element that overlaps with the first user interface element, the first user interface element is visually attenuated relative to the environment in a first manner; and Based on determining that the second user interface element is a second type of user interface element different from the first type of user interface element that overlaps with the first user interface element, the first user interface element is not visually attenuated relative to the environment in the first manner.

159. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; Means for: while displaying a first user interface element in the environment from the current viewpoint of the user of the computer system via the display generation component, detecting via the one or more input devices that the first event has occurred; And Means for: in response to detecting that the first event has occurred, displaying a second user interface element different from the first user interface element in the environment, where: From the current viewpoint of the user of the computer system, the second user interface element at least partially overlaps with the first user interface element; Based on determining that the second user interface element is a first type of user interface element that overlaps with the first user interface element, the first user interface element is visually attenuated relative to the environment in a first manner; and Based on determining that the second user interface element is a second type of user interface element different from the first type of user interface element that overlaps with the first user interface element, the first user interface element is not visually attenuated relative to the environment in the first manner.

160. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any one of the methods according to claims 133 to 156.

161. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 133 to 156.

162. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And Apparatus for performing any of the methods according to claims 133 to 156.

163. A method comprising: At a computer system in communication with one or more input devices and a display generation component: Simultaneously display a first virtual object and a second virtual object in a three-dimensional environment visible via the display generation component; While simultaneously displaying the first virtual object and the second virtual object via the display generation component, detect a first input via the one or more input devices, the first input including a request to move the first virtual object relative to the second virtual object; and In response to detecting the first input: Move the first virtual object relative to the second virtual object according to the first input; Reduce the opacity of a corresponding portion of the second virtual object based on determining that moving the first virtual object causes the current position of the first virtual object to overlap the current position of the second virtual object relative to the viewpoint of a user of the computer system, while the first virtual object is further from the viewpoint of the user than the second virtual object.

164. The method according to claim 163, further comprising: In response to detecting the first input: Abstain from reducing the opacity of the corresponding portion of the second virtual object based on determining that moving the first virtual object does not cause the current position of the first virtual object to overlap the current position of the second virtual object relative to the viewpoint of the user.

165. The method according to claim 163, further comprising: In response to detecting the first input: Abstain from reducing the opacity of the corresponding portion of the second virtual object based on determining that moving the first virtual object causes the current position of the first virtual object to overlap the current position of the second virtual object relative to the viewpoint of the user, while the first virtual object is closer to the viewpoint of the user than the second virtual object.

166. The method according to any one of claims 163 to 165, wherein an overlap between the first virtual object and the second virtual object when the first virtual object is further from the viewpoint of the user than the second virtual object includes a first degree of overlap relative to the viewpoint of the user, the method further comprising: While the current position of the first virtual object overlaps the current position of the second virtual object relative to the viewpoint of the user, and the first virtual object is further from the viewpoint of the user than the second virtual object and the opacity of the corresponding portion of the second virtual object is reduced: Detect a second input via the one or more input devices to move the first virtual object relative to the second virtual object; In response to detecting the second input: Move the first virtual object according to the second input; And Based on determining that the movement of the first virtual object causes the current position of the first virtual object to overlap with the current position of the second virtual object relative to the user's viewpoint, including a second degree of overlap relative to the user's viewpoint, and while the first virtual object is further away from the user's viewpoint than the second virtual object, reducing the opacity of an additional portion of the second virtual object that is different from the corresponding portion of the second virtual object.

167. The method according to any one of claims 163 to 166, wherein: In response to detecting the first input, and based on determining that moving the first virtual object causes the current position of the first virtual object to overlap with the current position of the second virtual object relative to the user's viewpoint, while the first virtual object is further away from the user's viewpoint than the second virtual object: Based on determining that the distance between the first virtual object and the second virtual object relative to the user's viewpoint is a first distance, the corresponding portion of the second virtual object has a first size relative to the three-dimensional environment; and Based on determining that the distance between the first virtual object and the second virtual object relative to the user's viewpoint is a second distance different from the first distance, the corresponding portion of the second virtual object has a second size different from the first size relative to the three-dimensional environment.

168. The method according to claim 163, wherein: The corresponding portion of the second virtual object includes a first portion and a second portion different from the first portion, The first portion corresponds to the visual overlap region between the first virtual object and the second virtual object relative to the user's viewpoint, and The second portion corresponds to the region around the first portion of the second virtual object relative to the user's viewpoint.

169. The method according to claim 168, wherein: Based on determining that when the first virtual object overlaps with the current position of the second virtual object relative to the user's viewpoint, the first virtual object is at a first distance from the second virtual object, the second portion of the second virtual object extends beyond the boundary of the first virtual object by a first amount, and Based on determining that when the first virtual object overlaps with the current position of the second virtual object relative to the user's viewpoint, the first virtual object is at a second distance different from the first distance from the second virtual object, the second portion of the second virtual object extends beyond the boundary of the first virtual object by a second amount different from the first amount.

170. The method according to any one of claims 163 to 169, further comprising: During the first input being performed, detecting the termination of the first input; And In response to detecting the termination of the first input: Add the first virtual object to the second virtual object based on determining that the current position of the first virtual object is within a threshold distance of the current position of the second virtual object when the first input terminates.

171. The method according to claim 170, further comprising: In response to detecting the termination of the first input, based on determining that the current position of the first virtual object is not within the threshold distance of the current position of the second virtual object when the first input terminates and the current position of the first virtual object is closer to the user's viewpoint than the current position of the second virtual object, refrain from adding the first virtual object to the second virtual object.

172. The method according to any one of claims 170 to 171, further comprising: In response to detecting the termination of the first input, based on determining that the current position of the first virtual object is farther from the user's viewpoint than the current position of the second virtual object, refrain from adding the first virtual object to the second virtual object.

173. The method according to any one of claims 163 to 172, wherein when the first virtual object and the second virtual object are simultaneously displayed, a first portion of the first virtual object is displayed at a first level of visual prominence relative to the three-dimensional environment, and a second portion of the second virtual object that is different from the corresponding portion of the second virtual object is displayed at a second level of visual prominence relative to the three-dimensional environment, the method further comprising: When the first portion of the first virtual object and the second portion of the second virtual object are simultaneously displayed at the second level of visual prominence, and when the first input is detected and the movement of the first virtual object satisfies one or more criteria: Based on determining that the current position of the first virtual object is within the threshold distance of the second virtual object and the current position of the first virtual object is closer to the user's viewpoint than the current position of the second virtual object, display the second portion of the second virtual object at a third level of visual prominence greater than the second level of visual prominence; And Based on determining that the current position of the first virtual object is farther from the user's viewpoint than the current position of the second virtual object, refrain from displaying the second portion of the second virtual object at the third level of visual prominence.

174. The method according to any one of claims 163 to 173, further comprising: When reducing the opacity of the corresponding portion of the second virtual object based on determining that the current position of the first virtual object is farther from the user's viewpoint than the current position of the second virtual object in response to detecting the first input, detect the termination of the first input; And In response to detecting the termination of the first input, and based on determining that the current position of the first virtual object is farther from the user's viewpoint than the current position of the second virtual object and the current position of the first virtual object causes the first virtual object to overlap the current position of the second virtual object with respect to the user's viewpoint, increase the opacity of the corresponding portion of the second virtual object.

175. The method according to any one of claims 163 to 174, further comprising: In response to detecting the first input, and when moving the first virtual object relative to the second virtual object according to the first input: Based on determining that the requested movement of the first virtual object includes a request to move the current position of the first virtual object from a position in front of the second virtual object as viewed from the user's viewpoint by a first amount through the second virtual object to a position behind the second virtual object as viewed from the user's viewpoint, move the first virtual object by a second amount; And Based on determining that the requested movement of the first virtual object includes a request to move the current position of the first virtual object from a position behind the second virtual object as viewed from the user's viewpoint by the first amount through the second virtual object to a position in front of the second virtual object as viewed from the user's viewpoint, move the first virtual object by a third amount less than the second amount.

176. The method according to claim 175, wherein, as viewed from the user's viewpoint, prevent the first virtual object from moving through the second virtual object in a direction from in front of the second virtual object to behind the second virtual object.

177. The method according to any one of claims 175 to 176, wherein moving the first virtual object from the position in front of the second virtual object to the position behind the second virtual object with respect to the user's viewpoint includes: Based on determining that the requested movement of the first virtual object corresponds to a speed of movement of the first virtual object greater than a threshold speed, when the first virtual object is within a threshold distance of the second virtual object, move the first virtual object through the second virtual object without snapping the first virtual object to the second virtual object, and based on determining that the requested movement of the first virtual object corresponds to a speed of movement of the first virtual object less than the threshold speed, when the first virtual object is within the threshold distance of the second virtual object, snap the first virtual object to the second virtual object while moving the first virtual object through the second virtual object.

178. The method according to claim 177, wherein moving the first virtual object from the position behind the second virtual object to the position in front of the second virtual object includes snapping the first virtual object to the second virtual object when the first virtual object is within the threshold distance of the second virtual object.

179. The method according to any one of claims 177 to 178, wherein the speed of the movement of the first virtual object is the average speed of the movement of the first virtual object.

180. The method according to any one of claims 163 to 179, further comprising: While moving the first virtual object relative to the second virtual object according to the first input: Snapping the first virtual object to the second part of the second virtual object in the three-dimensional environment according to determining that the current position of the first virtual object is within the threshold distance of the second part of the second virtual object displayed at a visual saliency level greater than the threshold visual saliency level with respect to the three-dimensional environment; And Giving up snapping the first virtual object to the second part of the second virtual object according to determining that the current position of the first virtual object is within the threshold distance of the position corresponding to the second part of the second virtual object, but the second part of the second virtual object is not displayed at a visual saliency level greater than the threshold visual saliency level with respect to the three-dimensional environment.

181. The method according to claim 180, further comprising: When displaying the second virtual object: Displaying the second part of the second virtual object at a first visual saliency level less than the threshold visual saliency level according to determining that the first part of the corresponding virtual object has a position corresponding to the part of the three-dimensional environment that is the same as the second part of the second virtual object; And Displaying the second part of the second virtual object at a second visual saliency level greater than the threshold visual saliency level according to determining that no object has the position corresponding to the part of the three-dimensional environment that is the same as the second part of the second virtual object.

182. The method according to any one of claims 180 to 181, further comprising: When displaying the second virtual object: Displaying the second part of the second virtual object at a first visual saliency level less than the threshold visual saliency level with respect to the three-dimensional environment according to determining that the angle of view between the user's viewpoint and the corresponding angle of view of the second virtual object is greater than the threshold angle; And Displaying the second part of the second virtual object at a second visual saliency level greater than the threshold visual saliency level with respect to the three-dimensional environment according to determining that the angle of view between the user's viewpoint and the corresponding angle of view of the second virtual object is less than or equal to the threshold angle.

183. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: One or more processors; Memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for: Simultaneously displaying a first virtual object and a second virtual object in a three-dimensional environment visible via the display generating component; While simultaneously displaying the first virtual object and the second virtual object via the display generating component, detecting a first input including a request to move the first virtual object relative to the second virtual object via the one or more input devices; and In response to detecting the first input: Moving the first virtual object relative to the second virtual object according to the first input; Reducing the opacity of a corresponding portion of the second virtual object based on determining that moving the first virtual object causes the current position of the first virtual object to overlap with the current position of the second virtual object relative to the user's viewing point of the computer system, while the first virtual object is farther from the user's viewing point than the second virtual object.

184. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generating component and one or more input devices, cause the computer system to perform a method including: Simultaneously displaying a first virtual object and a second virtual object in a three-dimensional environment visible via the display generating component; While simultaneously displaying the first virtual object and the second virtual object via the display generating component, detecting a first input including a request to move the first virtual object relative to the second virtual object via the one or more input devices; and In response to detecting the first input: Moving the first virtual object relative to the second virtual object according to the first input; Reducing the opacity of a corresponding portion of the second virtual object based on determining that moving the first virtual object causes the current position of the first virtual object to overlap with the current position of the second virtual object relative to the user's viewing point of the computer system, while the first virtual object is farther from the user's viewing point than the second virtual object.

185. A computer system in communication with a display generating component and one or more input devices, the computer system including: One or more processors; Memory; Means for: simultaneously displaying a first virtual object and a second virtual object in a three-dimensional environment visible via the display generating component; Means for: while simultaneously displaying the first virtual object and the second virtual object via the display generating component, detecting a first input including a request to move the first virtual object relative to the second virtual object via the one or more input devices; And Means for: in response to detecting the first input: Move the first virtual object relative to the second virtual object according to the first input; Reduce the opacity of the corresponding portion of the second virtual object based on determining that moving the first virtual object causes the current position of the first virtual object to overlap the current position of the second virtual object relative to the viewpoint of a user of the computer system, while the first virtual object is further from the viewpoint of the user than the second virtual object.

186. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods according to claims 163 to 182.

187. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform any of the methods according to claims 163 to 182.

188. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: One or more processors; A memory; And Means for performing any of the methods according to claims 163 to 182.